Showing posts with label performance. Show all posts
Showing posts with label performance. Show all posts

Monday, September 23, 2019

Avoid full table scan in MySQL

Stats

SHOW GLOBAL STATUS LIKE 'Select%';





Counter Select_scan shows how many full table scans were done since last MySQL restart.
Counter Select_full_join is even worse as MySQL has to perform a full table scan against a joined table which is even slower.
With that being said, we need to try the best to avoid full table scan when writing queries. 

Use index

Apart from PK and foreign keys, add index to columns
  • Columns frequently used to join tables
  • Columns that are frequently used as conditions in a query
  • Columns that have a high percentage of unique values
Without index on the column appears in where clause or sort by, MySQL will walk through the entire table to filter rows one-by-one.

Best practice

Avoid using function or math

SELECT * FROM table WHERE func(a) = 100
SELECT * FROM table WHERE a + 3 < 100

Avoid using Not equal and NOT IN

SELECT * FROM table WHERE a <> 1
SELECT * FROM table WHERE a NOT IN (1,2,3)

Avoid Bitwise on numeric column

SELECT * FROM table WHERE (a & 4) = 0

Avoid putting a wild-card before the first characters of the search criteria

SELECT * FROM table WHERE a LIKE '%abc'

Avoid the OR Operator

SELECT * FROM table WHERE a = 1 OR a = 2 OR a = 3
Try to replace it with an IN operator, something like SELECT * FROM table WHERE a IN (1,2,3)

Avoid using Having

Avoid using Order by if possible

Avoid using Group by if possible

Avoid using DISTINCT if possible

Avoid using ORDER BY RAND()

Avoid SELECT COUNT(*) FROM table

InnoDB doing a full table scan for this statement.

Tuesday, March 19, 2019

APM (Application Performance Management)

The heart of APM solutions is understanding why transactions in your application are slow or failing.

10 Critical Application Performance Management Features for Developers
  1. Performance of every web request and transaction
  2. Code level performance profiling
  3. Usage and performance of all application dependencies like databases, web services, caching, etc (Distributed Tracing)
  4. Detailed traces of individual web requests or transactions
  5. Basic server monitoring and metrics like CPU, memory, etc
  6. Application framework metrics like performance counters, JMX MBeans, etc
  7. Custom applications metrics created by the developer team or business (tracking, reporting, alerting)
  8. Application log data
  9. Application errors
  10. Real user monitoring (RUM)
Here's a list of New Relic Alternatives/Replacements of 2019:
  1. AppOptics
  2. Pingdom Server Monitor (formerly Scout App)
  3. Stackify Retrace
  4. DynaTrace (formerly Ruxit)
  5. Atatus
  6. Datadog
  7. LogicMonitor
  8. AppDynamics
  9. Elastic APM
For me, I care about where is slow and why it is slow with prompt alerting support.

Tuesday, May 8, 2012

Increase initcwnd for TCP Performance

MTU & MSS
MTU (maximum transmission unit): Nearly all IP over Ethernet implementations use the Ethernet V2 frame format, which is 1500 bytes. Linux ifconfig output can show MTU size.

MSS (The maximum segment size)  is a parameter of the TCP protocol that specifies the largest amount of data, specified in octets, that a computer or communications device can receive in a single TCP segment, and therefore in a single IP datagram. It does not count the TCP header or the IP header.

Therefore, TCP/IP Headers + MSS ≤ MTU
MTU = 1500
TCP Header = 20
IP Header = 20
TCP Option = 12 (optional)
MSS = 1460 (1448 if there is TCP Option)

Two TCP Windows
Congestion Window (cwnd) controls the number of packets a TCP flow may have in the network in any given time. cwnd is dynamically adapting to changing network condition. TCP Slow-start is one of the algorithms that TCP uses in its quest to control congestion inside the network and it is also known as the exponential growth phase.When TCP reaches a certain threshold (also known as sstrsesh) it will enter the linear growth, Congestion avoidance. Linux 2.6.39 increased the initial congestion window to 10 packets, previous versions are 3.

The Receiver Advertised Window (rwnd) is the buffer size sent in each ACK from TCP receiver to TCP sender. The window size is 65535 (64K) bytes on Windows/Mac/iOS.

The purpose of sliding window is to prevent from the sender to send too many packets to over flow the network resource or the receiver's buffer. The "sliding window size" is the maximum amount of data we can send without having to wait for ACK.

Therefore, with suggested 10 initcwnd size, sliding window can have 10 * MSS data flowing the network without ACK, it will definitely improve TCP performance, eliminating TCP slow start.

How to Change initcwnd/initrwnd
ip route show
sudo ip route change default via 192.168.1.1 dev eth0 initcwnd 10
ip route show

ip route show
sudo ip route change default via 192.168.1.1 dev eth0 initrwnd 10
ip route show

Notes:
1. The advertised receive window on Linux is called initrwnd. It can only be adjusted on linux kernel 2.6.33 and newer
2. This changes the initcwnd and initrwnd until the next reboot.
3. To persist the changes, try out below script (copied from cdnplanet.com, but not verified)
cp /etc/sysconfig/network-scripts/ifup-post /etc/sysconfig/network-scripts/ifup-post.bak; sed -i -e "/^exit 0/d" /etc/sysconfig/network-scripts/ifup-post; echo "ip route change " $(ip route show | grep '^default' | sed 's/initcwnd [0-9]*//') " initcwnd 10" >> /etc/sysconfig/network-scripts/ifup-post; echo "exit 0" >> /etc/sysconfig/network-scripts/ifup-post

Other Tunings
disable net.ipv4.tcp_slow_start_after_idle
setting tcp_slow_start_after_idle to 0 (for disabling it) to speed up initial connections, otherwise it will cause your keepalive connection to return to slow start after TCP_TIMEOUT_INIT (3 seconds)

disable Nagle algorithm
TCP implementations usually provide applications with an interface to disable the Nagle algorithm. This is typically called the TCP_NODELAY option.

Other options on CentOS
sysctl -w net.core.rmem_max=16777216
sysctl -w net.core.wmem_max=16777216
sysctl -w net.ipv4.tcp_rmem=4096 87380 16777216
sysctl -w net.ipv4.tcp_wmem=4096 87380 16777216
sysctl -w net.core.netdev_max_backlog=30000
sysctl -w net.ipv4.tcp_congestion_control=htcp

Reference
http://kernelnewbies.org/Linux_2_6_39
http://www.cdnplanet.com/blog/tune-tcp-initcwnd-for-optimum-performance/
http://monolight.cc/2010/12/increasing-tcp-initial-congestion-window/
http://www.osischool.com/protocol/Tcp/slidingWindow/index.php

Friday, September 23, 2011

jQuery 10 performance tips

This is a summary of 10 jQuery performance tips from http://addyosmani.com/jqprovenperformance/
  1. Use the latest jquery release
  2. Know your selectors (ID, element, class, pseudo & attribute)
  3. Use .find() -> $parent.find('.child').show() is the fastest than others (scoped selector)
  4. Don't use jQuery unless it's absolutely necessary -> this.id is fater than $(this).attr('id')
  5. Caching -> storing the result of a selection for later re-use, no repeat selection
  6. Chaining
  7. Event delegation -> if possible, use delegate instead of bind and live
  8. Each DOM insertion is costly -> keep the use of .append(), .insertBefore() and .insertAfter() to a minimum; .data() is better than .text() or .html(); $.data('#elem', key, value) is faster than $('elem').data(key, value)
  9. Avoid loops -> Javascript for or while loop is faster than jQuery .each()
  10. Avoid constructing new jQuery object unless ncessary -> use $.method() rather than $.fn.method(); $.text($text) is faster than $text.text(), $.data() is faster than $().data
Update (11/30/2011):
In recent jQuery performance discussion, I summarized the following points based on different messages.
  1. Scope jquery selectors (use find)
  2. Know selectors performance => ID > element (tag) > class > pseudo & attribute
  3. Avoid loops =>  if loop is inevitable, consider for/while > $.each()  if possible
  4. Cache the result of a selection for later use (using local variables if possible)
  5. DOM operation is expensive (esp. in a loop)
  6. Don't use $ unless it's necessary => this.name > $(this).attr('name’) 

    Thursday, September 1, 2011

    Summary about High Performance Mobile Meetup

    This Tuesday (Aug 30, 2011) I attended the SF web performance meetup at LinkedIn (Mountain View). Steve Souders presented a great talk "High Performance Mobile". Apart from this year Velocity conference HttpArchive lighting demo, this is the second time I joined his session. I read twice his previous talk slide "High Performance HTML5" at SF performance meet up, but could not join in person. Of course, I borrowed and read his two famous books.

    The event started at 6:30pm, welcomed attendees with portable plastic water bottle and Mexican food. Steve began his talk around 7pm after some introduction from meetup organizer (Aaron Kulick) and LinkedIn performance lead (they are hiring performance engineers). The talk consisted of 4 parts, you can find details from http://www.slideshare.net/souders/high-performance-mobile-sfsv-web-perf

    Part 1: WPO
    Nothing new, but he reiterated the importance and role of WPO. The Web Is Dead (http://www.wired.com/magazine/2010/08/ff_webrip/all/1), will this be true and what is the fate of WPO? It might be too early to conclude "The web is dead". One thing is true, the web is evolving with HTML5.

    Benefits of WPO
    1. drives traffic
    2. improves UX
    3. increases revenue
    4. reduces costs
    Part 2: Why Mobile
    Data and analysis show that Mobile is very important, but slow. Also the "Road is not clear" for performance optimization

    Part 3: Mobile Best practices
    Most desktop Web performance rules still apply. Steve mainly shared 5 items to which he thinks Web performance engineers need pay more attention.
    1. Reduce Http request (sprite/dataURI/CSS3/Canvas) - this should be the golden rule
    2. Responsive images (sencha.io src/DeviceAtlas/adaptive-images.com) - previously read an article tweeted by Stoyan, have some general idea about this
    3. script async & defer (execute when available, execute when parsing finished. He also mentioned his controlJS) - heard this many times, are all new browsers supporting them? should we include javascript using async & defer
    4. Appcache (5M+ limit) - from HTML5
    5. Local storage (window.localStorage) - from HTML5
    Part 4: Mobile tools
    The impressive one of this part was his demo after he briefed following 4 tools. Steve also mentioned his bookmarklet for mobile performance but he didn't add into his deck.
    1. pacpperf
    2. jdrop
    3. blaze.io
    4. weinre (WEb INspector REmote)
    After that, there was book giveaway and e-book lottery. Somehow I didn't have a number, so left before venue was clear up. (I hope I can get a signed copy of his book)

    Key takeaways
    1. Mobile is important but very slow
    2. There are challenges to make mobile fast
    3. There are tools to assist mobile performance
    4. Mobile winners will be fast
    Thanks Steve/Aaron/Sponsors to make this meetup so successful. Looking forward to next meetup in south bay area.

    Different Testings

    Testing is an art than science. There are different testings for different purposes to ensure the software quality. The goal is same, but the process, strategy and methodology are different from different testings.

    Correctness testing
    Correctness is the minimum requirement of software, the essential purpose of testing. It is for software quality, also called function testing.

    Black-box testing
    test data are derived from the specified functional requirements without regard to the final program structure. It is also termed data-driven, input/output driven, or requirements-based testing.

    White-box testing
    the structure and flow of the software under test are visible to the tester.

    Performance testing
    a process that focuses on testing individual components of the web app, such as databases, algorithms, network infrastructure, and cache layers under certain load

    Load testing
    a process of determining how an application under specific volumes of load, usually a range of the upper and lower limits expected by the business. Endurance testing is also part of this testing type.

    Stress testing
    a process of identifying when and how systems fail (and recover) under extreme levels of load. Also known as negative testing or destructive testing.

    Reliability testing

    a process of finding the probability of failure-free operation of a system.

    Security testing
    identifying and removing software flaws that may potentially lead to security violations, and validating the effectiveness of security measures. Simulated security attacks can be performed to find vulnerabilities.

    I summarized above testings based on below excellent blogs and paper. It is interesting to understand these terms better while working with QA (quality assurance) team in daily work.
    http://agiletesting.blogspot.com/2005/02/performance-vs-load-vs-stress-testing.html
    http://agiletesting.blogspot.com/2005/04/more-on-performance-vs-load-testing.html
    http://blog.browsermob.com/2008/12/performance-vs-load-vs-stress-testing/
    http://www.ece.cmu.edu/~koopman/des_s99/sw_testing/

    Thursday, August 4, 2011

    Cache related Http Headers

    Http headers can instruct what kind of cache mechanism browser and proxy should obey along the request/response chain. For static resources (Javascript, CSS, images, flash etc), it is suggested to apply cache on browser or proxy side to reduce # of Http requests.

    Response Headers:
    Cache-Control     Tells all caching mechanisms from server to client whether they may cache this object (public/private to control if browser or proxy cache, no-store to control if save to disk, max-age is to control how long to cache)
    Expires     Gives the date/time after which the response is considered stale (suggested 1 year from now, for aggressive static resource cache)
    Date     The date and time that the message was sent (It is useful for Expires by date time)
    ETag     An identifier for a specific version of a resource, often a message digest (Suggest to disable it for performance, or re-configure ETag to remove server specific info. See If-None-Match. ETag takes precedence over Last-Modified if both exist)
    Last-Modified     The last modified date for the requested object, in RFC 2822 format (for conditional get, 304 Not Modified, see If-Modified-Since)
    Pragma     Implementation-specific headers that may have various effects anywhere along the request-response chain. (Http1.0, example is Pragma: no-cache)
    Vary     Tells downstream proxies how to match future request headers to decide whether the cached response can be used rather than requesting a fresh one from the origin server. (The most common case is to set Vary: Accept-Encoding, so that proxy knows if return cached compressed data to browser)

    Request Headers:
    Cache-Control     Used to specify directives that MUST be obeyed by all caching mechanisms along the request/response chain
    If-Modified-Since     Allows a 304 Not Modified to be returned if content is unchanged
    If-None-Match     Allows a 304 Not Modified to be returned if content is unchanged
    Pragma     Implementation-specific headers that may have various effects anywhere along the request-response chain

    Recommendations:
    It is important to specify one of Expires or Cache-Control max-age, and one of Last-Modified or ETag, for all cacheable resources. It is redundant to specify both Expires and Cache-Control: max-age, or to specify both Last-Modified and ETag.

    You use the Cache-control: public header to indicate that a resource can be cached by public web proxies in addition to the browser that issued the request.

    Avoiding caching
    HTTP version 1.1 -> Cache-Control: no-cache
    HTTP version 1.0 -> Setting the Expires  header field value to a time earlier than the response time

    Reference
    http://tools.ietf.org/html/rfc2616
    http://en.wikipedia.org/wiki/List_of_HTTP_header_fields
    http://code.google.com/speed/page-speed/docs/caching.html#LeverageBrowserCaching
    http://code.google.com/p/doctype/wiki/ArticleHttpCaching

    Tuesday, August 2, 2011

    Mobile Web Optimization Webinar Takeaway

    Last week I attended a webinar from Compuware talking about mobile web optimization using page speed. The talk was very clear and well organized. It first went through the importance of mobile web performance, then discussed key difference between mobile and desktop, then detailed page speed rules about mobile web.

    Browser is the entry point to mobile web
    The browser is becoming the integration platform
    The browser is becoming more complex
        - # of hosts per user transaction
        - Many RIA frameworks
        - Performance differences across devices

    Free performance tools
    Page Speed
    WebPagetest
    dynaTrace Ajax Edition

    Mobile web page load process
    Mobile channel establishment
    DNS lookup
    TCP connect
    Http request
    Parse & Layout (subrequests)

    Key differences between mobile and desktop
    Networks:
        round-trip time (High channel establishment time, lower RTT)
        bandwidth (3G vs Cable)
    Devices:
        CPU (JS execution times, layout times, 10x JS runtime cost, 1 ms per kb parsing)
        memory (more code/objects - more GC, more DOM, more memory)
    Interaction model
        (touch vs click, mobile click event with 300-500ms delay)

    Page Speed Rules
    Use an application cache (how about localstorage?)
    Defer JS parsing (how about deferring JS download?)
    make landing page redirects cacheable (Cache-control: private, max-age > 0) (How long? 301 and 302 are both cacheable?)
    prefer touch events (why not disable click event on mobile?)

    Reference:
    http://slidesha.re/qgC8n3
    http://www.slideshare.net/Gomez_Inc/optimizing-web-and-mobile-site-performance-using-page-speed

    90% line in JMeter aggregate report

    I considered 90% line is the response data 90% users will see, but recently realized this was wrong. JMeter online Help has detailed explanation about every item in Aggregate report.

    The 90% line tells you that 90% of the samples fell at or below that number. However, it is more meaningful than average in terms of SLA. We expect it within 2x of average time. That is, if average time is 500ms, we expect 90% line is less than 1000ms. Otherwise the system fluctuates a lot.

    Friday, July 29, 2011

    Takeaway from High Performance HTML5

    http://www.slideshare.net/souders/high-performance-html5-sf-html5-ug

    speed matters
    WPWG (web performance working group)
    Web Timing (Navigation/User/Resource timing/window.performance)
    prefetching, prerendering
    async & defer
    app cache (update is served on 2nd reload?!)
    localStorage (5M max) as cache
    ! @font-face

    Friday, October 29, 2010

    High performance web site - reading notes (4)

    11. Avoid Redirects
    Redirects hurt performance
    Response status code is 3xx for redirects (300-307, and 304 is for conditional Get)
    Redirects delays html doc, CSS impacts rendering, JS impacts rendering and parallel download
    Missing Trailing Slash
        Apache alias/mod_rewrite/DirectorySlash
        autoindexing
    Connecting Web Sites
    Tracking internal traffic - referer logging
    Tracking outbound traffic - beacon (http request contains tracking info in the URL)
    Prettier URL - Avoid redirect using Alias, mod_rewrite, DirectorySlash and directly linking

    12. Remove duplicate scripts
    Unnecessary HTTP requests
    Wasted JS execution
    Implement a script management module in templating system
    Script has a getVersion() function



    13. Configure (Avoid) ETags
    Entity tags - a mechanism that web servers and browsers use to validate cached components
    ETag is a string, must be quoted, introduced in Http 1.1
    If-Non-Match takes precedence over If-Modified-Since
    ETag is typically constructed using attributes that make them unique to a specific server hosting a Web site
    Apache ETag uses inode-size-timestampe, and FileETag directive removes inode
    IIS ETag uses Filetimestamp:ChangeNumber (# of configuration changes to IIS)

    14. Make ajax cacheable
    Web2.0, DHTML, Ajax
    Yahoo Mail caches ajax result
    Use packet sniffer to monitor active/passive Ajax requests

    Tuesday, October 26, 2010

    High performance web site - reading notes (3)

    6. Put (java)scripts at the bottom
    Parallel downloads
    Limiting parallel downloads to two per hostname is a guideline, new browsers expand to 4 or more for HTTP/1.1
    Scripts block download
    Use deferred scripts (DEFER attribute indicates the script does not contain document.write)

    7. Avoid CSS Expressions
    CSS expressions are a powerful and dangerous way to set CSS properties dynamically
    CSS expressions are evaluated more frequently than most people expect.
    One-Time Expressions
    Event handlers

    8. Make javascript and CSS external
    In raw terms, inline is faster, but we need consider three metrics (page views, empty cache vs.primed cache, and component reuse).
    post-onload download (document onload event,firebug highlights DOMContentLoaded, load events)

    9. Reduce DNS lookup
    Reduce the number of unique hostnames reduces the number of DNS lookups
    Reduce the number of unique hostnames reduce the amount of parallel downloading
    Use keep-alive to reuse an existing connection by voiding TCP/IP overhead

    10. Minify javascript
    Use minification instead of obfuscation (due to bugs, maintenance, debugging etc concerns)
    Minify javascript using JSMin or dojo compressor (shrinksafe)

    High performance web site - reading notes (2)

    1. Make Fewer HTTP Requests
    Image maps
    CSS Sprites
    Inline images (data: URL scheme)
        e.g. <img alt="red star" src="data:image/gif;base64,THE-BASE64-DATA-OF-IMAGE">
    Combined javascripts and stylesheets

    2. Use a Content Delivery Network
    Akamai
    Mirror Image
    Limelight
    SAVVIS (specialized in video content delivery)
    Use keynote.com or gomez.com to test geographic locations

    3. Add an Expires header
    Expires
    Cache-Control (max-age) which take precedence over Expires
    Apache mod_expires
    Empty cache vs. primed cache
    Last-Modified
    revving filenames (add build version number), don't use query string

    4. Gzip components
    Accept-Encoding (Content-Encoding in response)
    Image/PDF should not be gzipped (Gzip your scripts and stylesheets)
    Gzip reduce by about 70%
    Apache mod_gzip (mod_deflate)
    Proxy caching uses Vary header (e.g. Vary: Accept-Encoding,User-Agent)
    Update (5/17/2012)
    Compress the Embedded OpenType font files used by Internet Explorer. EOT is a binary format, but it is not natively compressed
    Compress favicon, while an image file, is not natively compressed

    5. Put stylesheets at the top
    Use Link instead of @import as @import rule causes unexpected ordering in how the components are downloaded
    FOUC = Flash of unstyled content
    Put stylesheets in the document HEAD using the LINK tag

    High performance web site - reading notes (1)

    Background:
    Somehow I was assigned a new task to investigate front-end performance for an important project, and majorly about web site performance. This is a hot topic in Web2.0 era, and I did join the Velocity 2010 conference this June @Santa Clara. Most sessions were all about Web site performance, and behind the scene how to make web pages faster while dealing with HTML/JS/CSS/Images/Flash etc old friends. However, I have not worked on this layer for years, and almost forgot how to write CSS/JS efficiently, so need pick up quickly by reading.

    Why High performance web site?
    Two main reason: Steve is the author of YSlow and once was Chief Performance Yahoo! to lead a team focusing on yahoo performance (yahoo also published the best practices), and now he is with Google for performance. The book is well organized and easy to read and understand. - I am preparing to read his second book "Even faster Web site" now.

    Top 14 rules:
    There are many rules (best practices) regarding Web site performance from Yahoo, Google or other companies. But in this book, Steve listed top 14 rules and explained the ins and outs of these rules with examples and case study. I will not repeat his points word by word here, but as a reading notes, I will write down key take away from each rule. Therefore, the notes might not be complete sentence, or without context, or hard to fully understand. If you are interested, get one copy and read Steve's original words.
    1. Make few Http requests
    2. Use a CDN
    3. Add an Expires header
    4. Gzip components
    5. Put stylesheets at the top
    6. Put (java)scripts at the bottom
    7. Avoid CSS Expressions
    8. Make javascript and CSS external
    9. Reduce DNS lookup
    10. Minify javascript
    11. Avoid Redirects
    12. Remove duplicate scripts
    13. Configure (avoid) ETags
    14. Make ajax cacheable

    Thursday, September 16, 2010

    Oracle redo log slows down application

    Symptom:
    J2EE application on tomcat server suddenly has high latency, and has very slow server side processing response, and all 150 threads are used up. Regular request even takes more than 30 seconds while in normal case it takes around 100ms.

    Production Info:
    1. SW: Oracle 10g RAC, Tomcat6.0, JDK1.6, CentOS4.4
    2. No production outage or HA failover/failback
    3. No stress test or peak load
    Root Cause:
    NFS mount point hung which in turn slowed the archiving the logs to the NFS mount point, so the redos were not getting archived fast enough, and caused the latency.