Leveraging Queueing Theory and OS Profiling to Reduce Application Latency
Anshul Gandhi, Amoghvarsha Suresh · 2019
Request latency is a critical metric in determining the usability of online services, such as web applications and databases. Most existing approaches to improve application latency start by detecting bottlenecks in the application deployment; this typically entails determining stages of processing where the application spends most of its time. Inspired by queueing theory, we present an alternative approach to detect and mitigate bottlenecks âĂrŞ using variability of processing time as a guiding principle when designing applications.