Integrated Network Performance Diagnostics

Srikanth Kandula, Anees Shaikh, Erich Nahum · 2002

One of the most challenging parts of running a networkbased service is monitoring and managing performance. End-to-end performance may be influenced by numerous factors, and problems are not easily attributable to their correct source [1]. When a performance problem arises, the service provider or customer has available a number of tools to aid in diagnosing performance issues, each of which test different aspects of the network service. Client-perceived performance, for example, requires the use of application-layer tests to measure the application response time. Such tests can be conducted for Webbased applications using tools like PageDetailer [2] or services provided by companies such as Keynote Systems [3]. Low-level network properties such as connectivity and latency may be tested using tools like traceroute or ping. Network protocol behavior can be examined in detail using packet sniffing tools, as exemplified by tcpdump. Loss rate and bandwidth along paths can be measured using tools such as sting [4] and pathrate [5], respectively. While these low-level tools are useful for detecting relatively simple conditions, such as a server being unavailable, or the absence of a network path to the destination, it is not straightforward to directly relate the information from such tools to application behavior. A common performance management approach involves monitoring the application (i.e., user-perceived) performance at a relatively coarse level, and then conducting further, more detailed, tests when a potential problem is detected. As discussed above, such an approach requires the use of multiple techniques and tools, operating at multiple levels. Hence, the problem remains of how to correlate the information to construct a more complete view of low-level network events (e.g., packet retransmission) and the application actions that triggered them (e.g., HTTP request). In this abstract, we focus on the problem of integrating performance-related information from multiple network layers for the purpose of network performance diagnosis. Our approach is to provide application-level measurement tools with direct access to pertinent information from lower layers. This enables a diagnostician to detect specific packet-level events in various contexts of the application in an automated way, without having to unify multiple traces. Our initial objective is to embody this approach in a measurement and monitoring tool for identifying the causes of performance problems in Web-based applications. Below we describe our approach in further detail, with a discussion of our initial design and ongoing implementation. We also list some of the advantages and limitations of this scheme.

Read the paper · More papers on PaperTik