Application Specific, Automatic Distributed Evaluation of Performance Data on the Grid

Hamza Mehammed · mediaTUM – the media and publications repository of the Technical University Munich (Technical University Munich) · 2009

The exploitation of distributed evaluation of data in parallel tools has been one of the main concerns of tool developers in their effort to meet the ever-increasing demand on scalability of online parallel tools.Recently, the inherent complexity of parallel machines is enlarged with the sharing, transparency and dynamicity of the available resources in the environment like the Grid introducing a new challenge on performance analysis.In a grid computing environment which is devoted for a coordinated utilization of geographically distributed large parallel computing resources, most of the available parallel tools (like performance analyzers, debuggers and load balancers), which relied on a centralized evaluation of data, often fail to scale well since the centralized evaluation in such modestly sized computing environment mostly leads to a bottleneck at the front-end.Specially, in on-line parallel tools, the results computed at the back-ends must all arrive at the front-end without any loss of data in order to be evaluated properly.This is hard to achieve when the number of back-ends increases and/or they produce computed data more frequently.This thesis addresses a novel method of automatic, distributed evaluation of performance data in order to decrease the frequency and the amount of data transferred from the back-ends to the front-end to ensure the scalability of the tool.The main research topic is the automatic generation of application specific virtual networks used for the distribution of the subtasks that can be computed on a given location independently.Building such virtual networks depend merely on the metrics specification provided by the users.These specifications are used to describe applicable objects and partner objects (sites, hosts and processes), code regions (portion of source code or location of certain methods) and time specifications (time interval, a point in time or virtual time) which are going to be used for a computation to enhance a very flexible multicast reduction network.For the realization of the virtual networks, an augmented dataflow model based on the classical dataflow model is developed which supports a hierarchical parallel execution of subtasks and enables also the reassembly of the measurement results asynchronously.The reassembly process includes correlation, aggregation and synchronization of the result data values.Through its hierarchical nature, the virtual network allows the evaluation of data at their origin (instead of transferring them to a central location) by taking the abstraction level of the defined applicable objects into consideration.The feasibility of the distributed evaluation methodology is performed in a real parallel tool called Grid Performance Measurement Tool (GPM).In GPM, the number of communications between the front-end and the back-ends increases rapidly since the performance measurements are resolved in time, location in the system and location in the code.Using the automatic, distributed evaluation of the performance data, a well manageable front-end is achieved which increases the scalability of the parallel tools.This thesis wouldn't have been possible without the assistance of many people to whom I would like to express my sincere appreciation.First of all, I would like to express my special gratitude to Prof. Dr. Arndt Bode for providing me with an excellent research environment to work on my thesis.I am

Read the paper · More papers on PaperTik