Performance analysis and efficient processing of grid and scientific datasets on multi-core processors
Madhusudhan Govindaraju, Rajdeep Bhowmik · 2012
The microprocessor industry is rapidly moving towards chip multi-processors (CMPs), where multiple cores can independently execute different threads. This change in computer architecture requires corresponding design modifications in programming paradigms, including grid and scientific middleware tools and applications, to harness the opportunities presented by multi-core processors. Naive implementations of grid and scientific middleware and applications on multi-core systems can severely impact performance because of limitations of shared bus bandwidth, cache size and coherency, and communication between threads. Therefore, it is important to design and develop middleware libraries and applications to harness the opportunities presented by emerging multi-core processors that are available on grid and cloud environments; those not adhering or adapting to this programming paradigm can suffer from severe performance limitations. In this thesis, we present performance results and analysis for processing XML based data on emerging multi-core systems. The thesis also presents analysis of various processing and scheduling techniques on multi-core architectures based on HDF5 based scientific data characteristics and access patterns. We also focus on the utilization of the L2 cache, a critical shared resource on CMPs. The access pattern of the shared L2 cache, which is dependent on how the application schedules and assigns processing work to each thread, can either enhance or undermine the ability to hide memory latency on a multi-core processor. While processing grid and scientific application datasets, it is essential to conduct fine-grained analysis of cache utilization to make informed scheduling decisions in multi-threaded programming. Using the McGrid framework and TAU toolkit for performance feedback, we present performance analysis and recommendations on how processing threads can be scheduled on multi-core nodes to enhance the performance of a class of grid middleware and scientific applications that requires processing of XML and HDF5 datasets respectively. We discuss the gains associated with the use of Cache Affinity and Balanced-Set based scheduling algorithms to improve L2 cache performance, and hence the overall application execution time. We also present a dynamic marking scheme that keeps track of the progress of threads on each core to determine work allocation and optimized usage of L2 cache for HDF5 dataset processing.