Automatic Tuning and Software Modularization Techniques for Optimized Large Scale Data Analytics
Miyuru Dayarathna, Miyuru Dayarathna · Institutional Repositories DataBase (IRDB) · 2013
Large scale systems such as Exascale computing systems will be the key compute resources that enable high-end discoveries in the next half century of sciences.Data intensive large scale software systems need optimized software architectures focusing on concurrency and resiliency.Currently MPI, MapReduce, and stream programming models are famous for data intensive tasks.However, they are less optimal in terms of programmability, maintainability, and execution performance which are fundamental problems that need to be addressed in future systems.It is a technically challenging problem because it is very hard to achieve a single generic solution in variety of systems and programming environments.This thesis proposes new software modularization, automatic code transformation, and benchmarking techniques as solutions for improving productivity and execution performance of large scale data analysis systems.Specifically, it focuses on optimizing the current state of the art stream/graph data processing paradigms via the aforesaid techniques.First half of the thesis studies about automated techniques for performance improvement of stream processing systems.The thesis investigates on methods for workload characterization and benchmarking of large scale data analytics.The thesis evaluates performance of stream processing system architectures using System S, S4, and Esper.It introduces operator graph transformation, a notion for improving the concurrency of stream analytic tasks via automatic software tuning.The techniques were implemented on optimization frameworks called Hirundo, Tahitica, and Albatross.The second half of the thesis describes graph analytics optimization.In order to understand performance features of graph database systems and also to create realistic benchmark framework for graph database systems, the thesis implemented XGDBench following the Multiplicative Attribute Graph model.The thesis describes the importance of having well-defined abstractions for batch/hybrid execution of large graph analysis tasks.Especially it presents optimized software abstractions, libraries, and middleware that can withstand software architectural challenges introduced by heterogeneous hardware.It created the architecture of ScaleGraph 1.0, a Partitioned Global Address Space (PGAS) library for analyzing graphs in the scale of billions of vertices and edges.Furthermore, it implemented a domain specific language for large graph analysis called Exerda, and showed how such languages can support specification of optimized graph data analysis tasks efficiently.Thesis made contributions to large scale stream computing systems by introducing an automatic optimization mechanism for data stream programs.It created a scalable data generator (StreamFarm) and a benchmark (XGDBench) for supporting performance optimizations.Thesis contributed for the area of large scale graph analysis by introducing ScaleGraph library and Exedra which is a domain specific language for defining large graph analysis processes.Thesis identifies future expansions in the areas of high productive software development, massive graph analysis, and software maintainability.Thesis describes the software architectural aspects of Hirundo, Albatross, XGDBench, ScaleGraph, and Exedra focusing on their contributions for performance optimization of large scale data analytics.