Efficient Storage Design and Query Scheduling for Improving Big Data Retrieval and Analytics

Zhuo Liu · 2015

With the perpetually increasing requirement and generation of digital data, the human being has been stepping into the Big Data era. To efficiently manage, retrieve and exploit such gigantic amount of data continuously generated by all individuals and organizations of the society, a rich set of efforts has been invested to develop high-performance, scalable and fault-tolerant data storage systems and analytics frameworks. Recently, flash-based solid state disks and byte-addressable non-volatile memories have been developed and introduced into computer system storage hier- archy for substituting traditional hard drives and DRAM due to the faster data accesses, higher density and energy-efficiency. Along with the trend, how to systematically integrate such cutting edge memory technologies for fast system data retrieval becomes a highly concerned issue. In addition, from the users’ point of view, some mission-critical scientific applications are suffering from inefficient I/O schemes, thus not able to fully utilize the underlying parallel storage systems. This fact makes the development of more efficient I/O methods appealing. Moreover, MapReduce has emerged as a powerful big data processing engine that supports large-scale complex analytics applications. Most of them are written in declarative query languages such as Hive and Pig Latin. Therefore, it requires efficient coordination of Hive compiler and Hadoop runtime for fast and fair big data analytics. This dissertation investigates the research challenges mentioned above and contributes effi- cient storage design, I/O methods and query scheduling for improving big data retrieval and ana- lytics. I firstly aim at addressing the I/O bottleneck issue in large-scale computers and data centers. Accordingly, in my first study, by leveraging the advanced features of cutting-edge non-volatile memories, I have presented and devised a Phase Change Memory (PCM)-based hybrid storage architecture, which provides efficient buffer management and novel wear leveling techniques, thus achieving highly improved data retrieval performance and at the same time solving the PCM’s

Read the paper · More papers on PaperTik