Towards hybrid programming in big data

Peng Wang, Hong Jiang, Xu Liu, Jizhong Han · 2015

Within the past decade, there have been a number of par-allel programming models developed for data-intensive (i.e., big data) applications. Typically, each model has its own strengths in performance or programmability for some kinds of applications but limitations for others. As a result, multiple programming models are often com-bined in a complimentary manner to exploit their mer-its and hide their weaknesses. However, existing models can only be loosely coupled due to their isolated runtime systems. In this paper, we present Transformer, the first sys-tem that supports hybrid programming models for data-intensive applications. Transformer has two unique con-tributions. First, Transformer offers a programming abstraction in a unified runtime system for different programming model implementations, such as Dryad, Spark, Pregel, and PowerGraph. Second, Transformer supports an efficient and transparent data sharing mech-anism, which tightly integrates different programming models in a single program. Experimental results on Amazon’s EC2 cloud show that Transformer can flexi-bly and efficiently support hybrid programming models for data-intensive computing. 1

Read the paper · More papers on PaperTik