SETL: A scalable and high performance ETL system

Kunjian Sun, Yuqing Lan · 2012

In order to extract, transform and load large scale data from heterogeneous data sources into data warehouse efficiently, the SETL system is designed and implemented in this paper. By using PERL subroutine attribute and data partition, SETL can implement ETL job easily and perform ETL job efficiently, and the plug-in design makes SETL with high scalability, and the design that performing one ETL job in one ETL pipeline makes SETL with distribution environment support. For illustration, one ETL job example is utilized to show the scalability in designing ETL job and show the high efficiency in processing large scale data. Experiments prove that SETL can extract, transform, and load large scale data into data warehouse efficiently. The SETL system simplified the ETL job design and implementation and can deal with heterogeneous data sources flexibly. It is a light-weighted, scalable and high-performance ETL system.

Read the paper · More papers on PaperTik