Performance of the ETL processes in terms of volume and velocity in the cloud: State of the art
Papa Senghane Diouf, Aliou Boly, Samba Ndiaye · 2017
The ETL (Extract-Transform-Load) consists of extracting data from various sources, transforming and loading them into a place called datawarehouse. ETL is a mandatory step in the projects which implement decision-making information systems or knowledge management systems within organizations. But it is also a long and costly step in the use of human and IT resources. However, in the context of big data, characterized by 4V (Variety, Velocity, Volume and Veracity), the speed of processing has become a decisive factor in search of competitiveness. In order to facilitate the implementation of the ETL the solution is then to use the infrastructures of cloud computing whose resources in computation and storage are unlimited. This has resulted in considerable progress in terms of availability and scalability for the success of projects. But it remains a major problem: the cost can quickly become prohibitive with “pay-per-use” model of the cloud. So, in this case, how to find ETL solutions built on the cloud at a lower cost? A great deal of suggestions have been made. In this article, we have reviewed these works by highlighting the performance aspects of data processing in terms of volume and velocity.