Data Trasfer From MySQL To Hadoop
Sanjeev Kumar Pippal, Shiv Pratap Singh, Dharmender Singh Kushwaha · 2014
Processes use on-hand database management tools or traditional data processing applications. The challenges include capture, curation, storage, search, transfer and analysis. Data sets grow in size by gathering logs on servers, cameras and wireless sensor networks etc. Every day approximately 2.5 quintillion bytes of data is created and nearly 90% of the data is created in last two-three years. Hadoop is designed to scale up from single servers to thousands of machines, each offering local computation and storage. It is the core platform for structuring Big Data, and making it useful for solving problems related to analytic purposes. HDFS links together file systems on many local nodes to make them into one big file system. It assumes nodes will fail, so it achieves reliability by replicating data across multiple nodes. This paper proposes a way to make a cluster on low configuration machines and uses it efficiently for data processing and storage. This implementation integrates open source tool named Sqoop for importing data from MySQL server to cluster of Hadoop and store data in distributed file systems across various nodes. The results establish that as the number of Map-tasks increases Hive import time and Loading time decreases by 20% and over 100 % respectively.. The results show that as the number of Map-tasks increase transfer rate improves by over 27% but at the cost of failed/killed tasks. It happens because as the number of tasks increase, resources start getting exhausted and tasks start getting failed or killed.