Building a version control system in the Hadoop HDFS

Tsozen Yeh, Tingyu Chien · 2018

The cloud computing has been widely used in recent years. It facilitates the realization of many cutting-edge studies including big data, Internet of Things, and many others. The success of cloud computing cannot be achieved without reliable cloud infrastructure to store and handle the enormous volume of data stored therein. It is common that the contents of individual data files consist of data inserted at different periods of times in the cloud environment. In other words, data files often have chronological versions of contents since their creation. Unfortunately, file contents could be contaminated by bad data or viruses resulting in errors during the course of data processing. It will be easier and faster for users to identify the cause of the error if they could examine and process prior versions of data files in question. Consequently, by keeping versions of data files, cloud systems could help users solve problems more rapidly when errors occur. Hadoop is literally one of the most popular platforms adopted in the community of cloud computing. We designed and implemented an efficient scheme in HDFS, the default file system in Hadoop, to automatically maintain versions of individual data files when changes made to them. As a result, our system can retrieve prior versions of data files and display discrepancy between versions to ameliorate the data management in cloud centers.

Read the paper · More papers on PaperTik