Enhancing Hadoop System Dependability Through Autonomous Snapshot

Tsozen Yeh, Yipin Wang · 2018

The cloud computing has successfully facilitated many cutting-edge studies including Internet of Things, Big Data, and many others in recent years. The accomplishment of cloud computing cannot be realized without reliable infrastructure to store and maintain the colossal amount of data stored therein. The integrity of data stored on cloud systems could be compromised by viruses or human errors. As a result, it will be desirable for the cloud system to recover its file system to its prior states to enhance its dependability. Hadoop is one of the most popular cloud platforms utilized in the community of cloud computing. Its default file system Hadoop Distributed File System (HDFS) provides a tool, Snapshot, to take point-in-time copies of the entire or parts of the file system for system recovery in the future. Ideally, a new snapshot should be generated automatically when there are changes made to parts of the file system covered in snapshots so the system can reinstate those parts to any one of their prior states. Unfortunately, the HDFS Snapshot needs the user's involvement for each snapshot taken. We improved the Snapshot to autonomously retake snapshots in real time when changes made to parts of the file system already included in snapshots. Furthermore, the experimental results show that our design and implementation largely outperform what the current HDFS can do in the process of making changes to the file system and taking corresponding snapshots.

Read the paper · More papers on PaperTik