DiskReduce: RAIDing the Cloud

Bin Fan, Wittawat Tantisiriroj, Lin Xiao, Garth A. Gibson · 2010

Data-Intensive Scalable Computing (DISC) file systems such as HDFS employs replication for reliability, typically delivering users with only about a third of the storage capacity of the raw disks. In this project, we investigate DiskReduce, a framework for integrating RAID into these replicated storage systems to lower storage capacity overhead, for example, from 200 % to 25 % when triplicated data is dynamically replaced with 8+2 RAID 6 encoding. We gathered usage data from large HDFS DISC systems and find that DISC files are huge relative to traditional and HPC file systems, but because DISC blocks are also huge, perfile RAID wastes significant capacity. We chose to encode blocks across files. We also studied the implication of reading RAIDed data to MapReduce job performance. We measured read performance benefits from replication that will be lost

Read the paper · More papers on PaperTik