Hierarchical RAID: Organization, Operation, Reliability and Performance

Alexander Thomasian, Yujie Tang · 2011

We consider two level Hierarchical RAID (HRAID) arrays with erasure coding at both levels. The main advantage of HRAID is tolerating disk array controller failures in addition to disk failures, so that it is suitable for storage clouds based on bricks. We consider an HRAID with N nodes consisting of independent RAID controllers and M disks per node. Each node tolerates the failure of l disks and the controllers use a software protocol to coordinate their activities to implement an inter-node RAID level tolerating k node failures. We postulate MDS coding which incurs the minimal level of redundancy at both levels. HRAIDk/l can fail with a minimum of (k+1)(l+1) disk failures, regardless of N and M , but the probability of this event is very small. The maximum number of disk failures with no controller failures is N × l + (M − l)k. We specify examples of HRAID organization, describe distributed transactions to ensure the integrity of updates, and assess the small write penalty in HRAID and compare it with single level RAID. We report Monte-Carlo simulation results to obtain the Mean Time to Data Loss (MTTDL) assuming repeated successful restripings via parity sparing until such space is exhausted. For fixed k × l, an approximate reliability analysis, which does not take into account controller failures, shows that k < l yields the higher reliability, but simulation results show that the opposite is true for less reliable controllers.

Read the paper · More papers on PaperTik