Analysis of repair algorithms for mirrored-disk systems
Hannu H. Kari, Heikki Saikkonen, N. Park, Fabrizio Aghini Lombardi · IEEE Transactions on Reliability · 1997
This paper analyzes the effects of several disk-repair algorithms (DRA) for a mirrored disk subsystem (RAID-1). The main interest is in disk faults and how the repair-process copies data, for user requests, 'from a fault-free disk to a spare disk' with the least performance-degradation. This study compares how various DRA affect system performance. Two DRA are compared and two access patterns (uniform and nonuniform) are studied to establish their effects on the repair process and performance. Sector faults are repaired using the reassign block facility in the SCSI protocol. When the 'mean load of the disk subsystem is moderate' and the 'sector repair time is of the same order of magnitude as the mean disk request processing time', then the differences between various DRA are minor. Simulation results indicate that the performance degradation of user disk requests can be reduced by introducing a short delay in the repair algorithm, A new algorithm (DRA 3) for detecting sector faults is presented. It scans the disk space, while no user disk-requests are issued, and using the advanced statistics of SCSI disks detects deteriorated media. Its advantage is that it can repair the disk subsystem before data are actually lost due to a media defect.