Modeling and tolerating heterogeneous failures in large parallel systems

E. M. Heien, Derrick Kondo, Ana Gainaru, Daniel LaPine, Bill Kramer, Franck Cappello · 2011

As supercomputers and clusters increase in size and complexity, system failures are inevitable. Different hardware components (such as memory, disk, or network) of such systems can have different failure rates. Prior works assume failures equally affect an application, whereas our goal is to provide failure models for applications that reflect their specific component usage. This is challenging because component failure dynamics are heterogeneous in space and time.

Read the paper · More papers on PaperTik