Performance and reliability

Daniel A. Reed · ACM SIGMETRICS Performance Evaluation Review · 2006

Legend says that Archimedes remarked, on the discovery of the lever, "Give me a place to stand and I can move the world." Today, computing pervades all aspects of society. "Science" and "computational science" have become largely synonymous, and computing is the intellectual lever that opens the pathway to discovery in diverse domains. As new discoveries increasingly lie at the interstices of traditional disciplines, computing is also the enabler for scholarship in the arts, humanities, creative practice and public policy. Equally importantly, computing supports our critical infrastructure, from monetary and communication systems to the electric power grid.With such pervasive dependence, computing system reliability and performance are ever more critical. Although the mean time before failure (MTBF) of commodity hardware components (i.e., processors, disks, memories, power supplies and networks) is high, their use in large, mission critical systems can still lead to systemic failures. Our thesis is that the "two worlds" of software -- distributed systems and sequential/parallel systems -- must meet, embodying ideas from each, if we are to build resilient systems. This talk surveys some of these challenges and presents possible approaches for resilient design, ranging from intelligent hardware monitoring and adaptation, through low-overhead recovery schemes, statistical sampling and differential scheduling and to alternative models of system software, including evolutionary adaptation.

Read the paper · More papers on PaperTik