Efficient Error-Detection and Recovery Mechanisms for Reliability and Resiliency of Multicores

Sandip Kumar Kundu, Omer Khan · 2016

With increasing density of power, traditional frequency scaling of processors came to an end. The power wall forced the industry to seek performance from parallel processing - shifting focus away from instruction throughput to thread throughput. This led to the development of multicore processors. However, extracting performance from multicore is not an easy problem. There are two approaches towards utilizing the multicores: scale-out parallelism that seeks performance by running independent threads and scale-up parallelism where multiple threads of a program execute in parallel to improve performance. In scale-out parallelism, the memory traffic increases linearly with the number of threads until the memory bandwidth is fully saturated at which point increasing the number of threads yield no benefit. Scale-up parallelism is contingent on programmers writing parallel software that takes advantage of the multicores - a task that fraught with perils. Data race in multithread programming is a well-known issue. What is less-known is that a majority of the resiliency and recovery solutions that have been developed for single cores do not scale for multicores.Reliability of processors is a major concern today. Catastrophic failures such as power loss can erase months of computation. System wide check-pointing provides a measure of protection against loss from such failures. Unfortunately, system wide check-pointing is not practical for errors that occur for individual threads, as such errors are far more prevalent today. Single event upsets (SEU) are caused by radiation-induced soft errors as well as timing errors that stem from signal integrity problems. Power supply noise caused either by circuit switching or power converters is a major source of such errors as are thermal issues and other sources of electrical noises. If a thread experiences an error, (i) we must have detection mechanisms in place as well as (ii) correction mechanisms. In general it is not possible to roll back a thread without rolling back inter-thread messages that may trigger cascading roll-backs of other threads. Such roll-backs result in wasted computation and wasted power.In this tutorial, we will explore this issue at depth and present state of the art industrial and academic solutions to address these problems. This tutorial also covers the resiliency trade-offs in modern processors that arise from increasing error-rates with decreasing supply voltage which is the preferred solution for reducing processor power that enables parallelism and multicores in the first place. We will explore multi-objective optimization problem that seek to satisfy dual objectives of performance-per-watt and lower error-rate. We will describe novel solutions for ensuring total store order (TSO) in processors in presence of errors. The tutorial presenters are academics with long experience in industry. They have studied these problems at-depth in their research. The tutorial is targeted for a wide audience and does not require attendees to be fully-versed in design or microarchitecture. The proposed tutorial will introduce the problem both qualitatively and quantitatively; discuss microarchitecture alternatives and trade-offs for resiliency. Although this tutorial primarily focuses on hardware sources of errors, it will touch upon programming errors, and debug support for identification of software bugs.

Read the paper · More papers on PaperTik