Digital Cluster Circuits for Reliable Datacenters

Davide Rovelli, Patrick Th. Eugster · 2025

Promptly reacting to failures is key for highly available datacenter services. At the core, failure detectors (FDs) have to rely on timeouts which are hard to set correctly on top of heavily contended processors and network resources. To overcome unpredictable interaction jitter, designers still have to resort to large timeouts to preserve safety (i.e. false positives) at the cost of availability. The shift toward hybrid architectures – where applications span accelerators, disaggregated memory, and network devices – exacerbates this issue. With common-case latencies dropping to the μs-scale or lower, conservative timeouts on top of best-effort software make FDs and coordination services impractically slow and unreliable.Motivated by the growing adoption of programmable network devices (e.g. smartNICs, FPGA-switches), we propose a new paradigm to tackle this issue by pushing self-contained, time-sensitive services to hardware. We introduce the idea of modeling a distributed system as a large digital circuit characterized by a stable "clock signal" consisting of periodic packets delivered with ultra-low reliable latency. We highlight how current datacenter technologies can be used to enable synchronous interactions in practice and discuss how a reliable FD trivially built on top can be used to increase robustness and performance of both hardware and software distributed applications.

Read the paper · More papers on PaperTik