SINA: A Server-Assisted In-Network Repair Acceleration for Erasure-Coded Storage Systems
Geyao Cheng, Junxu Xia, Haibo Mi, Deke Guo, Kun Wang, Zhenyi Wang · 2025
In the erasure-coded storage systems, multiple related blocks have to be retrieved from other surviving nodes to repair a failed block. This incurs significant communication overhead with the surging scale of distributed storage systems. To mitigate the bandwidth bottleneck, in-network repair (INR) has emerged as a promising transport paradigm, which migrates the aggregation operations from the repair node to the programmable hardware, such as Intel Tofino switches. However, due to the limited on-chip memory size of these switches, the INR can degrade to the most primitive incast-type transmission, leading to massive traffic volume and hindered repair throughput. While we notice that, there are spare CPU cores in the storage servers that can be leveraged as alternative computing resources. With this intuition, we propose SINA, a Server-assisted In-Network repair Acceleration framework in this paper, which leverages the spare servers to assist aggregation operations when the programming switches' memory size is scarce for failure repair. We formulate this problem by adjusting the aggregation modes across the involved racks and solve this NP-hard problem using the Gurobi optimization solver. For all we know, this is the first work exploring spare servers for assisting the memory-scarce INR in erasure-coded storage systems. We have implemented SINA on an FPGA-based prototype system, and the experimental results show that SINA can ensure fault tolerance and accelerate failure repair by$5.0 \times$compared to the conventional methods.