Hostmesh: Monitor and Diagnose Networks in Rail-optimized RoCE Clusters

Kefei Liu, Jiao Zhang, Zhuo Jiang, Xuan Zhang, Shixian Guo, Yangyang Bai, Yongbin Dong, Zhang Zhang, Xiang Shi, Lei Wang, Haoran Wei, Z. B. Wang, Yongchen Pan, Tian Gong Pan, Tao Huang · 2024

RoCE services are sensitive to failures and bottlenecks, which become more common as the RoCE network scales. To effectively detect and locate these problems independent of service traffic, RoCE networks require a monitoring and diagnostic system based on active probing. However, existing active probing schemes typically rely on a controller to design the probing plan for each server, which is difficult to deploy and has high synchronization overhead in multi-tenant clusters. Fortunately, rail-optimized clusters have become more common in recent years to improve network performance. In these clusters, the controller is unnecessary.

Read the paper · More papers on PaperTik