Investigating the Impact of LT-like Fountain Codes Degree Distribution on Distributed Machine-Learning Applications
Nadhir Ibrahim Abdulkhaleq, Ahmed Saad Hussein, Basim Mahbooba · 2025
This work proposes using Luby Transform (LT) codes for efficient, fault-tolerant distributed machine learning (ML) training. Current distributed ML practices follow uniform data partition across worker nodes, leading to inefficiency in handling stragglers and system crashes. To address this, we employ LT-coded task assignment to represent training data in redundant chunks with different redundancy levels, rendering computation resilient. Among the primary research focuses of this work is studying how different LT degree distributions affect convergence rate and model accuracy. We implement and test LT-coded distributed linear regression on MATLAB, encoding data across worker nodes with varying degree distributions. Nodes compute local gradients based on input received and forward resulting gradients to a central server for global summation. Our findings demonstrate that degree distributions of various degrees significantly impact the stability, convergence rate, and resultant parameter values. We further offer visualization and convergence analysis to indicate the trade-offs between redundancy and learning efficiency and thus allow the design of optimal degree distributions for future distributed ML applications. Our results show the impact of various degree distributions on model convergence and effectiveness and shed light on learning efficiency vs. redundancy trade-offs. This work establishes grounds for future improvements in distributed learning infrastructure by opening the gate to improved LT-coded ML strategies.