Enhancing Gradient Descent with Parallel Computing: A Scalable Optimization Using Federated Learning
Deepthy K Bhaskar, B Minimol, V. P. Binu · 2025
Traditional Stochastic Gradient Descent (SGD) follows a sequential update process, which can be slow and inefficient for large-scale distributed learning tasks. Parallel computing offers a powerful way to accelerate gradient descent by enabling simultaneous computation of gradients across multiple workers. In this paper, we explore Parallel Stochastic Gradient Descent (Parallel SGD) as an optimization strategy by make use of Federated Learning (FL). We compare the efficiency of Parallel SGD against traditional gradient descent methods, highlighting its advantages in convergence speed and scalability. Furthermore, we discuss how different synchronization strategies impact model accuracy and training efficiency. Our experimental results demonstrate that Parallel SGD significantly reduces training time while maintaining accuracy, making it a promising optimization method. The results show speedups of up to 1.3x with 5 workers while preserving 92.7% accuracy compared to centralized SGD's 94.3 %, validating its effectiveness for distributed machine learning scenarios.