Adaptive Fault Tolerance Using Machine Learning for Dynamic Distributed Systems
Dr Manoj Kumar Niranjan · INTERANTIONAL JOURNAL OF SCIENTIFIC RESEARCH IN ENGINEERING AND MANAGEMENT · 2025
. Distributed systems are foundational to modern computing paradigms like cloud, edge, and IoT environments, yet their dynamic nature poses significant fault tolerance challenges. Traditional fault tolerance mechanisms often lack adaptability and scalability in real-time environments. This paper explores the integration of machine learning (ML) techniques to develop adaptive fault tolerance for distributed systems. Through a comprehensive literature review and a simulation based implementation, we demonstrate the efficacy of ML models—specifically a Random Forest classifier—in predicting node failures and proactively redistributing tasks. The proposed framework combines reactive and proactive strategies to enhance system resilience. Results from the simulation highlight an 85% failure prediction accuracy and reduced disruption through intelligent task redistribution. The work addresses key research gaps in end-to-end adaptive frameworks, lightweight ML models, and practical validation. Future enhancements include real-world testing, hybrid ML integration, and explainable AI techniques for greater trust and applicability. Keywords: fault tolerance, distributed systems, machine learning, proactive recovery, Random Forest