A Review of Resilience Testing in Microservices Architectures: Implementing Chaos Engineering for Fault Tolerance and System Reliability
Akalanka Bandara Mailewa, Arunkumar Akuthota, Thivanka M. Dissanayake Mohottalalage · 2025
This paper evaluates the impact of chaos engineering on the resilience of microservices architectures. By simulating controlled failures, this research highlights specific findings, including reduced system downtime, improved fault detection, and faster incident resolution. A comparative analysis of tools like Chaos Monkey, Litmus, and Azure Chaos Studio is included, offering insights into best practices for implementation. The core principle of chaos engineering is that by inducing controlled failures, vulnerabilities within a system can be identified and addressed proactively, thereby mitigating risks in the production environment. This research focuses on microservices-small, independent code components that perform specific tasks within larger applications. Due to their distributed nature, often hosted on various servers and managed by different teams, microservices are inherently more fragile compared to monolithic applications. Additionally, their use of diverse programming languages complicates the understanding and testing of these systems for bugs. The primary objective of this study is to evaluate the efficacy of chaos engineering in enhancing the resilience of microservices, specifically by identifying common failure types and assessing the time required for their resolution.