Chaos Engineering in the Wild: Findings from GitHub
Joshua Owotogbe, Indika Kumara, Dario Di Nucci, Damian A. Tamburri, Willem‐Jan van den Heuvel · arXiv (Cornell University) · 2025
Chaos Engineering (CE) enhances system resilience by deliberately injecting faults to uncover weaknesses before they cause outages. Despite the availability of many CE tools, little is known about how they are adopted and maintained in open-source software (OSS) projects. This study empirically characterizes the adoption, evolution, and practical use of CE tools in OSS projects, examining who adopts them, how they are maintained, and which fault types they target. We conducted a large-scale mining study of GitHub repositories associated with ten widely used CE tools. Starting from 5,410 candidate repositories, we systematically filtered and manually validated 1,275 records and analyzed repository metadata, commit histories, and documentation. We found that adoption is concentrated around a few tools, with Toxiproxy, Chaos Mesh, and Chaos Monkey accounting for 68.86% of the validated repositories. In terms of adopter context, development is the predominant repository purpose (56.55%), while industry represents the largest ownership category (34.82%), closely followed by personal repositories (33.49%). At the project level, activity varies substantially, with 51.42% of repositories having at most 50 commits. In terms of fault coverage, network faults (44.85%) and instance termination (29.96%) together account for 74.81% of the 2,410 observed fault instances, whereas application-level faults account for only 2.57%. Taken together, these findings suggest that practitioners should consider repository activity and fault coverage when selecting CE tools, while researchers should investigate the factors behind concentrated adoption and limited application-level experimentation.