Empirical analysis of gherkin quality and common mistakes by novice testers: Implications for sustainable software testing
Luthfi Satrio Wicaksono, Joe Lian Min, Asri Maspupah, Rio Agasta, Jelang Anugrah Raharjo, Reqi Jumantara Hapid, Muhammad Rafif Genadratama, Yahya Alfon Sinaga · Springer Link (Chiba Institute of Technology) · 2025
Behavior-Driven Development (BDD) relies on quality Gherkin scenarios to bridge communication between technical and non-technical stakeholders. However, novice software testers often struggle to transition from procedural writing to declarative scenarios, resulting in low-quality scenarios and wasted automated testing resources. This research proposes a systematic rubric-based methodology with 12 assessment aspects to measure the quality of Gherkin scenarios created by novices software tester. The focus of the study is an empirical comparative analysis of Gherkin file artifacts from two representative systems, JTK-Learn (novices) and WebsiteOne (professionals), to identify quality gaps and common mistakes. The research stages include data collection, rubric development, Gherkin quality measurement, and comparative descriptive quantitative analysis. The results show significant gaps in conceptual and collaborative aspects, while novices are strong in structural aspects, such as scenario focus and step structure. Five common mistakes made by novices were identified, namely the use of technical vocabulary, procedural writing, inconsistency in domain terminology, mixing of abstraction levels, and step duplication that reduces maintainability. These mistakes cause a waste of computational resources and energy. However, with the validated rubric as a reliable evaluation tool, there is a clear potential for more effective BDD training development and sustainable software testing practices, which can significantly improve the efficiency of automated test execution.