Automated test generation for Scratch programs

Adina Deiner, Patric Feldmeier, Gordon Fraser, Sebastian Schweikl, Wengran Wang · Empirical Software Engineering · 2023

Abstract The importance of programming education has led to dedicated educational programming environments, where users visually arrange block-based programming constructs that typically control graphical, interactive game-like programs. TheScratchprogramming environment is particularly popular, with more than 90 million registered users at the time of this writing. While the block-based nature ofScratchhelps learners by preventing syntactical mistakes, there nevertheless remains a need to provide feedback and support in order to implement desired functionality. To support individual learning and classroom settings, this feedback and support should ideally be provided in an automated fashion, which requires tests to enable dynamic program analysis. In prior work we introducedWhisker, a framework that enables automated testing ofScratchprograms. However, creating these automated tests forScratchprograms is challenging. In this paper, we therefore investigate how to automatically generateWhiskertests. Generating tests forScratchraises important challenges: First, game-like programs are typically randomised, leading to flaky tests. Second,Scratchprograms usually consist of animations and interactions with long delays, inhibiting the application of classical test generation approaches. Thus, the new application domain raises the question of which test generation technique is best suited to produce high coverage tests capable of detecting faulty behaviour. We investigate these questions using an extension of theWhiskertest framework for automated test generation. Evaluation on common programming exercises, a random sample of 1000Scratchuser programs, and the 1000 most popularScratchprograms demonstrates that our approach enablesWhiskerto reliably accelerate test executions, and even though manyScratchprograms are small and easy to cover, there are many unique challenges for which advanced search-based test generation using many-objective algorithms is needed in order to achieve high coverage.

Read the paper · More papers on PaperTik