From Fine-tuning to Output: An Empirical Investigation of Test Smells in Transformer-Based Test Code Generation
Ahmed Aljohani, Hyunsook Do · 2024
Researchers have recently leveraged transformer-based test code generation models to improve testing-related tasks (e.g., assert completion and test method generation). One such model, AthenaTest, has been recently introduced, and it has been well accepted by developers due to its ability to generate test cases similar to the ones written by developers. While the AthenaTest model provides adequate test coverage and improves test code readability, concerns remain regarding the quality of its generated test code, particularly the presence of test smells, which could degrade the test code's comprehension, readability, performance, and maintainability. In this paper, we investigated whether test cases generated by a transformer-based test code generation model - AthenaTest, contain test smells, including the presence of test smells in its fine-tuning dataset (Methods2Test). We evaluated seven test smells in AthenaTest's and Methods2Test's test cases. Our results reveal that 65% of Methods2Test's test cases contain test smells, which influence the output of AthenaTest, where 62% of its test cases contain at least one test smell. We also examined the design (test size and assertion frequency) of AthenaTest and Methods2Test test cases. Our findings show that AthenaTest tends to generate more assertions than the Methods2Test test case, which influenced the model to increase the occurrence rate of Assertion Roulette and Duplicate Assert smells.