Efficient, effective regression testing using safe test selection techniques

Mary Jean Harrold, Gregg Rothermel · 1996

Regression testing is an expensive procedure performed on modified software to establish confidence that the software performs correctly. Most research on regression testing addresses one or both of two problems: how to select regression tests from an existing test suite (the test selection problem), and how to determine portions of a modified program to retest (the coverage identification problem). Techniques for solving these problems are called selective retest techniques, because they reduce the cost of regression testing through selective reuse of tests, or selective retesting of program components. Most selective retest techniques focus on coverage identification, and use information gained in pursuit of coverage to guide test selection. We take an opposite tack. We first define the regression test selection problem precisely, and investigate the theoretical difficulty of the problem. We then present a framework for comparing and evaluating test selection techniques, and apply that framework to existing methods to determine where additional work is needed. Our analysis demonstrates a need for a test selection technique that has several qualifications. First, the technique must be safe: it must select every test, from the original test suite, that could reveal an error in the modified program. Second, the technique must be sufficiently precise: it must identify enough unnecessary tests to offset its own expense. Third, the technique must be general: applicable to a wide class of programs and modifications, both intraprocedurally (on single procedures), and interprocedurally (on programs and subsystems). Fourth, the technique must be efficient. Acting on these results, we present a family of test selection algorithms that meet the aforementioned qualifications: safe, efficient algorithms that function intraprocedurally and interprocedurally, and select more precise test sets than existing safe test selection algorithms. We discuss our prototype implementations of the algorithms, and the empirical studies we performed using the implementations. Our results suggest that for single procedures, regression test selection is not cost-effective. For entire programs or subsystems, however, our test selection algorithms offer substantial savings. Finally, we examine the coverage identification problem, and extend our test selection algorithms to assist with coverage identification.

Read the paper · More papers on PaperTik