Automated regression testing of database applications
Erik Rogstad · NORA - Norwegian Open Research Archives · 2016
Ensuring the functional quality of database applications is a very important problem in software testing, yet few innovative solutions and empirical studies are reported on the subject. Database applications are widely adopted, for example in public administrations and banks, as they need to process large amounts of transactions efficiently for a large population and store large amounts of data. Such applications are often highly automated in order to efficiently cope with a large number of transactions, and are usually difficult to maintain and change. In order to preserve system quality through frequent system releases, a thorough, systematic, and automated regression test approach is needed for such applications, as they tend to provide core business value to their organizations. The objective of this thesis is to help scale functional system-level regression testing of database applications through cost-effective automation. We propose a black-box approach that relies on classification trees to model the input domain of the system under test. We use the classification tree models as basis for automatically selecting abstract regression test cases, and either generate test data automatically according to the model specifications or rely on production data (data cloned from the production environment) that match the model specifications. Regression testing is carried out by running the selected test cases on consecutive versions of the system under test, while automatically capturing changes in the database state during system execution. The captured database manipulations for each test case are automatically compared across system versions, and test cases that deviate between two system versions are either due to anticipated changes in the release, or regression faults. The resulting deviations from a regression test are clustered according to their output characteristics, so that deviations resulting from the same change or fault (ideally) are contained in the same cluster. These clusters are then used to minimize the effort required to analyze deviations. In order to evaluate our approach, we conducted a large-scale case study in a real development environment at the Norwegian Tax Department. The Norwegian Tax Department maintains several database applications, which are built on standard and widely used database technology and are representative of many such applications in public administrations. Together with the Norwegian Tax Department, we developed an industry strength tool in accordance with our proposed regression test approach. We applied the tool to the regression testing of their tax accounting system, thus evaluating its applicability and scalability on a large-scale database application with real changes and regression faults. The results of our study showed that our proposed solution to regression testing helps mitigate risk when releasing new versions of a system, as it is more thoroughly, yet efficiently tested, causing less regression faults to be released. When our tool was applied for regression testing of eight consecutive releases of the subject system, it helped identify 60% additional faults to those found through regular testing. We regard this as a substantial contribution in terms of increased fault detection power. Furthermore, we made a thorough assessment of various strategies for selecting test cases based on classification tree models. When using existing production data sets as basis for regression testing, carefully selecting test cases according to their model partition coverage, can help reduce test effort dramatically. This is important in cases when running regression test cases are expensive, or when test results have to be manually inspected. For example, when selecting test cases according to our proposed selection approach, nearly 80% of the regression faults were captured when selecting only 5% of the test cases for execution. This is important in order to scale the regression test effort when the input domain, and thus the number of possible test cases are very large. To further add on the scalability of the regression test approach, our results show that clustering regression test deviations based on their output characteristics can help reduce the effort spent on analyzing such deviations. The clustering strategy turned out to be very accurate as it resulted in homogenous clusters (all deviations in each cluster match one change or fault) for all regression test campaigns assessed. This implies that testers can inspect one deviation only from each cluster and still remain confident of finding all regression faults. Moreover, we assessed the cost-effectiveness of various strategies for regression testing and found that combining combinatorial test suites with test suites conforming to the operational profile of the system under test was effective, as neither one alone is sufficient to find all kinds of regression faults. In conclusion, we have proposed a novel and holistic solution to functional systemlevel regression testing of database applications, that automates many steps in the test process. The system under test is tested in a structured and systematic manner as we rely on model specifications to drive the test process. The regression test approach has been evaluated in a real and representative development environment and proved to be both effective in detecting regression faults and to scale for testing large database applications.