Building a Reusable Defect Resolution Time Prediction Model Based on a Massive Open-Source Dataset: An Industrial Report
Murad Mamedov, Ksenia Vorontsova, Elena Treshcheva, Iosif Itkin · 2021
This research describes an attempt to build a model aimed to reproduce predictive ability of an experienced human tester and provide analytical insights beyond statistical dependencies. An important step in this direction is working out an approach to create a dataset that has a sufficient size and is organized as a set of features allowing to train a model yielding reliable results without being dependent on data from specific projects. In this paper, we outline our experience of designing a general approach for time-to-resolve prediction for defect reports to assist in defect management tasks. We describe the process of dataset preparation based on the massive open-source data and the experiments using bi- and multi-labeling classification. The article covers the evolved understanding of the task, approaches to classification, and comparison of the results.