Effort Estimation via Text Classification And Autoencoders
Rodrigo G. F. Soares · 2018
The estimation of the effort required for the production of a software or the correction of a software issue is an important activity in software development methodologies. Such estimates are essential to the delivery of high-quality products within a time frame and budget. Intuitively, these estimates should be produced as early as possible in the development process in order to improve planning. Issue reports are the earliest documents available for the correction of a software. We tackle effort estimation in early stages of agile software development as text classification of issue reports. This task consists of assigning categories to such documents, often written in natural language and programming code. Typically, text classification is performed with feature extraction techniques that produce informative features for a text classifier. Autoencoders generate input representations in increasing levels of abstraction, which may represent meaningful semantics in issue reports. We propose a study on the effectiveness of autoencoders for estimating the effort required for the agile correction of software functionalities in the context of real-world open-source projects. Our experiments provided significant evidences that autoencoders are able to encode noisy documents and provide informative features to a text classifier and improve its generalization.