Comparative analysis of FastText and Doc2Vec for semantic feature oriented software defect prediction
Priya Singh, Gaurav Sharma · 2025
Embeddings are valued for their capacity to grasp semantic relations, streamline dimensionality, and discern data patterns. These are used across various domains of machine learning and artificial intelligence. In the area of software defect prediction, choosing an appropriate embedding technique is very important. This study aims to compare the effectiveness of two prominent word embedding methods, FastText and Doc2Vec, when applied to detecting bugs in software. The entire analysis is conducted by using different sets of Java projects taken from an open-source Promise repository. Through rigorous training and evaluation of several deep learning models, generally designed for defect detection in software, a comprehensive evaluation was conducted. Evaluation metrics, such as Matthews correlation coefficient, specificity, and accuracy, along with other important performance indicators, were used to assess the effectiveness of both techniques. The results indicate that Doc2Vec performs significantly better than FastText in terms of multiple metrics, indicating its superiority in predicting defects in the software.