Hybrid Software Defect Prediction Based on LSTM (Long Short Term Memory) and Word Embedding

Rizal Broer Bahaweres, Detia Jumral, Irman Hermadi, Arif Imam Suroso, Yandra Arkeman · 2021

Defects often occur in software so it can cause various problems for users. The running system may have defects, although they are not immediately apparent. Defects can be UI or display defects or defects in logic code. In predicting software defects several methods have been developed focusing on software metrics (Line of Code, Cyclomatic Complexity, etc.). But in capturing program syntax and semantics and the ability to build accurate predictive models, software metrics often fail. We propose a method that combines deep learning methods and word embedding to predict software defects. We will map tokens using an abstract syntax tree from the source code and then build an LSTM network to predict software defects. The dataset used is taken from an open source Java project, namely the Apache Project Repository. The evaluation results show that the LSTM method combined with word embedding can get accuracy 98%, precision 89%, recall 92% and f1-score 90%. This is greater than other software defect prediction methods.

Read the paper · More papers on PaperTik