DETECTION of INFRASTRUCTURE ANOMALIES in BUILD LOGS USING MACHINE LEARNINGText classification on Continous Integration log files.
Didrik Lindqvist · DiVA at Umeå University (Umeå University) · 2019
Continuous integration is a practice where software developers integrate their code to a bigger codebase multiple times per day. Before the integration, the code is built and tested by e.g open source build tools such as Jenkins, and the information produced during this process is stored in a log file. Sometimes these builds fail, and the cause can be either user or infrastructure related. A user related error may be that the code cannot compile due to syntax error and an infrastructure error could be a DNS problem. This thesis evaluated how well machine learning can be used to label the cause on failed build logs as either user or infrastructure. This thesis compared the performance of three machine learning algorithms: support-vector machine, random forest, and gradient boosting classifier. Two different datasets are used in this study. A balanced dataset used for training and validation and another dataset used for testing. The preprocessing step, including feature selection, is done using term frequency-inverse document frequency, which converts the text from the build log to a machine learning friendly format. The study also evaluated three different sizes of n-grams for each algorithm and dataset. The performance for the three machine learning algorithms is evaluated by comparing the precision, recall, and F1-score for each model. The three machine learning algorithms and the methodology around preprocessing and evaluation are explained in this study. The results show that machine learning can be used as a tool to help the CI-owners, but may not be used to fully replace the classification done manually today. The machine learning algorithm that performed the best was gradient boosting classifier with an bag of 1 and 2-grams, with a precision, recall and F1-score of 0.87, 0.73 and 0.79.i