Leveraging XGBoost Machine Learning Algorithm for Common Vulnerabilities and Exposures (CVE) Exploitability Classification

M. Hareesh Babu, Namruth Reddy S, Minal Moharir, Mohana · 2024

The problem of exploitation is highly critical in cybersecurity. Therefore, it is very important for the prioritization of security patching efforts and risk mitigation to accurately predict the exploitability of a given vulnerability using the Common Vulnerabilities and Exposures System. In this paper XGBoost ML model is used for the classification of CVEs. According to the analysis of the data, compared with others, network-based vulnerabilities are still predominant at 85%, while local and adjacent network vectors stood at 12.9% and 1.9%, respectively. The estimate of the most affected operating systems was on Linux and Microsoft, with about 37,500 and 11,000 CVEs, respectively. Most of the vulnerabilities resided in cryptographic issues, cross-site scripting, restriction on memory buffers, and SQL injection. The training and testing used in the XGBoost model was based on a CVE dataset obtained from the National Vulnerability Database where major features impacting exploitability were targeted. Proposed model obtained good model accuracy, precision, recall, and an F1-score, of 94%, 92%, 91%, and 92%, respectively, thereby proving to be a very efficient model in distinguishing between exploitable and non-exploitable vulnerabilities. This work contributes to the domain of vulnerability analysis by presenting the potential of XGBoost on CVE classification, along with several actionable insights into vulnerability management.

Read the paper · More papers on PaperTik