Enhancing Cybersecurity Risk Assessment with Text Mining of Security Bulletins and Advisory Data

Srinivasan Suresh Kumar, Chandrasekhar Rohith Bhat · 2025

The potential improvement on the specifics of cybersecurity risk evaluation is the subject of this paper with a set methodology that includes text mining on security bulletins and advisory data besides machine learning algorithms and anomaly detection. The data collection and preprocessing involves crawling textual data and cleaning the intense textual data and the process further involves feature extraction is done then using TF-IDF, Word Embeddings and certain technical values such as CVSS scores. Therefore, the machine learning models such as XGBoost, Random Forest, SVM, and Logistic Regression classify vulnerable features according to the degree of risk severity. Therefore, while developing the framework, it is demonstrated that XGBoost attains the overall accuracy of 92%, precision of 0.91, recall of 0.94, and AUC-ROC of 0.95; Random Forest results in rather good accuracy of 89% and 0.92 AUCROC. Through autoencoders based anomaly detection, new vulnerabilities that may not be detected using other approaches are recognized. Diagnosis criteria including accuracy, precision, recall, F1 score, as well as AUC of the ROC curves estimate the model performance to guide further remedial actions. Feature importance analysis is used to find out which variables affect the predictions of the risks most. This facilitates identification of the extent of vulnerability and distances that require risk control so that effective risk management can be effected. It enhances the process of detecting and handling new threats in cybersecurity since it unites features of machine learning, text mining and anomaly detection.

Read the paper · More papers on PaperTik