Conformal prediction for labelling and updating online models in the presence of concept drift in cybersecurity

David Escudero García, Noemí DeCastro‐García · Journal of Information Security and Applications · 2025

Machine learning is used for detecting malicious activity in cybersecurity contexts since it provides more adaptable models than signature-based solutions. One of the main challenges in applying machine learning to detect malicious activity is the presence of concept drift, which is a change in data distribution over time. Online models that are updated dynamically are usually applied to handle drift. However, these models require new labelled instances to be updated. Reliable labels are typically scarce, expensive to obtain, and not immediately available, which makes building an effective model difficult. In this work, we propose applying online models with conformal prediction , which provides statistical guarantees, to obtain reliable pseudo-labels to update the model and mitigate the absence of ground truth in new data. Although the use of conformal pseudo-labels produces significant improvements in some cases, these are inconsistent across datasets and models, which limits the applicability of the approach.

Read the paper · More papers on PaperTik