Demystifying the Role of Publicly Available Up-to-Date Benchmark Intrusion Datasets: A Case Study of Web Security
Oumaima Chakir, Yassine Sadqi · River Publishers eBooks · 2025
With the proliferation of web application usage, the danger of web-based attacks has also increased. Thus, various machine learning (ML)-based web application firewalls (WAFs) have been proposed to enhance the security of web applications. Although the vast majority of studies focus on investigating algorithms that can improve detection performance, a few studies have been devoted to evaluating the reliability of benchmark datasets in the context of web security. To fill this gap, this article provides valuable information about the publicly available web benchmark datasets to beginner web security researchers by exploring (1) the role of benchmark datasets in developing ML-based WAF, (2) the relation between benchmark datasets and ML-based WAF’s performance, (3) the shortcomings of the currently available web-based attack datasets, and (4) the primary factors to consider while assembling appropriate benchmark datasets for ML-based WAF evaluation. The results of this study highlighted the need for up-to-date and representative benchmark datasets for ML-based WAFs evaluation since the currently available datasets are obsolete and do not meet the current-world web security.