Network intrusion monitoring based on margin distance pruning and RF algorithm

Xiangming Gou, Md Gapar Md Johar, Jacquline Tham · Results in Engineering · 2025

• An enhanced random forest and K-means algorithm-based network intrusion monitoring technique is suggested to increase the effectiveness of the system. • First, margin optimal distance pruning is used, and feature ordering and K-means++ clustering are introduced to improve the random forest and build an intrusion detection model based on an improved random forest algorithm. • Then, density-based spatial clustering and local anomaly factors are combined with applications with noise to design anomalous data detection algorithms, and the Apriori and K-means algorithms are combined to design improved cluster analysis algorithms to build a data mining model based on an improved K-means algorithm. • The findings demonstrated that the accuracy of the intrusion detection model in the LUFlow and CIDDS datasets was 79.4 % and 91.2 %, respectively. • The model constructed in this study has a good application effect and helps ensure network security. With the rapid development of network technology, network intrusion events are becoming increasingly frequent and complex. Traditional intrusion detection systems often have low detection accuracy, posing a serious threat to information security. To address the shortcomings of traditional methods in handling large-scale imbalanced data and improve the accuracy and efficiency of network intrusion monitoring, an enhanced random forest and K-means algorithm-based network intrusion monitoring technique is suggested. Firstly, margin optimization distance pruning is adopted, and feature sorting and K-means++clustering are introduced to improve the random forest algorithm. Then, combining density based spatial clustering and local anomaly factor detection of abnormal data, and combining Apriori algorithm and K-means algorithm for clustering analysis. The results showed that the accuracy of the proposed intrusion detection model in the LUFlow and CIDDS datasets was 79.4 % and 91.2 %, respectively, with modeling times of 80 s and 118 s, respectively. The accuracy and F1 value of the anomaly detection algorithm are 0.88 and 0.90, respectively. In summary, the model constructed in this study has good accuracy in intrusion detection and abnormal data detection, which helps ensure network security. However, it may also lead to an increase in modeling time. In the future, modeling efficiency can be improved through introducing distributed and parallel computing.

Read the paper · More papers on PaperTik