DATA CLUSTERING ALGORITHM FOR PROTECTING CONFIDENTIAL INFORMATION ON THE INTERNET
I.S. Bereshpolov, Y.A. Kravchenko, Aleksandr Sleptsov · Известия Южного федерального университета. Технические науки · 2023
The article is devoted to solving the scientific problem of protecting confidential informationin the Internet based on the algorithm for clustering significant amounts of data. The protection ofa computer network confidential information is a hot topic for research, especially in connectionwith the growing use of information technology and the increase in data of valuable informationstored in the Internet. With the growth of information responsibility, the need for effective methodsof computer networks information security has become critical. In this scientific article, the authorspropose a solution to the problem of protecting computer networks confidential informationbased on the big data clustering algorithm. Traditional intrusion detection methods have limitationssuch as the ability to work only with one- or two-dimensional data, and also have a strongreliance on prior knowledge. To eliminate these limitations, the authors propose a heuristic intrusiondetection algorithm that uses clustering based on a cloud model. The proposed algorithmtakes advantage of both labeled and unlabeled samples for data clustering, thereby reducing relianceon a priori knowledge. The results of a computational experiment carried out on the proposedalgorithm were compared with several canonical intrusion detection algorithms. The resultsshowed that the proposed algorithm improved the performance of the intrusion detection system,increased the accuracy of detection, reduced the false alarm rate, and enhanced the reliability ofthe system. The dynamic weighting method used in the algorithm removed the complexity of highleveldata processing and allowed the algorithm to learn itself, resulting in a relatively stablecloud model. Despite the significant improvement in the performance of the proposed algorithmcompared to the canonical clustering algorithms, the results of the study also showed that thealgorithm has some limitations, such as a high false positive rate and sensitivity to data with certaintypes of distribution. To eliminate these shortcomings, further improvement of the algorithm isrequired. In general, the proposed heuristic clustering intrusion detection algorithm based on thecloud model is a promising solution for protecting computer networks confidential information.