Analysis of Network Attacks and Security Events using Modern Data Visualization Techniques
Paulo Macedo Pereira · Portuguese National Funding Agency for Science, Research and Technology (RCAAP Project by FCT) · 2015
Data visualization techniques comprise crucial resources in many research and professional areas. Effective representations often contribute to the understanding of the overall picture behind a large volume of data, sometimes leading to novel discoveries or to an ef cient synthesis. Due to the large amount of data that computers handle nowadays, many modern data visualizations techniques were designed to deal with such large data sets, exhibiting unique characteristics. In the information era, computers (and their operators) and networks are also amongst the biggest sources of raw data, though they are also used in its processing and storage. Many network monitoring systems and security appliances make usage of traditional data visualization techniques in reporting functionalities or to provide practitioners with status information. The scope of this work falls within the intersection of the elds of network security and data visualization techniques. Its objectives are to study modern approaches to represent data, which may be currently being used in other areas, and apply one of those approaches in the visualization of network traf c and attacks. Assessing the usefulness of the visualizations was also an objective, along with the constitution of a large data set of representations for several traf c classes and classical network attacks. A technique known as Circos, widely used for genomic representations, was the one applied for achieving the objectives of this masters program. Many representations for at least 18 different traf c traces were produced along this work, with many analyzed with detail in this dissertation. These traces, containing traf c generated by contemporary applications and classical network attacks or probing activities, were selected from two datasets. In order to produce the Circos, a minimal set of traf c characteristics was identi ed,and several scripts for automating the processing were implemented. Towards the nal part of this work, an experiment based on the (human) comparison between nine labeled and nine unlabeled Circos was set up to demonstrate that the obtained representations were useful up to the point of being used to identify traf c classes or attacks. During the experiment, it was possible to correctly identify eight, out of the nine, traces (one of the attacks was incorrectly classi ed as HTTP traf c), proving the usefulness of this technique in this eld.