Identifying Significant Features for Network Forensic Analysis Using Artificial Intelligence Techniques.

Srinivas Mukkamala, Andrew H. Sung · International journal of digital evidence · 2003

Network forensics is the study of analyzing network activity in order to discover the source of security policy violations or information assurance breaches. Capturing network activity for forensic analysis is simple in theory, but relatively trivial in practice. Not all the information captured or recorded will be useful for analysis. Identifying key features that reveal information deemed worthy for further intelligent analysis is a problem of great interest to the researchers in the field. The focus of this paper is the use of artificial intelligent techniques for offline intrusion analysis, to protect the integrity and confidentiality of the information infrastructure. An effective forensic tool is essential for ensuring information assurance by updating the newly identified security breaches in to the organizations protection and detection mechanisms. Two artificial intelligent techniques are studied: Artificial Neural Networks (ANNs) and Support Vector Machines (SVMs). We show that SVMs are superior to ANNs for network forensics in three critical respects: 1. SVMs train, and run an order of magnitude faster; 2. SVMs scale much better; and 3. SVMs give higher classification accuracy. We also address the related issue of ranking the importance of input features, which is a problem of great interest in modeling. Since elimination of the insignificant and/or useless inputs leads to a simplification of the problem and may allow faster and more accurate detection, feature selection is very important in network forensics. Two methods for feature ranking are presented; the first one is independent of the modeling tool, while the second method is specific to SVMs. The two methods are applied to identify the important features in the 1999 DARPA intrusion data. It is shown that the two methods produce results that are largely consistent.

Read the paper · More papers on PaperTik