A Deep-dive on Machine Learning for Cyber Security Use Cases

R. Vinayakumar, K. P. Soman, Prabaharan Poornachandran, Vijay Menon · 2019

Conventional methods, such as static and binary analysis of malware, are inefficient in addressing the escalation of malware because of the time taken to reverse engineer the binaries to create signatures. A signature for the malware is accessible and the malware might have made the significant damages to the system. The on-going malicious activities cannot be detected by the anti-malware system. Still, there is a chance of detecting the malicious activities by analyzing the events of DNS protocol in a timely manner. As DNS was not made by considering the security, it suffers from vulnerabilities. On the other direction, malicious Uniform Resource Locator (URL) or malicious website is a primary mechanism for internet malicious activities, such as spamming, identity theft, phishing, financial fraud, drive-by exploits, malware, etc. In this paper, the event data of DNS and malicious URLs are analyzed to create cyber threat situational-awareness framework. Previous studies have used blacklisting, regular expression and signature-matching approaches for analysis of DNS and malicious URLs. These approaches were completely ineffective at detecting variants of existing domain name/malicious URL or entirely newly-found domain name/URL. This issue can be mitigated by proposing the machine learning-based solution. This type of solution requires an extensive research on feature engineering and feature representation of security ’artefact type’, e.g. domain name and URLs. Moreover, feature engineering and feature representation resources must be continuously reformed to handle the variants of existing domain name/URL or an entirely new domain name/URL. In recent times, with the help of deep learning, artificial intelligent (AI) systems have achieved human-level performance in several domains. They have the capability to extract optimal feature representation by themselves by taking the raw inputs. To leverage and to transform the performance improvement of them towards the cyber security domain, we propose a method named as christened Deep-DGA Detect/Deep-URL, in which raw domain names/URLs are encoded using character-level embedding. Character- level embedding is a state-of-the-art method amplified in natural language processing (NLP) research [1]. Deep learning layers extract features from character-level embedding followed by feed-forward network with a non-linear activation function estimating the probability that the domain name/URL is malicious. For comparative study, various deep-learning layers are used. The optimal parameters for network and network structure are selected by conducting experiments. All the experiments are run for 1,000 epochs with a learning rate of 0.001. The deep-learning methods performed well in comparison to the traditional machine learning classifiers in all the test cases. Moreover, convolutional neural network-long short-term memory (CNN-LSTM) has performed well in comparison to other methods of deep learning. This is due to the fact that the deep-learning algorithms implicitly obtain the hierarchical features and long-range dependencies in character sequence in the domain name/URL. The proposed framework is highly scalable and which is named as Deep Cyber Threat Situational Awareness Framework (DCTSAF). DCSAF can handle 2 million events per second, analyzing the large volume and variety of data to perform near real-time analysis that is crucial to providing early warning about the malicious activities.

Read the paper · More papers on PaperTik