Automated Dynamic Detection of Ransomware using Augmented Bootstrapping

M Sukul, S. Aswin Lakshmanan, R Gowtham · 2022 6th International Conference on Trends in Electronics and Informatics (ICOEI) · 2022

Ransomware is a subset of malware that blocks the access of infected computers by encrypting the user files and locking the screen. To restore the computer's functionality and user files, the malware demands a ransom, which must be paid in the form of cryptocurrency within a stipulated period. The deleterious effects caused by ransomware must be detected and addressed promptly. It is often difficult to perceive the metrics and parameters to look for, particularly when conducting experiments with unfamiliar ransomware variants. Traditionally, analysis carried out on these suspicious files can be categorized as static or dynamic based analysis. Static-based signature analysis is tremendously affected because of its nature of generating hard-coded policies or rules. Attackers often attempt to breach them by employing various techniques like code obfuscation, encryption, packaging, etc. While the dynamic analysis provides more insights into the analyzed files. The proposed work observes the behavioral aspects of ransomware and builds a generic model using machine learning techniques to detect the variants of ransomware families. Such threats must be detected early, i.e., before encryption takes place. Early detection, however, is hampered by a lack of sufficient information in the early stages of an attack, which leads to low detection accuracy and a high rate of false alarms. Currently, available solutions assume that there is complete knowledge of the behavior of such attacks at the time of detection, but that is not the case. The same is not true for early detection which takes place during an attack when the data are not fully accessible. Our paper proposes a novel technique to address these limitations called Augmented Bootstrapping and this method analyses the API calls made by the ransomware along with their timestamps. These timestamps indicate the API call sequences, which in turn help in constructing the N-gram model. The names of API calls are concatenated together to generate the bigram, trigram, 4-gram, and 5-gram models. Further, TF-IDF scores will be calculated for each of these N-grams sequences, and a machine learning model will be trained based on these values. The method analyses API call sequences of each of the active applications in the system and confirms the ransomware attack when it matches a malicious call sequence pattern. Our empirical results confirm that the proposed work could detect ransomware attacks with high accuracy and fewer false predictions.

Read the paper · More papers on PaperTik