Enhancing Machine Learning in Information Security: Power-Law Distribution and Dragon King
Yuan Wang, Xiaofan Chen · 2023
This paper addresses challenges in applying machine learning to information security, focusing on the pivotal role of training data quality. We investigate outlier removal, particularly in heavy-tailed distributions like those in malicious software families. Through a case study on malicious file identification, we identify a power-law distribution, posing a unique challenge. Inspired by Sornette's ‘dragon king’ concept, our analysis reveals these exceptional entities within malicious software families, which, when recognized as defaults, impact labeling consistency. Removing these outliers in power-law distribution improves machine learning model performance.