Mitigating False Positives in DGA Detection for Non-English Domain Names

Huiju Lee, Huy Kang Kim · 2024

The existing machine learning and deep learning-based domain generation algorithm (DGA) detection methods often encounter false positives due to domain names representing non-English. To this end, we propose a DGA detection method that includes a domain name embedding approach capable of effectively representing the linguistic patterns in domain names. We here focus on Chinese domain names among non-English-based domain names. The proposed method consists of three steps as follows: (1) subword-based domain name embedding, (2) statistical feature extraction, and (3) deep learning-based detection. Experimental results demonstrate that our method overcomes misclassification of legitimate Chinese domains as DGA domains, thereby enhancing the overall performance of the DGA detection model.

Read the paper · More papers on PaperTik