DNS Detection Under Distribution Shift: Stable, Calibrated, and Deployable
Nizamuddin Maitlo, Samina Rajper, Nooruddin Noonari, Kaleem Arshid, Hafsa Iqbal · IEEE Access · 2026
Security Operations Centers (SOCs) increasingly use Domain Name System (DNS)-layer detectors to triage malicious domains under strict alert budgets. In this setting, excellent ranking scores do not necessarily imply deployable performance because the decision threshold must operate at ultra-low false-positive rates (FPRs). This study proposes a deployment-oriented DNS detection framework using an author-collected malicious-domain dataset. The raw file contains 2,000,000 domain records and six metadata fields: Domain, Label, threat_family, first_seen_utc, last_seen_utc, and source. The additional timestamp and source fields support provenance auditing, while the threat_family field is treated as a proxy annotation rather than analyst-verified forensic attribution. The processed feature file contains 1,048,575 samples, 23 numerical domain features, and a near-balanced benign/malicious class distribution. We audit duplicates, missing values, label structure, and rounded-hash proxy groups; the processed file contains 471,470 exact feature-label duplicates and 576,294 unique proxy groups, with zero train-test overlap under the primary rounded-hash split. XGBoost, multilayer perceptron (MLP), and Feature Tokenizer Transformer (FT-Transformer) baselines are compared using Area Under the Receiver Operating Characteristic Curve (AUROC), Area Under the Precision-Recall Curve (AUPRC), Expected Calibration Error (ECE), Brier Score, Negative Log-Likelihood (NLL), and Recall@FPR at 1.0%, 0.5%, and 0.1%. Under the rounded-hash group split, XGBoost achieves AUROC = 0.999934 and Recall@FPR = 0.001 of 0.992590, outperforming MLP and FT-Transformer at the strictest operating point (all differences statistically significant at $p\lt 0.001$ , $B=300$ bootstrap repetitions). Under an intentionally severe high-drift stress test (mean Kolmogorov-Smirnov (KS) statistic = 0.578, Population Stability Index (PSI) = 4.897), Recall@FPR = 0.001 collapses to 0.118 for XGBoost and to zero for MLP, demonstrating that deployment reliability cannot be inferred from AUROC alone. We also evaluate temperature scaling, Platt scaling, isotonic regression, Quantile-Matched Thresholding (QMT), and Entropy-based Test-Time Temperature (T-TTT), and report runtime, memory, bootstrap confidence intervals, SHapley Additive exPlanations (SHAP)-based feature importance, and an attack-spike stress test. The results show that QMT stabilizes alert budgets but suppresses recall when genuine malicious prevalence surges; it should therefore be deployed with SOC guardrails and prevalence monitoring.