FT-Transformer-Based IoT Network Attack Detection and Cross-Dataset Generalization Analysis
Fapeng Li, Yatong Tao, Leilei Qu · Electronics · 2026
With the large-scale deployment of Internet of Things (IoT) devices in smart homes, smart healthcare, and industrial internet scenarios, network attacks against IoT environments have become increasingly sophisticated, making reliable intrusion detection increasingly important. Focusing on tabular traffic statistical features, this study systematically evaluates an FT-Transformer-based IoT network attack detection framework across primary-dataset classification, feature-aligned external validation, feature-union validation, standardized NetFlow external validation, domain-aligned training, multi-class classification, and SHAP-based interpretability analysis. CICIoT2023 is used as the primary dataset for binary and multi-class attack detection, while CICIoMT2024 is used for feature-aligned external validation. In addition, NF-ToN-IoT and NF-BoT-IoT are introduced as standardized NetFlow IoT datasets to provide an additional external validation scenario. Random Forest, XGBoost, MLP, TabNet, and a TabTransformer-style numerical Transformer baseline are included for comparison. Experimental results show that FT-Transformer achieves competitive performance on the CICIoT2023 binary classification task, with an accuracy of 0.980675, an attack-class F1-score of 0.980366, a ROC-AUC of 0.995014, and a PR-AUC of 0.996261. Under the controlled 38-feature aligned CICIoT2023-to-CICIoMT2024 validation setting, FT-Transformer shows better ROC-AUC stability than Random Forest and XGBoost. However, the feature-union validation experiment reveals that this advantage does not necessarily extend to less constrained feature-space settings, indicating that FT-Transformer should not be interpreted as a universal cross-dataset generalization solution. In the additional standardized NetFlow NF-ToN-IoT-to-NF-BoT-IoT validation experiment, FT-Transformer achieves the strongest external validation performance among the compared models. Furthermore, the domain-aligned FT-Transformer-CORAL experiment shows that explicit source-target representation alignment can improve external validation performance, especially in terms of ROC-AUC. For the multi-class task, FT-Transformer achieves an accuracy of 0.809554, a Macro-F1 score of 0.724006, and a Weighted-F1 score of 0.806680. SHAP analysis further indicates that key features such as Number, HTTPS, Header_Length, Rate, and AVG have meaningful correspondence with known attack behaviors. Overall, this study provides a systematic and reproducible empirical evaluation of FT-Transformer for tabular IoT network attack detection. The results suggest that meaningful cross-dataset robustness requires not only suitable tabular model architectures but also carefully designed validation protocols and representation-learning strategies.