TI-NERmergerV2: Automating the integration of threat intelligence NER datasets via STIX standard

Inoussa Mouiche, Sherif Saad · Computers & Security · 2026

Quality-labeled data are essential for developing accurate AI models in cybersecurity, particularly for threat intelligence named entity recognition (TI-NER), which automates the extraction of threat indicators and entities from unstructured reports. While several annotated datasets exist, their isolated use hinders scalability due to inconsistent tagging schemes, label names, and non-standard entity categories. This paper introduces TI-NERmergerV2, a robust, semi-automated framework for integrating heterogeneous TI-NER datasets into a unified, high-quality corpus aligned with the structured threat information expression (STIX) standard (e.g, STIX 2.1). Building upon its predecessor, TI-NERmerger, which is limited by its reliance on strict string matching and a narrow cyber lookup space, TI-NERmergerV2 incorporates string normalization, fuzzy fallback matching, and alias expansion using the MITRE ATT&CK knowledge base to resolve lexical variation and annotation inconsistencies. We validate its effectiveness by comparing it with a manual integration of two public datasets (DNRTI and APTNER), producing a unified dataset called AAPTNER. TI-NERmergerV2 achieves over 94% alignment with the manual process, reducing months of expert effort to minutes. Evaluations using a RoBERTa-based NER model further confirm that TI-NERmergerV2 enhances annotation quality and effectively disambiguates key entity types in the resulting DNRTI-STIX2.1 and AAPTNER datasets. The framework generalizes across datasets that adopt STIX domain and observable objects, providing a scalable and reproducible foundation for cyber threat intelligence research. Both the framework and resulting datasets are publicly released to support broader efforts in standardizing and enriching TI-NER resources.

Read the paper · More papers on PaperTik