Tail-Calibrated Transformer Autoencoding with Prototype- Guided Mining for Open-World Object Detection
Muhammad Ali Iqbal, Yeo Chan Yoon, Soo Kyun Kim · Applied Sciences · 2025
Open-world object detection (OWOD) aims to build detectors that can recognize known categories while simultaneously identifying unknown objects and incrementally learning novel classes. Despite recent advances, existing OWOD approaches still struggle with two critical challenges: the severe bias toward head classes in long-tailed data distributions and the misclassification of unknown objects as background. To address these issues, we introduce TAPM (Tail-Calibrated Transformer Autoencoding with Prototype-Guided Mining), a novel framework that explicitly enhances tail-class representation and robustly reveals unknown objects. TAPM integrates three core innovations: (1) a transformer-based autoencoder that reconstructs region features to calibrate embeddings for rare categories, mitigating the dominance of frequent classes; (2) a prototype-guided mining strategy that leverages class prototypes to localize both overlooked tail instances and candidate unknowns; and (3) an uncertainty-aware soft-labeling mechanism that assigns probabilistic supervision to pseudo-unknowns, reducing noise in incremental learning. Extensive experiments on the MS-COCO and LVIS benchmarks demonstrate that TAPM significantly improves unknown-object recall while maintaining strong known-class accuracy, achieving state-of-the-art performance across both the superclass-separated (S-OWODB) and superclass-mixed (M-OWODB) benchmarks. In particular, TAPM achieves a +20.4-point gain in U-Recall over the strong PROB baseline, underscoring its effectiveness in detecting novel objects without sacrificing mean Average Precision (mAP). Furthermore, TAPM achieves better generalization on cross-dataset evaluations, highlighting its robustness in diverse open-world scenarios.