A Hybrid Detection Framework for Traffic Signs Using YOLOv12, Detectron2, and Vision Transformers

Antonio Makary, Issam J. Dagher · 2025

Traffic sign detection is essential for autonomous driving but remains challenging due to small sign sizes, distant signs, varying lighting conditions, and occlusions. This paper proposes a hybrid detection approach combining YOLOv12 (onestage detector), Detectron2 Faster R-CNN (two-stage detector), and a Vision Transformer (ViT) classifier. The detectors are trained on the Self-Driving Cars Dataset (SDCD) from Roboflow, and predictions are merged using overlapping box fusion. A ViT classifier further validates the detections and resolves label conflicts. Evaluated on the SDCD test set, our hybrid model achieves 97.7% mAP@50. Qualitative analysis demonstrates that the fusion significantly reduces false positives and missed detections in challenging scenarios. While this ensemble approach increases computational complexity, it markedly improves detection reliability. Future work will focus on optimizing inference speed and extending validation to larger datasets such as the Tsinghua-Tencent 100K (TT100K) dataset.

Read the paper · More papers on PaperTik