Vehicle Detection in Videos Leveraging Multi-Scale Feature and Memory Information

Yanni Yang, Shengnan Lu · Electronics · 2025

Vehicle detection in videos is a critical task in traffic monitoring. Existing vehicle detection tasks commonly use static detectors. Since video frames are processed as discrete static images, static detectors neglect the temporal information of vehicles when detecting vehicles in videos, leading to a reduction in detection accuracy. To address the above shortcoming, this paper improves the detection performance by introducing a video vehicle detection method that combines multi-scale features with memory information. We design a Multi-scale Feature Generation Network (MFGN) to improve the detector’s self-adaptation ability to vehicle scales. MFGN generates features with two scales and predefines multi-scale anchors for each feature scale. Based on MFGN, we propose a Memory-based Multi-scale Feature Aggregation Network (MMFAN), which aggregates historical features with current features through two parallel memory networks. The multi-scale feature and memory based method enhances the features of each frame in two perspectives, thus enhancing the vehicle detection accuracy. On the commonly adopted vehicle detection dataset UA-DETRAC, the mAP of our method is 7.4% higher compared to its static detector. The proposed approach is further validated on the well-known ImageNet VID benchmark. It demonstrates comparable performance with the memory-driven state-of-the-art frameworks.

Read the paper · More papers on PaperTik