AI4TEN: Synthetic-to-Real Transfer for Acoustic Vehicle Classification Using Physics-Based and AI-Generated Training Data
Jibran Rasheed Khan, Ivan Ryzhikov, Mikko Kolehmainen · Applied Sciences · 2026
Traffic noise monitoring increasingly relies on automated vehicle type classification from audio recordings for source-specific noise analysis. However, training acoustic classifiers requires large, labeled datasets that are costly to collect. This study compares two synthetic data generation approaches for training a CNN-based vehicle classifier (car, truck, motorcycle), physics-based simulation (pyroadacoustics) and AI text-to-audio generation (AudioLDM). A cross-dataset evaluation protocol trains on one dataset (1006 real samples, supplemented with synthetic data) and tests on two independent datasets (1019 samples) recorded in different countries with different microphones. Eight training configurations are evaluated across five random seeds. Combining real data with both synthetic sources achieves a balanced F1 of 0.39, a 55% improvement over the real-only baseline. A controlled experiment with matched sample sizes (200 per class) shows that the two synthetic sources perform comparably as standalone training data. When combined with real recordings, physics-based augmentation retains a modest advantage (F1 = 0.29 vs. 0.24), suggesting that parametric variation contributes beyond sample size effects. A spectral analysis reveals that AudioLDM produces sounds closer to real recordings, yet combining both sources provides the largest gain. A congestion stress test reveals 43% F1 degradation with two overlapping vehicles, accompanied by increasing model confidence. All data, code, and trained models are publicly available.