Fusion of audio visual cues for vehicle classification
C Divina Paul Daniel, Leena Mary · 2016
Road planning and traffic monitoring is conducted based on survey of traffic volume. In recent years many of researchers have developed vision and audio based techniques for detection and classification of moving vehicles. Audio based technique suffers from low accuracy but has low computational cost. Then there is visual based approach which has significantly higher accuracy but demands high computational resources. This paper proposes a new approach which utilizes both audio and video of traffic data to perform traffic volume survey. Vehicle detection can be done from audio signal. Video frames around audio peaks are selectively extracted. Then visual feature vectors are extracted from the binary image of the vehicle. Audio features represented using Mel Frequency Cepstral Coefficients (MFCC) are extracted from the regions around the vehicle peak. Classification is done using multilayer feed forward neural network which gave an overall classification accuracy of 92.67% for seven vehicle classes with the chosen set of audio visual features.