Vehicle Classification with Audio and Video Modalities Using CNN and Decision-Level Fusion
Vijay Viswanath, Ben P. Babu · 2020
The paradigm of vehicle classification is ever evolving, with new requirements and challenges adding up at every juncture. It remains an important task in the efficient management of traffic and road infrastructure. Much work has been done over the years to perfect this task. Complementarity of information has been used before to improve classifier performance using various classifier fusion techniques. This has however been not explored in the case of a completely neural network based multi-modal classification system. In this work, a multi-modal vehicle classification system is proposed which utilizes the useful property of complementarity to achieve improved performance. Classification of vehicles is performed with the two modalities separately, using Convolutional Neural Network (CNN) classifiers. In the case of audio modality, sets of Mel Frequency Cepstral Coefficients (MFCC) are the feature vectors. The individual predictions from the base classifiers are fused at the decision-level to arrive at a final prediction. The results of both single and multi-modal classification are compared. The results show that decision-level fusion improves the accuracy of classification.