Hybrid deep learning ensemble model for improved large-scale car recognition
Abhishek Verma, Yu Liu · 2017
Smart video based traffic monitoring and surveillance systems for improved security rely upon sophisticated deep learning based computer vision algorithms. In this paper we propose a novel deep learning model to automatically recognize cars on large-scale grand challenge CompCars dataset. Such a task for a computer is difficult due to its fine-grained nature and achieving recognition accuracy close to human experts remains a challenge due to lack of big dataset and machine learning model to detect delicate nuances. Our proposed deep CNN hybrid architecture outperforms previously published classification accuracy by 2.52%, which is a significant improvement considering the challenging nature of the dataset. Our method suggests several novelties and advantages over existing methods: First, it uses the GoogLeNet's key architecture - inception modules to efficiently exploit the inception's dimension reduction power and to lower the network cost. Second, inspired by the VGG's uniform and powerful architecture, the method replaces GoogLeNet's auxiliary classifiers into deeper networks with 3×3 convolution components to increase its recognition capability. Third, it is a powerful and efficient network in the way that it represents the ensembles of multiple short and medium depth networks. We believe our method could be useful in other domains that require finegrained recognition.