Traffic Captioning: Deep Learning-based Method to Understand and Describe Traffic Images
Parniya Seifi, Abdolah Chalechale · 2022
The world around us is full of pictures. Although it seems easy for humans to compress much visual information, it is still a problem for computer systems with low output accuracy. In this paper, a method is introduced to convert traffic images into their descriptions. The presented description is grounded on prominent objects from images with deep learning models and includes three fundamental ways. First, data processing is performed on training images. Second, functional features are extracted by two deep neural networks named EfficientNet and InceptionV3. Finally, two neural networks, Gated Recurrent Unit and Transformer, are used to convert image features into text. Eventually, the optimal solution will be introduced, significantly increasing the quality of the output sentences. The MS-COCO dataset is used to evaluate the proposed methods. For this purpose, a subset including 8000 images and ten classes of traffic objects in the MS-COCO dataset are used and pre-processed. The accuracy of proposed model using BLEU evaluation is 66.9%.