Vi-DiSC: A novel dataset and framework for extracting information from Vietnamese signposts images

Minh Tam Nguyen, Hoang Anh Tran, Truong Thinh Nguyen, Gia-Phu P. Tran, Co-Thai Quach, Trong-Hop Do · Research Square · 2023

Abstract With the development of autonomous vehicle technologies, a vision-based vehicle guidance system should be able to understand traffic signs, especially directional signposts. Directional signposts are not as specific as the warning and compulsory signs since they include many different variations. More specifically, they have special arrangements, but still follow some principles. It is difficult for captioning models to tell precisely what information the sign road includes. From the above difficulty, we propose to build a framework that simulates the process of reading a human traffic sign, which we named Vi-DiSC. The framework includes a directional sign recognition model, an optical character recognition system in parallel with a directional arrow recognition model, and finally a rule-based model based on the information extracted from the above models. We also built two Deep Learning models for the Image Captioning task to compare the generated results with the rule-based system. Our team also collects, filters, labels, and augments the directional signposts images to create the dataset for the image captioning task. In addition, we also propose TRID, a suitable metric for evaluating the directional signposts image captioning task. The framework is trained and evaluated on our dataset using various metrics, including our proposed metric TRID, and achieved promising resutls. The dataset is available at \url{https://github.com/thinhnt19393/ViDirectionSignpostCaptioning}.

Read the paper · More papers on PaperTik