A Crowdsourcing Data Annotation System For Vietnamese Scene Text Detection

Bich-Van Nguyen-Thi, Triet Le Phan, Thua Nguyen, Thanh Duc Ngo · 2022

Most of the recent breakthroughs in AI-related research and applications are data driven. Annotated data plays the main role in improving AI/ML-based solutions. The lack of high-quality annotated data is one of the main obstacles to AI adoption, especially in Vietnamese scene text detection. Additionally, manual annotation is expensive yet not scalable. This project aims to tackle the problem by developing a crowdsourcing data annotation system at word-level annotations, vCaptcha, in the form of an anti-bot widget to integrate to some web pages. After that, results can be derived from the workers prior to evaluating and aggregating the crowdsourced labels. Also, we propose a novel human-in-the-loop approach for incorporating the state-of-the-art scene-text detection models and crowdsourcing system so that in the end, a fully annotated dataset with optimal and high quality can be generated with significant cost and time savings. Usage document to integrate vCaptcha to any web pages is available at https://bichvan2810.github.io/vCaptcha-end-user-info.

Read the paper · More papers on PaperTik