AADGen: Automatic Annotated Data Generation for Training Text Detection and Recognition Models
Mithilesh Kumar Singh · 2020 IEEE International Conference for Innovation in Technology (INOCON) · 2020
_One of the biggest challenges faced in deep learning community is the image data annotation, we know how data is essential for any machine learning problems but when it comes to object detection and segmentation (ODS) problems the data preparation process becomes extremely difficult and time-consuming. The data preparation becomes two steps process, first the data collection itself and second the data annotation. In this paper I present AADGen, a novel way of automatic data collection and annotation (DCA) for text detection and segmentation (TDS) problems. This method reduces the time taken for DCA to almost zero and provides error-free data annotation process. The DCA process will help us in building online learning model training by providing on demand continuous stream of training data. I have done extensive experiments on the data annotation process and achieved more than 99% annotation accuracy. I present a fast, scalable, customizable and error free data annotation approach.