Descriptor: Multilingual Visual Font Recognition Dataset

Moshiur Rahman Tonmoy, Md. Akhtaruzzaman Adnan, Aloke Kumar Saha, M. F. Mridha, Nilanjan Dey · IEEE data descriptions. · 2024

With advancements in deep learning (DL) and computer vision-based applications, the visual font recognition (VFR) field has evolved rapidly. From browser extensions to mobile and web apps, several efficient systems now exist for identifying fonts from images. However, progress in languages other than English has been limited, largely due to insufficient data availability. To address this obstacle, we created the multilingual visual font recognition (MVFR) dataset, an image dataset for the VFR domain encompassing four different languages: Bangla, Hindi, Russian, and Spanish. Our MVFR dataset comprises 50 000 images for recognizing ten distinct font styles for each language, resulting in a substantial corpus of 200 000 VFR images overall by accumulating four languages. Furthermore, we have also provided our developed language-agnostic Python generator script employed to generate the dataset, which can be extended to generate VFR image data for other languages, fueling the advancements of the VFR domain in languages with limited resources.IEEE SOCIETY/COUNCILComputational Intelligence Society (CIS)DATA TYPE/LOCATIONImagesDATA DOI/PID10.17632/cnd2wh65my.1

Read the paper · More papers on PaperTik