Kannada ImageNet: A Dataset for Image Classification in Kannada
G. Ramesh, W. Srihari, Atul M. Bharadwaj, H N Champa · 2021
Most of the ongoing research for image classification happens in English and there is very little research about classification of images in regional Indian languages like Kannada. There are several uses and advantages of classifying images natively in Kannada. A diverse dataset of 425,723 images which is accurately labelled in Kannada is obtained through this study, which adds to the contribution this paper makes to image classification literature, that is verified and corrected by a human judge. It consists of 1,083 unique classes which was constructed by a combination of randomly selecting classes from ImageNet, a database commonly used for Image classification in English as well as 193 manually thought out classes. The classes were translated first using an online translation service and later corrected by a human judge to achieve a translation accuracy of 97.15% on the entire dataset.