Sinhala Text Extraction from YouTube Thumbnails using Convolutional Spiking Neural Networks

Heshan De Silva, Supunmali Ahangama, Suresha Perera · 2021

Social media platforms have been fueled by the advancement of Information Technology and the human impulse to communicate. YouTube has grown to become one of the most popular video sharing and social media platforms, with users able to watch, like, share, comment on, and create their own videos. An increasing body of research emphasizes the significance of YouTube data, where billions of monthly viewers and millions of users talk and share their views. Among those data, thumbnails play a vital role where users can select their own or automatically generated image, which provides an overview of a video to the viewers when uploading a video. Those thumbnails images have a significant value in social media data analytics. However, text information extraction from thumbnail images has become more challenging as the text is normally printed against a complex background. This becomes more challenging with the Sinhala language. In this paper, we proposed a novel approach to extract Sinhala text rapidly from YouTube thumbnail images using Convolutional Spiking Neural Networks. The network consisted of three convolutional layers. The rate base convolutional spiking neural network approximation is used to train the network. The proposed solution can be divided into three main steps, pre-processing, prediction and post-processing. This method extracts the Sinhala text from thumbnails with an accuracy of approximately 85% in a short period of time.

Read the paper · More papers on PaperTik