Devising a Distinguishable Feature Set for Sinhala and English Script Separation on Social Media Images

K. S. A. Walawage, L. Ranathunga · 2020

In the last decade, social media has become the most popular and ubiquitous facilitator for creating, and sharing content, accessing information, and connecting people around the world. Nowadays, social media content generation is rapidly growing in the Sri Lankan community. Some content generation was a threat to peace, which led to unpleasant situations in the country within the last two years. Content moderation is the best solution for this problem, but manually processing millions of posts is a difficult task. For automation, separation and identification of Sinhala and English character in social media images is an essential project in content identification. In this work, Sinhala and English characters are separately identified from image posts and video thumbnails which are on public Facebook pages. This paper introduces a novel algorithm to separate Sinhala and English characters. It has utilized topological features, contour-based technique, and water reservoir-based technique. The proposed algorithm achieved 93% accuracy. Separated characters are recognized using the Convolutional Neural Network in TensorFlow framework.

Read the paper · More papers on PaperTik