Kannada Manuscript Digitization through OCR and Machine Learning
Asha Rani Borah, Abhilash Vijapur, B. Mahesh Kumar · 2024
This research paper deals with the challenges of digitizing ancient Kannada manuscripts, addressing all the complex features of the Kannada script which has different types of handwriting with different styles depending on the historical context. The main focus is on development of a special Optical Character Recognition (OCR) model which can decipher Kannada characters and convert them into digital characters. This process follows CNN algorithm which accepts pre-processed data and predicts each character based on the trained and tested database. Another idea is to predict the approximate age of the manuscript depending on the linguistic features of the Kannada that is written in different eras. Through this, the research can preserve and make old Kannada manuscripts more accessible ensuring future generations can access these manuscripts.