FlaskOCR: Building a Framework for Text Extraction From Image
Puneet Puneet, Deepika Deepika, Abhishek Kumar · 2025
Text recognition has become a favored topic in the field of computer vision. Research in this field has great potential to benefit many domains worldwide. Examples include text visualization, document classification, and license plate identification. Many researchers working on this topic have developed methods for identifying text using Optical Character Recognition (OCR). OCR serves as a diagnostic tool in fields such as machine vision, pattern recognition, and machine intelligence. Handwriting recognition, a key area of OCR, is widely used to evaluate a computer's ability to translate human-written text into digital form. This can be achieved through scanning handwritten text or typing directly into a designated input field. Although OCR has nearly solved the problem of text recognition for Latin characters, challenges remain for non-Latin scripts. Recognizing non-Latin text is particularly difficult because its appearance often differs significantly in contour and layout compared to the Latin text. This study aims to compile a comprehensive collection of OCR databases. A total of 5,900 characters were prepared and processed in various ways using Tesseract OCR tools. The FlaskOCR project was developed to create an OCR system capable of recognizing both online and offline handwriting.