Information Extraction From Images Using Pytesseract and NLTK
Akash V Pavaskar, Akshay S Achha, Anoop R Desai, K L Darshan · Journal of Emerging Technologies and Innovative Research · 2017
Images are used in various fields such as advertisements, business purpose, and spreading awareness. Text data present in these images contain useful and helpful information like contact details, hyperlinks, QR codes. Extraction of this information involves detection, localization, tracking, extraction, enhancement, and recognition of the text from a given image. However, variations of text due to differences in size, style, orientation, and alignment, as well as low image contrast and complex background make the problem of automatic text extraction extremely challenging in the computer vision research area. But the difficulty in implementation proves to be useful and fruitful. This project aims at using computer vision (Pytesseract) to extract useful information like text, contact details and hyperlinks from images. The android based app would allow user to upload a photo and enable user in storing the contact details, set a remainder, provide summary of the content of the image, opening of hyperlinks directly from the app without needing to type the URL inside the browser. Thus, making the images a more productive and making the job of the user more easy and convenient.