Automated Extraction of Company Names from Product Label Images: A Text Mining Approach
Gaurav Kholiya, Vimal Joshi, Sanyam Pandey, Sashank Thapa, Vikrant Sharma, Satvik Vats · 2024
Nowadays, it is becoming more and more important to get information from various sources. One of the key details that provide us with information about the product is its name, which is prominently printed on it. In this study, we have introduced an autonomous text-mining technique to extract the company name from the images. For machine learning and natural language processing applications, this method uses the Python language and packages like scikit-learn, py-tesseract, and NLTK. The process involves employing stemming, tokenization, stop-word removal, and optical character recognition (OCR) to prepare the text that was extracted from the image. Then, using the TF-IDF vectorizer, the text is transformed into numerical vectors. Cosine similarity is used to compare the vectorized input text to a corpus of previously processed paragraphs that have the well-known description associated with different organizations. The name of the business that most closely matches the text is the outcome. The recommended method appears to have potential for accurately identifying the brand name from images of product labels, facilitating efficient information retrieval and assisting consumers in making decisions. This research can be helpful to companies in the analysis of a certain product in a given location in relation to other brands and how their product item is doing in terms of consumption in that area may also be valuable to companies.