Stop word detection in compressed textual images: An experiment on indic script documents
Utpal Garain, Amit Kumar Das · Proceedings - International Conference on Pattern Recognition/Proceedings/International Conference on Pattern Recognition · 2008
Stop word detection is attempted in this work in the context of retrieval of document images in the compressed domain. Algorithms are presented to identify text lines and words and to cluster similar words to count word occurrence frequencies. A list of words with their occurrence frequencies is generated from a corpus of textual images. As stop words in any language show high occurrence frequencies, such words occupy the upper positions in the sorted word list. Experiments have been carried out on two major indic scripts (Devanagari (Hindi) and Bangla). Test results using 150 document images consisting of about 12 K words in each script show the promising potential of the proposed approach.