Document image retrieval system using character candidates generated by character recognition process
Shuji Senda, Michihiko Minoh, Kayo Ikeda · 2002
The authors have implemented a document image retrieval system; it automatically stores document images in the form of character candidates obtained by character recognition process. When we give some keywords to the system, it can retrieve the images containing at least one of the keywords. The strategy of using character candidates remarkably lowers the rate of missing the keywords on retrieval, because it includes several hypotheses in character segmentation and in character recognition. For finding the keywords from the storage of the character candidates, the authors have developed a fast and efficient searching algorithm. It is able to locate all occurrences of any finite number of keywords in the character candidates. The improvement of using character candidates and the efficiency of the searching algorithm are also described.>