FILLING THE GAPS USING GOOGLE 5-GRAMS CORPUS
Costin-Gabriel Chiru, Andrei Hanganu, Traian Eugen Rebedea, Ștefan Trăușan-Matu · 2010
In this paper we present a text recovery method based on a probabilistic post-recognition processing of the output of an Optical Character Recognition system. The proposed method is trying to fill in the gaps of missing text resulted from the recognition process of degraded documents. For this task, a corpus of up to 5grams provided by Google is used. Several heuristics for using this corpus for the fulfilment of this task are described after presenting the general problem and alternative solutions. These heuristics have been validated using a set of experiments that are also discussed together with the results that have been obtained.