Error detection in character recognition using pseudosyllable analysis

Ricardo Alexandrino Garcia, Yannis A. Dimitriadis, F. Merino Pastor, Juan López Coronado · 2002

In modern document management systems it is difficult to include large vocabularies (more than 150,000 words long) to detect on-line errors. The main drawback lies in the manipulation of the great amounts of data. This difficulty becomes critical if the system incorporates character recognition modules. In this paper we propose a new technique that stems from a written text segmentation based on phonetical and etymological criteria. The procedure we use integrates dictionary n-gram techniques. It checks whether the given character sequence matches a sequence of pseudosyllables (using a dictionary) and simultaneously checks if the pairs of pseudosyllables is admissible (through n-gram techniques). The results obtained from the proposed method using lists of words of various sizes, as well as a corpus in Spanish, are better than the n-gram methods typically used in error detection. Furthermore, it requires less memory and processing time, as compared with dictionary look-up methods.

Read the paper · More papers on PaperTik