Robust Error Detection: A Hybrid Approach Combining Unsupervised Error Detection and Linguistic Knowledge
Johnny Bigert · 2002
This article presents a robust probabilistic method for the detection of context-sensitive spelling errors. The algorithm identifies lessfrequent grammatical constructions and attempts to transform them into more-frequent constructions while retaining similar syntactic structure. If the transformations result in lowfrequency constructions, the text is likely to contain an error. A first unsupervised approach uses only information derived from a part-ofspeech tagged corpus. This experiment shows a good error detection capacity but also a high rate of false alarms, in many cases due to phrase and clause boundaries. In a second approach, we combine the first method with robust phrase and clause recognition to avoid many of the false alarms in the first experiment. A comparative evaluation of the experiments shows that the introduction of linguistic knowledge dramatically increases the precision of the error detection method. 1