Statistics and phonotactical rules in finding OCR errors
Stina Nylander · DSpace repository (University of Tartu) · 1999
This report describes two experiments in finding errors in optically scanned Swedish without lexicon.First, statistics were used to find unexpectedly frequent trigrams and correction rules were created.Second, Bengt Sigurds model of Swedish phonotax was used to detect words with phonotactically illegal beginning or end.The phonotax did not perform as well as the statictic rules did on their training material, but outscored them by far on new text.A correction tool was created with the phonotax as means of error detection.The tool displays every occurrence of an error string at the same time and gives the user the possibility to give different corrections to each occurrence.This work shows that it is possible to find errors in optically scanned text without relying on a lexicon, and that word structure can provide useful information to the correction process.