A Comparison of Spelling-Correction Methods for the Identification of Word Forms in Historical Text Databases

A. M. ROBERTSON · Literary and Linguistic Computing · 1993

This paper discusses the application of algorithmic spelling—correction techniques to the idenficiation of those words in databases of sixteenth-, seventeenth-, and eighteenth-century English texts that are most similar to a query word in modern English. The experiments involved the n-gram matching, phonetic and non-phonetic coding, and dynamic-programming methods for spelling correction, and demonstrate the general effectiveness of this approach to the identification of historical word forms. The best results are given by the editcost-distance and longest-common-subsequence methods, which use a dynamic-programming algorithm; however, these are only slightly more effective than the digram-matching method, which is far faster in operation

Read the paper · More papers on PaperTik