The Relevance of Frequency Lists for Error Correction and Robust Lemmatization

René Schneider, Ingrid Renz · 2000

In this paper we discuss the usefulness of frequency lists and the impact they have on a learning algorithm, named rank-and-similarity-based learning. The combination of frequency lists with a simple similarity measure leads to significant results that are useful for bootstrapping a frequency dictionary in which each lexical entry provides information about a stem and its well- and illformed variants. The modification of the frequency lists allows the construction of a collocation measure, determining the syntagmatic relationship of two or more words in a given domain. The results of the algorithm are applied to two different problems that arise in the area of information extraction from paperbound documents, namely error correction, and robust lemmatization. The paper finishes with some remarks concerning the validity and evaluation of the results.

Read the paper · More papers on PaperTik