ColLex.en: Automatically Generating and Evaluating a Full-form Lexicon for English
Tim vor der Brück, Alexander Mehler, Zahurul Islam · 2014
The paper describes a procedure for the automatic generation of a large full-form lexicon of English.We put emphasis on two statistical methods to lexicon extension and adjustment: in terms of a letter-based HMM and in terms of a detector of spelling variants and misspellings.The resulting resource, ColLex.EN, is evaluated with respect to two tasks: text categorization and lexical coverage by example of the SUSANNE corpus and the Open ANC.