Automatic Extraction of Frequently Confused Words in English Based on String Similarity Algorithm

Weijie Kang · IOP Conference Series Materials Science and Engineering · 2020

Abstract The calculation method for the form similarity of English words is carried out. An algorithm where the upper limit of string similarity parameters can be set is used for automatic extraction of words with similar spellings from a specified vocabulary range. The frequently confused words with similar spellings screened can enrich the English lexical knowledge base after duplicate removal and classification. The frequently confused word knowledge base is of application value in the fields of textbook writing, vocabulary training design, dictionary compilation, and real word misspelling correction.

Read the paper · More papers on PaperTik