Four types of context for automatic spelling correction.
Michael Flor · 2012
ABSTRACT. This paper presents an investigation on using four types of contextual information for improving the accuracy of automatic correction of single-token non-word misspellings. The task is framed as contextually-informed re-ranking of correction candidates. Immediate local context is captured by word n-grams statistics from a Web-scale language model. The second approach measures how well a candidate correction fits in the semantic fabric of the local lexical neighborhood, using a very large Distributional Semantic Model. In the third approach, recognizing a misspelling as an instance of a recurring word can be useful for reranking. The fourth approach looks at context beyond the text itself. If the approximate topic can be known in advance, spelling correction can be biased towards the topic. Effectiveness of proposed methods is demonstrated with an annotated corpus of 3,000 student essays from international high-stakes English language assessments. The paper also describes an implemented system that achieves high accuracy on this task. RÉSUMÉ. Cet article présente une enquête sur l’utilisation de quatre types d’informations contextuelles pour améliorer la précision de la correction automatique de fautes d’orthographe de mots seuls. La tâche est présentée comme un reclassement contextuellement informé. Le contexte local immédiat, capturé par statistique de mot n-grammes est modélisé à partir d’un modèle de langage à l’échelle du Web. La deuxième méthode consiste à mesurer à quel point une correction s’inscrit dans le tissu sémantique local, en utilisant un très grand modèle sémantique distributionnel. La troisième approche reconnaissant une faute d’orthographe comme une instance d’un mot récurrent peut être utile pour le reclassement. La quatrième approche s’attache au contexte au-delà du texte lui-même. Si le sujet approximatif peut être connu à l’avance, la correction orthographique peut être biaisée par rapport au sujet. L’efficacité des méthodes proposées est démontrée avec un corpus annoté de 3 000 travaux d’étudiants des évaluations internationales de langue anglaise. Le document décrit également un système mis en place qui permet d’obtenir une grande précision sur cette tâche.