A Partitional Clustering Approach to Persian Spell Checking
Fatemeh Hojati Kermani, Shirin Ghanbari · 2019 5th Conference on Knowledge Based Engineering and Innovation (KBEI) · 2019
Spell checkers find, suggest and correct incorrect words within documents. Complexities of languages increases the challenge in-hand. Due to the challenges of the Persian language, preprocessing and processing spell checkers requires further analysis and natural language processing. For the past decade, the majority of methods that identify errors in documents have been lexicon-based or probability-based. This paper describes a new approach for Persian spell-checking using a clustered dictionary to save time and storage. Clustering is optimally generated through k-medoids and during processing misspelt words are compared to the centers of the generated clusters rather than the entire dictionary. The proposed spell checker is evaluated using the Faspell spelling error corpus that contains Persian misspellings and is shown to provide suggestions with significant accuracy rates.