Word Suggestions for non-word Text Errors using Similarity Measure
Kavita Tukaram Patil, R. P. Bhavsar, B. V. Pawar · 2021
Spelling errors are a common phenomenon in any text writing system. It is due to various factors like memorization, spelling construction and training, and word ambiguity of that language. With the advent of machines keyboard layout also one of the causes for spelling mistakes. Natural Language Processing (NLP) is a prominent research area in the territory of human language. A spelling checker is a key aspect of many applications such as the MT framework, information retrieval, Desktop applications, and Office automation framework, etc. There are various approaches to spelling checking like rule-based, data-driven, etc. With the commencement of machine learning, a new dimension was introduced as a similarity measure. In this paper, we proposed the cosine similarity measure which handles the nonword error in the Marathi language text. As cosine similarity measure is an important strategy of spelling checking by using this strategy, we enhanced the quality of suggestions for non-word errors. This algorithm also boosts the user's proficiency in the situation when the users impotent to measure the accurate spelling by themselves. We have harvested about 9, 29, 663 unique words from various sources. The system results in suggestion generation accuracy of about 86.61%.