Towards Indian language spell-checker design

B.B. Chaudhuri · 2003

This paper deals with the development of a spell-checker in Indian languages using as an example Bangla, the second most popular language on the Indian Subcontinent. A brief review of problems and the current scenario of Indian language spell-checkers is described. The approach for the Bangla spell-checker is then elaborated. In this approach the technique works in two stages. The first stage takes care of phonetic similarity error. For that the phonetically similar characters are mapped into single units of character code. A new dictionary D/sub c/ is constructed with this reduced set of alphabets. A phonetically similar but wrongly spelt word can be easily corrected using this dictionary. The second stage takes care of errors other than phonetic similarity. A wrongly spelt word S of n characters is searched in the dictionary D/sub c/. If S is a nonword, its first k/sub 1//spl les/n characters will match with a valid word in D/sub c/. (if k/sub 1/=n then the word in D/sub c/ must be longer than n). A reversed word dictionary D/sub r/ is also generated where the characters of the word are maintained in a reversed order. If the last k/sub 2/ characters of S match with a word in D/sub r/ then, for a single error, it is located within the intersection region of first k/sub 1/+1 and last k/sub 2/+1 characters of S. We observed that this region is very small compared to word length for most cases and the number of suggested correct words can be drastically reduced using this information. We have used our approach in correcting Bangla text, where the problem of inflection is tackled by a simplified version of a morphological analyser. Another problem encountered in Indian languages is the existence of a large number of compound words formed by euphony and assimilation. The problem of compound words is also carefully tackled.

Read the paper · More papers on PaperTik