A Novel Probabilistic Framework for Modeling Spelling Errors by Term Length and Frequency
Can Özbey, Hatice Altinok, Mustafa Umut Demi̇rezen · 2022
In this paper, we present a probabilistic framework that expresses spelling errors as a joint distribution of term length and frequency by means of formulating probability functions of typographic and orthographic errors separately. The parameters of the model are obtained by a root finding algorithm with respect to one another yielding the desired number of misspellings. When provided with term length and frequency information of true misspellings, the optimal parameter pair can be found by minimizing the relative entropy, and the individual rate of error types can be induced. We evaluated the model’s performance by a chi-square goodness of fit test on the empirical probability mass function of term lengths by spelling errors detected via a rule-based spell checker in three different datasets. It has been observed that the proposed approach perceives the empirical distribution of spelling errors more accurately in comparison to both word and character-based random error processes. We consider that it will be particularly useful in synthetic error generation and unsupervised spelling detection tasks.