Character-Based String Matching Similarity Algorithms for Madurese Spelling Correction: A Preliminary Study
Noor Ifada, Fika Hastarita Rachman, Sri Intan Wahyuni · 2023
The Madurese language lacks of spelling correction study. Spelling correction suggests a list of corrections for misspelled words or strings. Its process consists of two steps, i.e., generating a list of suggestion strings likely to be the expected correct string and then ranking them. Henceforth, the ones considered the best match are those on the top list. This paper compares five character-based string matching similarity metrics for Madurese spelling correction. The metrics are also known as edit distance metrics, as they work by computing the similarity between two strings based on the number of edit operations needed to transform a string to another. Experiment results show that Hamming performs the best by achieving 90% correction accuracy, followed by Damerau-Levensthein at 88%, Levensthein at 86%, Jaro-Winkler at 73%, and Jaro at 72%. However, the computation time of Hamming is longer than Levensthein and more comparable to Damerau-Levensthein. Hence, implementing those three metrics is worthwhile for further study of Madurese spelling correction.