Introducing UWS - A fuzzy based word similarity function with good discrimination capability: Preliminary results

João Paulo Carvalho, Luísa Coheur · 2013

This paper introduces a novel word similarity function, the Uke Similarity Function (UWS), that fuses the most interesting characteristics of the two main philosophies in word and string matching: the edit distance and the n-gram similarity approach. It also uses fuzzy sets to integrate expert knowledge about typographical errors and to easily include phonetic and token related errors. The UWS was developed with the goal of automatic detection and correction of typographical and other word errors in unedited corpus data when creating word lists.

Read the paper · More papers on PaperTik