A Knowledge-Rich Approach to Measuring the Similarity between Bulgarian and Russian Words
Svetlin Nakov, Elena Paskaleva, Preslav Nakov · 2009
nakov @ comp.nus.edu.sg We propose a novel knowledge-rich approach to measuring the similarity between a pair of words. The algorithm is tailored to Bulgarian and Russian and takes into account the orthographic and the phonetic correspondences between the two Slavic languages: it combines lemmatization, hand-crafted transformation rules, and weighted Levenshtein distance. The experimental results show an 11-pt interpolated average precision of 90.58%, which represents a sizeable improvement over two classic rivaling approaches.