Automatic cognate identification with gap-weighted string subsequences.
Taraka Rama · 2015
In this paper, we describe the problem of cognate identification in NLP.We introduce the idea of gap-weighted subsequences for discriminating cognates from non-cognates.We also propose a scheme to integrate phonetic features into the feature vectors for cognate identification.We show that subsequence based features perform better than state-ofthe-art classifier for the purpose of cognate identification.The contribution of this paper is the use of subsequence features for cognate identification.