MDL-based Models for Alignment of Etymological Data
Hannes Wettig, Suvi Hiltunen, Roman Yangarber · 2011
This paper introduces several models for aligning etymological data, or for finding the best alignment at the sound or sym-bol level, given a set of etymological data. This will provide us a means of measur-ing the quality of the etymological data sets in terms of their internal consistency. Since one of our main goals is to devise automatic methods for aligning the data that are as objective as possible, the mod-els make no a priori assumptions—e.g., no preference for vowel-vowel or consonant-consonant alignments. We present a base-line model and successive improvements, using data from Uralic language family. 1