Advances in automated language classification

Eric W. Holman, Søren Wichmann, Cecil H. Brown, Viveka Velupillai, Andre Matthias Müller, Dik Bakker · UvA-DARE (University of Amsterdam) · 2008

The paper presents a method for the automatic reconstruction of language relationships taking the Swadesh (1955) 100-item word list as a point of departure.However, the method differs from the original lexicostatistical approach in two fundamental ways.First, the comparison between word forms is done by a computer program (ASJP; automated similarity judgment program) on the basis of Levenshtein's (1966) algorithm, resulting in a distance matrix between individual languages.And second, graphic branching structures illustrating language relatedness (family trees) are generated from this matrix by the way of standard software and algorithms originally developed for the use of biologists in studying phylogenetic relationships (Huson 1998).To accommodate wordlists originally published in a variety of more or less simplified orthographies, a special alphabet, called ASJPcode, was devised which makes use of the QWERTY keyboard symbols only.It contains just 34 consonant symbols and 7 symbols for vowels.These symbols are used for phonological segments defined by the most common points and manners of articulation.Rarer segments are represented by the symbol they most closely resemble in terms of point and manner of articulation.See Brown et al (to appear 2008) for details.Unlike most other approaches to automatic language classification, such as those described by Oswalt (1971), Atkinson et al. (2005), and Nakleh et al. (2005), the present method automates both the judgments of cognacy and the subsequent inference of phylogeny.We can therefore apply the same objective criteria worldwide to classify an unusually large sample of languages.This facilitates the large scale statistical study of overlaps in lexicons between languages and may reveal previously unknown phylogenetic relationships.To date, we have collected and transcribed a basic word set for close to 2000 languages of the world.The nearly 2 million language pairs in the database are compared by means of the Levenshtein Distance (LD: see Levenshtein 1966).For any pair of words represented in ASJPcode, LD is defined as the minimum total number of additions, deletions, and substitutions of symbols necessary to transform one word into the other.For any pair of languages L1 and L2, first the LD values are established for each of the N Swadesh words that L1 and L2 share (virtually always the full set that we consider).These LD values are then normalized by dividing each LD by its theoretical maximum giving the normalized LD (LDN).Finally, since lexical similarity may be influenced by chance resemblances, such as an overlap in the phoneme inventories or shared phonotactic preferences for the two languages involved, we correct each LDN by dividing by it the mean LDN of all N(N-1)/2 pairings of words with different meanings,

Read the paper · More papers on PaperTik