An algorithm to align words for historical comparison
Michael A. Covington · 1996
The first step in applying the comparative method to a pair of words suspected of being cognate is to align the segments of each word that appear to correspond. Finding the right alignment may require searching. For example, Latin dō ‘I give ’ lines up with the middle dō in Greek didōmi, not the initial di. This paper presents an algorithm for finding probably correct alignments on the basis of phonetic similarity. The algorithm consists of an evaluation metric and a guided search procedure. The search algorithm can be extended to implement special handling of metathesis, assimilation, or other phenomena that require looking ahead in the string, and can return any number of alignments that meet some criterion of goodness, not just the one best. It can serve as a front end to computer implementations of the comparative method. 1. The problem The first step in applying the comparative method to a pair of words suspected of being cognate is to align the segments of each word that appear to correspond. This alignment step is not necessarily trivial. For example, the correct alignment of Latin dō with Greek didōmi is--dō-didōmi and not d ō--didōmi d--ōdidōmi----dō didōmi or numerous other possibilities. The segments of two words may be misaligned because of affixes (living or fossilized), reduplication, and sound changes that alter the number of segments, such as elision or monophthongization. Alignment is a neglected part of the computerization of the comparative method.