Statistical Learning of Semitic Morphology Using Autosegmental Orthography

Paul Rodrigues · 2005

The root and pattern system, as well as the system of reduplication, are essential to the morphological analysis of Arabic words. (McCarthy 1979, 1981) Few computational morphology systems have been designed to parse concatenative morphology, as well as roots and reduplication simultaneously, without the help of a dictionary. By using simple statistics, we show an algorithm that can learn both the concatenative morphology as well as the roots and template. This paper shows an approach that is analogous to the the tier-based autosegmental approach developed by Goldsmith (1976), and applied to Semitic languages in McCarthy (1979). The root system is learned by comparing frequency statistics. Evidence is weighed for and against a triliteral being declared as the root. Positive evidence includes: the summation of the ratios between a letter being a root and being an affix, the summation of the frequency that a letter has shown up as a possible root, and the summation of the probabilities that the letter belongs to the root. This is then divided by the negative evidence: the summation of the probabilities that the letter is an affix, and the summation of the frequency that a letter has shown up as a possible affix yielding a "score " for the triliteral. (Elghamry, 2004) The triliteral that has the highest score is

Read the paper · More papers on PaperTik