Generalization Analysis of the CL and MM-based

Classifications L. Kovács, Péter Barabás · 2008

Computational linguistics covers the statistical and logical modeling of languages using computer-based softwarehardware tools. An important component in CL systems is the morphological parser. The scope of our study is to build a statistical method to learn the rules of word inflection. The pre-requirement regarding the language is that the language uses words which are sequences of characters. A key factor of the required clustering algorithm is the cost efficiency. After analysis of the alternatives, two methods were selected to perform further refinement and adaptation: the observable Markov Model method and the formal concept analysis method. I. INTRODUCTION Computational linguistics (CL) covers the statistical and logical modeling of languages using computer-based software-hardware tools. An important component in CL systems is the morphological parser. A morpheme is the minimal unit with meaning in the language. The key morpheme for a concept is the stem. All of the transformations are defined on the stems. The stems should determine the base concept. The application context of the concept is given with affixes. The affixes give additional meaning of various kinds. Depending on the location of the affix, the affix unit may be called prefix, suffix, infix or circumfix [1]. The application of affixes may result in a new concept. In this case, the transformation is called derivation. If the output belongs to the same concept family, the transformation is called inflection. In the agglutinative languages, the inflections are more complex, a stem can be extended with ten or more affixes. The problem of living languages is that the dictionary describing the rules and exceptions is huge and not static. The building and updating these dictionaries is an expensive process. Our goal is to investigate the automated dictionary generation.

Read the paper · More papers on PaperTik