Towards a new type of morphemic analysis
Eva Koktová · 1985
The present paper provides a report on a new system of an automated morphemic analysis of technical texts in Czech as a highly inflectional language, which is being prepared by the linguistic team of the Faculty of Mathematics and Physics in Prague, within the project of man-machine communication without a pre-arranged data base (TIBAQ). The kind of morphemic analysis presented here is based on a retrograde (right-to-left) analysis of words by means of morphemically unambiguous or irresolvably ambiguous word-ends, which do not coincide with the etymological word-endings but correspond to the structure of the accidental cases of morphemic ambiguity in an inflectional language (word-endings being accountable for in a certain way by word-ends). The algorithm of analysis can thus dispense with any dictionary (of morphemic irregularities and exceptions), economically accounting especially for productive word-endings. The word-ends of the analysis are assigned several kinds of morphemic information, concerning morphemic categories and lemmatization. The analysis is based on the absolute frequency of word-ends in technical texts and is able to interact with the semantic analysis.