Automatic dictionary organization in NLP systems for Oriental languages

V. Andrezen, Lior Kogan, W. Kwitakowski, Rinad S. Minvaleev, Rajmund G. Piotrowski, V. Shumovsky, E. Tioun, Yu Tovmach · 1992

This paper presents a description of automatic dictionaries (ADs) and dictionary entry (DE) schemes for NLP systems dealing with Oriental languages. The uniformity of the AD organization and of the DE pattern does not prevent the system from taking into account the structural differences of isolating (analytical), agglutinating and internal-flection languages.The "Speech Statistics" (SpSt) project team has been designing a linguistic automaton aimed at NL processing in a variety of forms.In addition to Germanic and Romance languages the system under development is to handle text processing of a number of Oriental languages. The strategy adopted by the SpSt group is characterized by a lexicalized approach: the NLP algorithms for any language are entirely AD dependent, i.e., a large lexicon database has been provided, its entries being loaded with information including not only lexical, but also morphological, syntactic and semantic data. This information concentrated in dictionary entries (DEs) is essential for both source text analysis and target (Russian) text generation.The DE structure is largely determined by the typological features of the source language. The SpSt group has hitherto had to deal with European languages and it was for these languages (inflective and inflective-analytical) that the prototype entry schemes were elaborated and adopted. No doubt, the typological characteristics of Oriental languages required certain modifications to be made to the basic scheme. Hence in the present paper each of the language types is given consideration. Agglutinating languages proved to be the most suitable to process according to the SpSt strategy. But an isolating language will be the first to be proposed for discussion.

Read the paper · More papers on PaperTik