Morphological Disambiguation of Hebrew

Danny Shacham · 2007

Morphological analysis is a crucial stage in a variety of natural language processing applications. When languages with complex morphology are concerned, even shallow applications such as search engines, information retrieval or question answering, let alone heavier applications such as machine translation, require morphological analysis and disambiguation as a first step. The lack of a morphological disambiguation module for languages such as Hebrew or Arabic handicaps the performance of many other applications. We present the HAifa morphological DisAmbiguation System (HADAS). This system uses the output of a morphological analyzer and a limited linguistic knowledge, for disambiguating Hebrew morphologically annotated text. HADAS consists of several (currently, 10) simple classifiers and a module which combines them. We build a classifier for each individual morphological feature. These classifiers are trained on feature vectors that are generated from the output of an annotated morphological analyzer, and ranks the possible analyses of each word. We investigate a number of techniques for combining and disambiguating the results produced by those classifiers. These techniques decrease the average level of ambiguity from 2.4 analyses per word to only 1.1 analyses per word. Our best result, 87.27% accuracy, was obtained using the simple supervised clasifiers, and a module for combining the results of those classifiers. This module consists of two phases; first we compose the confidence scores of all the analyses, as predicted by the classifiers. In the second phase we use a few context-dependent constraints to rule out some of the paths defined by the possible outcomes of the morphological analyzer on a sequence of words.

Read the paper · More papers on PaperTik