Hybrid Methods of Natural Language Analysis

Kilian A. Foth · 2007

The automatic retrieval of syntax structure has been a longstanding goal of computer science, but one that still does not seem attainable. Two different approaches are commonly employed: 1. A method of grammatical description that is deemed adequate is implemented in a way that allows the necessary generalizations to be expressed, and at the same time is still computationally feasible. General principles of the working language are expressed within this model. Analysing unknown input then consists in computing a structure that conforms to all principles that constitute the model. Often it is also necessary to select the desired solution from several possible ones. 2. Instead of linguistic assumptions, a large set of problems already solved is used to induce a probability model that defines a total ordering on all possible structures of the working language. The parsing task is thus transformed into an optimization problem that always selects the structure that is most similar to the previous solutions in some way. The exact nature of this similarity is defined by the algorithm used for extracting the model. There are obvious and important upand downsides to both approaches. Theorydriven parsers often cannot cover arbitrary input satisfactorily because not all occurring structures were anticipated. Among the analyzable sentences, the ambiguity of the results is often very high; thousands of analyses for a sentence of a dozen words are not uncommon. Worse yet, each of these problems can be solved only at the expense of the other. Finally, this kind of grammar development requires enormous effort by experts qualified in a particular language. The currently predominant statistical approaches exhibit largely complementary features (automatic extraction of grammar rules, ambiguity resolution via numeric scores, robustness through smoothing and interpolation). However, they lack the perspicuity of the rule-based approach: • Acceptable and less acceptable structures are processed indifferently. Unusual or wrong turns of phrase are only detected insofar as analysis sometimes fails altogether.

Read the paper · More papers on PaperTik