The design of a neural data-oriented parsing (DOP) system
Jan Scholtes, S. Bloembergen · 2003
In a data-oriented parsing (DOP) system, sentences are parsed on the basis of language examples from a large analyzed corpus. Parsing holds the derivation of the most probable structure from (fragments) that already exist in the corpus. The DOP paradigm uses statistical features of language in combination with a structured corpus. In the neural variant of such a system, the derivation of the 'probabilities' and the storage of the corpus is done with a Kohonen feature map. The actual data-oriented parsing is performed by a regular Von Neuman machine. The model is capable of processing incomplete sentences, wrong sentences and completely new sentences which contain some amount of unknown words or structures. The implicit features of the Kohenen model make the evaluation of possible parses and the selection of the most probable one a straightforward task.>