Stochastic methods for resolution of grammatical category ambiguity in inflected and uninflected languages

Steven J. DeRose · Americanae (AECID Library) · 1990

Grammatical category ambiguity (distinct from semantic and structural ambiguity) is extremely frequent in natural language. In the Brown Corpus (a million-word grammatically tagged sample of English prose) 11% of all word forms, or 48% of all word instances, occur as members of more than one grammatical category. These figures greatly under-represent actual categorial ambiguity, for instance because uncommon words may seem unambiguous when they are actually not. Such frequent ambiguity poses extreme problems of non-determinism for parsers. Therefore means of resolving such ambiguities are important to the progress of natural language processing systems. This thesis examines probabilistic strategies for resolving categorial ambiguity. I consider the contextual probability of a given category, given a categorial context, and the relative probability that a given word form represents a particular category. Up to 96% of all words can be assigned the correct category without morphological analysis, special handling of idioms, or other non-probabilistic features. Dynamic programming yields disambiguation time directly proportional to text length. Probabilistic methods are thus both faster and more accurate than previous methods, and overcome the non-determinism which renders many other methods unworkable. I apply these methods to the Brown Corpus (English) and to the Greek New Testament (140,000 words of Koine Greek). I discuss the effects of various parameters on the accuracy of category assignment, and analyze the types and frequencies of residual errors. I report control studies which help to predict the algorithm's effectiveness for unrestricted text, and investigate the amount of normalization text required to obtain reliable probability estimates. Analyses of related information-theoretic properties of natural language corpora are also included, for example, investigations of the effect of sample size on measurement of entropy.

Read the paper · More papers on PaperTik