Ambiguity in language learning: computational and cognitive models
Hinrich Schütze · Stanford University eBooks · 1996
This dissertation is concerned with how ambiguity and ambiguity resolution are learned, that is, with the acquisition of the different representations of ambiguous linguistic forms and the knowledge necessary for selecting among them in context. Despite much separate research on acquisition and ambiguity, there is little work on how successful acquisition is possible for ambiguous forms. The dissertation presents three models of ambiguity acquisition. In TAGSPACE, a model of syntactic categorization, unsupervised classification groups tokens into syntactic categories according to distributional similarity. An evaluation on the Brown corpus demonstrates successful acquisition of the major syntactic categories of English. In WORDSPACE, a model of semantic categorization, semantic categories are induced by unsupervised classification of associational representations of contexts. The model achieves high accuracy when evaluated on a set of ambiguous words in a New York Times corpus. The third acquisition problem addressed is subcategorization. A connectionist model predicts subcategorization frames from lexicosemantic representations. The model learns to form internal verb representations depending on context (i.e. disambiguate) and to generalize the subcategorization behavior of 178 English verbs in the Dative Alternation class. Ambiguity is important for theories of linguistic representation. Disambiguation is difficult with symbolic representations since all-or-none criteria like grammaticality or soundness eliminate few readings. In contrast, disambiguation in both TAGSPACE and WORDSPACE relies crucially on proximity relations that can only be modeled by gradient representations. In subcategorization learning, gradience solves the transition problem, the question of how the child makes the gradual transition from a state with little knowledge to adult performance. Ambiguity is also relevant for linguistic innateness. Innate categories do not explain learnability without an account of how they are grounded in perception. The grounding problem is hard for ambiguous forms with their many possible groundings. It is shown here that--contrary to much skepticism in contemporary linguistic theories--analogy, induction, distribution and association are powerful sources of information and can, in concert with general cognitive innate knowledge, but without language-specific innate knowledge, learn important properties of language that are thought to be innate by many.