Using Entropy to Learn OT Grammarsfrom Surface Forms Alone

Jason Riggle · 2006

The problem of ranking a set of constraints in Optimality Theory (Prince and Smolensky 1993) in a fashion that is consistent with an observed training sample comprised of 〈input, output〉 pairs has been solved with a variety of algorithms (e.g. Tesar 1995; Tesar and Smolensky 1993, 1996, 1998, 2002; Boersma 1997; Boersma and Hayes 2001). In real-world language learning scenarios, however, learners aren’t usually presented with 〈input, output〉 pairs but must instead learn grammars from output forms alone (possibly aided by conjectures about their meanings or morphology). The task of learning OT grammars from training samples that consist of output forms alone presents many challenges. Chief among these is the problem that there are often many 〈possible-input, possible-grammar〉 pairs that are consistent with any observed training sample of surface forms. For instance, the fully faithful ‘identity grammar’ is a perennially viable hypothesis under which the input forms are assumed to be basically identical to the observed output forms. In addition to this hypothesis, depending on the particular constraints in CON and the specific surface forms in the training sample, there are a potentially huge number of other 〈grammar, i/o-mapping-set〉 pairs, each one deriving a different set of unfaithful mappings from potential input forms to the outputs in the training sample. Knowledge of meaning and morphology can help the learner choose grammar hypotheses that map the same input to different surface instantiations of the same morpheme. But, even before the learner has any knowledge about the morphology and meanings of words, it is possible to make educated guesses about the structure of the grammar. Suggested strategies for this include principles like Prince and Tesar’s (1999) selectional preference for ranking hypotheses that are maximally ‘restrictive’ or Smolensky’s (1996) default MARKEDNESS >> FAITHFULNESS ranking. Both of these strategies restrict the search through the space of possible grammars by providing a heuristic that’s designed to prefer rankings that generate a tight fit with the observed data. In this paper I present yet another strategy for adjudicating among competing grammar hypotheses without recourse to morphological information or meanings. The strategy that I propose is not based on formal properties of the constraint rankings themselves, but instead is based on informationtheoretic properties of the set of inputs that each candidate grammar (ranking) associates with the training sample. The idea is this: if learners choose grammars whose associated input sets have the highest entropy (are least ordered) they will select grammars that maximally characterize patterns in the training sample as consequences of the grammar rather than as accidents of the lexicon.

Read the paper · More papers on PaperTik