Frequent Frames as Cues to Part-of-Speech in Dutch: Why Filler Frequency Matters - eScholarship

Richard E. Leibbrandt, David M W Powers · Proceedings of the Annual Meeting of the Cognitive Science Society · 2010

Frequent Frames as Cues to Part-of-Speech in Dutch: Why Filler Frequency Matters Richard Eduard Leibbrandt ([email protected]) David Martin Ward Powers ([email protected]) School of Computer Science, Engineering and Mathematics, Flinders University, Adelaide, Australia Abstract The Frequent Frames model (Mintz, 2003) attempts to assign words to word categories based on their distributional patterns of usage. This model is highly successful in categorizing words in child-directed speech in English, but has been shown by Erkelens (2008) to be less effective with Dutch material. We show that extending the amount of contextual information in a frame by making use of the full utterance context does not improve categorization performance, but that constraining the fillers of Frequent Frames to be relatively less frequently occurring words does improve categorization significantly. We connect the latter result to a basic dichotomy in some languages between function words and content words, and conclude that, at least for English and Dutch, paying attention to this dichotomy is of greater importance for distributional bootstrapping proposals than the specific distributional contexts that are used to categorize words. Keywords: Language learning; Distributional bootstrapping; Parts-of-speech; Function words; Frequent frames Introduction The parts-of-speech of a language (word classes such as nouns, verbs and adjectives) are of crucial importance in describing the grammar of the language. A vast amount of research has aimed to delineate the processes by which children learn to categorize words into the parts-of-speech of their native language. Researchers favouring semantic bootstrapping approaches (Grimshaw, 1981; Pinker, 1984) have proposed that early word categories are formed by grouping together words that refer to the same dimensions of concrete meaning, such as actions or objects. On the other hand, following early proposals by Maratsos & Chalkley (1980), proponents of distributional bootstrapping have argued that word categories can be induced by observing that certain groups of words are used in similar linguistic contexts, whether these contexts are defined at the level of words, morphemes, or even phonological or prosodic phenomena. In recent years, it has become feasible to implement specific distributional bootstrapping proposals as computer algorithms that attempt to categorize words purely by analysing distributional patterns in large corpora of natural utterances (Cartwright & Brent, 1997; Redington, Chater & Finch, 1998). For instance, Redington et al. (1998) found that words in child-directed English speech could be categorized with a high level of success by considering only very local utterance contexts made up of words that occur in close proximity to the target word. A particularly successful distributional model has been the Frequent Frames model of Mintz (2003, 2006a, 2006b). Frequent frames are defined as a disjunct frame occurring around a target word, made up of the word immediately preceding and the word immediately following the target, so that all frequent frames have the form a _ b, with a and b standing for specific words, and the underscore representing a slot that can accept a variety of filler words. For example, in the three-word sequence “a house and”, the frame is “a _ and”, and the filler is “house”. Once all frames of this form have been collected from a corpus, only the most frequent ones are retained for the purpose of categorization. This reflects the intuition that, if two words co-occur frequently on either side of another word across several utterances, this is likely to be due to some meaningful linguistic relationship between them. All words occurring in the same frequent frame are assigned to the same category, and frames that have more than 20% overlap in their set of slot fillers have their categories amalgamated into larger, more general categories. This amalgamation step is crucially important: by grouping together frames that accept similar sets of words, the child may be able to hypothesize that a word used in one verb frame may also be legitimately used in another verb frame; without amalgamation, this kind of generalization is not possible. Frequent Frames provide a very successful categorization of the words that occur in them, with Mintz (2003) reporting values greater than 0.9 for the evaluation measures accuracy and completeness when the model was implemented on a set of English corpora. However, recently Erkelens (2008) has shown that, in the case of child-directed speech in Dutch, Frequent Frames provide a less accurate basis for part-of- speech categorization than they do for English: whereas the use of Frequent Frames in English yielded an accuracy figure that exceeded the random baseline by 0.52 for tokens and 0.46 for types, a replication with a Dutch corpus could attain an improvement in accuracy over baseline of only 0.33 for tokens and 0.25 for types. Full-Utterance Frames As Distributional Contexts An important issue in distributional bootstrapping is to decide on the most appropriate usage contexts to consider for the purpose of categorization. One possible reason for the purported lower utility of Frequent Frames in Dutch

Read the paper · More papers on PaperTik