Extracting collocations from text. An application: language generation
Frank Smadja · 1992
Natural languages are full of collocations, arbitrary and recurrent combinations of words that co-occur more often than chance. Such combinations correspond to arbitrary word usages and are termed collocations. Recent work in lexicography indicates that collocations are pervasive in English; apparently, they are common in all types of writing, including both technical and non-technical genres. In the dissertation, we describe a set of techniques for retrieving and identifying collocations from large textual corpora. These techniques are based on statistical methods, and identify a wide range of collocations. A statistical filtering technique is described for identifying word pairs involved in a syntactic relation. The words can appear in any order and can be separated by an arbitrary number of other words. Another technique describes how n word collocations (or n-grams) can be identified in a simpler and cheaper way than other methods. The techniques also feature an original method for syntactically labeling and filtering collocations. These techniques have been implemented in a lexicographic tool, Xtract that automatically acquires collocations. Xtract identifies collocations of arbitrary length as well as more flexible collocations. The techniques are described and some results are presented on a 10 million word corpus of stock market news reports. A lexicographic evaluation of Xtract as a collocation retrieval tool estimated the precision of Xtract to be 80%. The evaluation is presented in the dissertation. As a performance task, we demonstrate how such collocations enhance the task of lexical selection in language generation. Previous language generation works were not able to account for co-occurrence knowledge for two principal reasons. They did not have the compiled information and the lexicon formalisms available were not able to properly handle collocational knowledge. The knowledge problem is handled with the use of lexicographic tools such as Xtract, and the representation problem is handled with Functional Unification Grammars (FUGs). We show how the use of FUGs allows to properly handle the interactions of collocational and various other constraints. Finally, we consider several other applications of our work such as computer assisted lexicography, information retrieval, machine translation and spelling correction.