Automatic extraction of collocations: a new Web-based method
Jean-Pierre Colson · 2010
The automatic extraction of collocations from large corpora or the Internet poses a daunting challenge to computational linguistics. Indeed, previous statistical methods based on bigram extraction have shown their limitations, and there is besides no theoretical consensus on the extension of parametric methods to trigrams or higher n-grams. This is a key issue, because the automatic extraction of significant n-grams has important implications for computer-aided translation, translation quality assessment, automated text correction, terminology and computational lexicography. This paper reports promising results that were obtained by using a totally different approach to the automatic extraction of significant n-grams of any size. Instead of having recourse to statistical scores, the method is based on the testing of proximity algorithms that corroborate the native speaker’s competence about existing collocations. It is argued that compound terminology and phraseology in the broad sense can be captured by algorithms based on linguistic co-occurrence phenomena. This is made possible by a subtle manipulation of the Application Programming Interface (API) of a Web search engine, in this case Yahoo. The algorithm presented here, the Web Proximity Measure (WPR), has been tested on about 4,000 collocations mentioned in traditional dictionaries and on 340,000 n-grams extracted from the Web 1T or ‘Google n-grams’. The results show precision and recall scores superior to 0.9.