Retrieving Domain‐Specific Collocations by Co‐occurrences and Word Order Constraints

Sayori Shimohata, Toshiyuki Sugio, Junji Nagata · Computational Intelligence · 1999

In this paper, we describe a method for automatically retrieving collocations from large text corpora. This method comprises the following stages: (1) extracting strings of characters as units of collocations, and (2) extracting recurrent combinations of strings as collocations. Through this method, various types of domain‐specific collocations can be retrieved simultaneously. This method is practical because it uses plain text with no specific‐language‐dependent information, such as lexical knowledge and parts of speech. Experimental results using English and Japanese text corpora show that the method is equally applicable to both languages.

Read the paper · More papers on PaperTik