Machine learning for collocation identification

Shouxun Yang · 2004

Previous works on automatic identification or extraction of collocations from large scale corpora generally make use of certain statistical measures to test for significance of association to yield n-best collocation candidates for human scrutiny, optionally with linguistic preprocessing and linguistic filtering. The drawback of these approaches is we can only take advantage of one single statistical test (optionally in association with simple frequency threshold), even though we often calculate the values of several statistical tests. Manually exploring a scheme to combine two or more different tests is out of the question. We report experiments with machine learning for collocation identification using a variety of statistical association measurements. In particular, we develop a new decision tree learning algorithm based on C4.5 to be used for learning tasks with unbalanced data. The experiment results are presented and briefly discussed.

Read the paper · More papers on PaperTik