Automatic compositionality detection from corpora

John Cristian Borges Gamboa · Lume (Universidade Federal do Rio Grande do Sul) · 2013

Phrasal verbs in English present varying levels of semantic idiosyncrasies. Aiming to detect some of these idiosyncrasies (in this case, how much of the meaning of a phrasal verb can be extracted from each of its words) a set of measures was proposed by MCC (2003), which use a thesaurus as input. This work reimplements those measures, focusing on checking how robust they are, by applying them on several thesauri. The thesauri were built using the method in Lin (1998). We evaluate our results using a gold standard, and the results suggest the PMI as the best way to filter the contexts the verbs are found in.

Read the paper · More papers on PaperTik