Breaking Bad: Extraction of Verb-Particle Constructions from a Parallel Subtitles Corpus
Aaron Smith · 2014
The automatic extraction of verb-particle constructions (VPCs) is of particular inter-est to the NLP community. Previous stud-ies have shown that word alignment meth-ods can be used with parallel corpora to successfully extract a range of multi-word expressions (MWEs). In this paper the technique is applied to a new type of cor-pus, made up of a collection of subtitles of movies and television series, which is par-allel in English and Spanish. Building on previous research, it is shown that a preci-sion level of 94 ± 4.7 % can be achieved in English VPC extraction. This high level of precision is achieved despite the dif-ficulties of aligning and tagging subtitles data. Moreover, many of the extracted VPCs are not present in online lexical re-sources, highlighting the benefits of using this unique corpus type, which contains a large number of slang and other informal expressions. An added benefit of using the word alignment process is that trans-lations are also automatically extracted for each VPC. A precision rate of 75±8.5 % is found for the translations of English VPCs into Spanish. This study thus shows that VPCs are a particularly good subset of the MWE spectrum to attack using word alignment methods, and that subtitles data provide a range of interesting expressions that do not exist in other corpus types. 1