NP Alignment in Bilingual Corpora
Gábor Recski, András Rung, Attlia Zséder, András Kornai · 2010
We created a simple gold standard for English-Hungarian NP-level alignment, Orwell's 1984, (since this already exists in manually verified POS-tagged format in many languages thanks to the Multex and MultexEast project) by manually verifying the automaticaly generated NP chunking (we used the yamcha, mallet and hunchunk taggers) and manually aligning the maximal NPs and PPs.The maximum NP chunking problem is much harder than base NP chunking, with F-measure in the .7 range (as opposed to over .94for base NPs).Since the results are highly impacted by the quality of the NP chunking, we tested our alignment algorithms both with real world (machine obtained) chunkings, where results are in the .35range for the baseline algorithm which propagates GIZA++ word alignments to the NP level, and on idealized (manually obtained) chunkings, where the baseline reaches .4 and our current system reaches .64.