Any-language frame-semantic parsing

Anders Johannsen, Héctor Martínez Alonso, Anders Søgaard · 2015

We present a multilingual corpus of Wikipedia and Twitter texts annotated with FRAMENET 1.5 semantic frames in nine different languages, as well as a novel technique for weakly supervised cross-lingual frame-semantic parsing. Our approach only assumes the existence of linked, comparable source and target lan-guage corpora (e.g., Wikipedia) and a bilingual dictionary (e.g., Wiktionary or BABELNET). Our approach uses a truly interlingual representation, enabling us to use the same model across all nine lan-guages. We present average error reduc-tions over running a state-of-the-art parser on word-to-word translations of 46 % for target identification, 37 % for frame identi-fication, and 14 % for argument identifica-tion. 1

Read the paper · More papers on PaperTik