Multilingual and Cross-Lingual ComplexWord Identification

Seid Muhie Yimam, Sanja Štajner, Martin Johannes Riedl, Chris Biemann · 2017

Complex Word Identification (CWI) is an important task in lexical simplification and text accessibility.Due to the lack of CWI datasets, previous works largely depend on Simple English Wikipedia and edit histories for obtaining 'gold standard' annotations, which are of mixed quality, and limited to English only.We collect complex words/phrases (CP) for English, German and Spanish, annotated by both native and non-native speakers, and propose language independent features that can be used to train multilingual and crosslingual CWI models.We show that the performance of cross-lingual CWI systems (using a model trained on one language and applying it on the other languages) is comparable to the performance of monolingual CWI systems.

Read the paper · More papers on PaperTik