Detecting Fine-Grained Cross-Lingual Semantic Divergences without Supervision by Learning to Rank

Eleftheria Briakou, Marine Jacinthe Carpuat · 2020

Detecting fine-grained differences in content conveyed in different languages matters for cross-lingual NLP and multilingual corpora analysis, but it is a challenging machine learning problem since annotation is expensive and hard to scale.This work improves the prediction and annotation of finegrained semantic divergences.We introduce a training strategy for multilingual BERT models by learning to rank synthetic divergent examples of varying granularity.We evaluate our models on the Rationalized English-French Semantic Divergences, a new dataset released with this work, consisting of English-French sentence-pairs annotated with semantic divergence classes and token-level rationales.Learning to rank helps detect finegrained sentence-level divergences more accurately than a strong sentence-level similarity model, while token-level predictions have the potential of further distinguishing between coarse and fine-grained divergences.ADV VERB ADJ NOUN how weak they are.BERT predictions { permission, attention, hand, mercy, story } WORDNET hypernyms { communication

Read the paper · More papers on PaperTik