Building a Non-Trivial Paraphrase Corpus Using Multiple Machine Translation Systems

Yui Suzuki, Tomoyuki Kajiwara, Mamoru Komachi · 2017

We propose a novel sentential paraphrase acquisition method.To build a wellbalanced corpus for Paraphrase Identification, we especially focus on acquiring both non-trivial positive and negative instances.We use multiple machine translation systems to generate positive candidates and a monolingual corpus to extract negative candidates.To collect nontrivial instances, the candidates are uniformly sampled by word overlap rate.Finally, annotators judge whether the candidates are either positive or negative.Using this method, we built and released the first evaluation corpus for Japanese paraphrase identification, which comprises 655 sentence pairs.

Read the paper · More papers on PaperTik