Using Multiple Metrics in Automatically Building Turkish Paraphrase Corpus

Bahar Karaoğlan, Tarık Kışla, Senem Kumova Metın, Ufuk Hürriyetoğlu, Katira Soleymanzadeh · Research in Computing Science · 2016

Paraphrasing is expressing similar meanings with different words in different order.In this sense it is viewed as translation in the same language.It is an important issue in natural language processing for automatic machine translation, question answering, text summarization and language generation.Studies in paraphrasing can be classified as paraphrase extraction, paraphrase generation, paraphrase recognition.In this paper we present automatic sentential paraphrase extraction from comparable texts downloaded from Turkish newspapers related to similar news.We applied seven text similarity metrics and assumed the two most similar ones as candidates.Through an interface these are shown to 3 human annotators to be labelled as paraphrase, entailing, entailed, opposite in meaning and not paraphrase.In this paper we only present results driven from a single topic.The sentences in the other topics will be processed based on the experience gained in the current work.This will be the first automatically built and golden standard tagged Turkish paraphrase corpus.

Read the paper · More papers on PaperTik