Apply Paraphrase Generation for Finding and Ranking Similar News Headlines in Punjabi Language
Arwinder Singh, Gurpreet Singh Josan · Journal of scientific research · 2022
Paraphrase generation is an important task in Natural Language Processing (NLP) and is successfully applied in various applications such as question-answering, information retrieval & extraction, text summarization and augmentation of machine translation training data.A lot of research has been carried out on paraphrase generation but it is missing to apply it for finding similar news headlines.However, no approach is available for finding similar news headlines by using paraphrase generation in Punjabi Language.Hence, this paper presents an approach for finding similar news headlines by applying a paraphrase generation system to plug in the gap.Similar news is being detected by comparing the sentence vectors of news headlines.The input headline is compared with all available headlines using a combination of cosine and Jaccard similarity to find similar headlines.The detected news headlines are then ranked by comparing them with the given headline.To represent news headlines as sentence embeddings, RNN based Seq2Seq with different hyper-parameter settings and attention mechanisms is used.The direct method can restrict the process, so the proposed approach has applied a paraphrase generation system to rephrase the given headline for finding more relevant news headlines.To generate a paraphrase of the given headline, the current state-ofthe-art transformer with an augmented encoder is being used as transformers can learn long-term dependencies.The effect of Jaccard and cosine similarity has also been tested along with the combination of both metrics but the combination of Jaccard and cosine similarity performs better for finding similar headlines.The news headlines extracted from four newspapers have been used for evaluation in this article.Further, different evaluation metrics have been applied for in-depth comparisons, the BLEU score has been calculated between detected and input headlines.Another evaluation is done by comparing embeddings of the detected headlines and inputs along with human judgements.The generation of sentence vectors further enhanced the evaluation measure and showing that the proposed approach produced state-of-the-art results.