Tweet Contextualization: a Strategy Based on Document Retrieval Using Query Enrichment and Automatic Summarization.

Jorge Vivaldi, Iria da Cunha · CLEF (Working Notes) · 2013

The aim of the tweet contextualization INEX (Initiative for the Evaluation of XML retrieval) task at CLEF 2013 (Conference and Labs of the Evaluation Forum) is to build a system that provides automatically information related with different tweets, that is, a summary that explains a specific tweet. In this article, our strategy and results are presented. The methodology for the task in English includes three stages. First, automatic reformulations of the initial queries provided for the task, that is, the tweets, are performed. In this research, we use words sequences that agree with the typical terminological patterns, name entities, hashtags and Twitter users accounts, since we consider that they are representative of tweets’ topics. Second, related documents are retrieved from Wikipedia with the search engine Indri, using the reformulated queries. Third, the obtained documents are summarized by using two different automatic summarization systems, in order to provide the final summary associated to each query. Regarding the pilot task for Spanish, our strategy includes a first stage where automatic reformulations of the initial queries provided for the task (similar to English) are carried out. However, it does not include neither the search engine Indri nor the summarization systems REG and Cortex. In this case, we directly extract relevant text passages from Wikipedia pages using the generated queries and we build the summary with the first sentences of these pages.

Read the paper · More papers on PaperTik