Triple based Background Knowledge Ranking for Document Enrichment

Muyu Zhang, Bing Qin, Ting Liu, Mao Yu Zheng · 2014

Document enrichment is the task of retrieving additional knowledge from external resource over what is available through source document. This task is essential because of the phenomenon that text is generally replete with gaps and ellipses since authors assume a certain amount of background knowledge. The recovery of these gaps is intuitively useful for better understanding of document. Conventional document enrichment techniques usually rely on Wikipedia which has great coverage but less accuracy, or Ontology which has great accuracy but less cover-age. In this study, we propose a document enrichment framework which automatically extracts “argument1, predicate, argument2 ” triple from any text corpus as background knowledge, so that to ensure the compatibility with any resource (e.g. news text, ontology, and on-line ency-clopedia) and improve the enriching accuracy. We first incorporate source document and back-ground knowledge together into a triple based document-level graph and then propose a global iterative ranking model to propagate relevance score and select the most relevant knowledge triple. We evaluate our model as a ranking problem and compute the MAP and P&N score to validate the ranking result. Our final result, a MAP score of 0.676 and P&20 score of 0.417 outperform a strong baseline based on search engine by 0.182 in MAP and 0.04 in P&20. 1

Read the paper · More papers on PaperTik