Report on the CLEF-IP 2011 Experiments: Exploring Patent Summarization.
Parvaz Mahdabi, Linda Andersson, Allan Hanbury, Fábio Crestani · 2011
Abstract. This technical report presents the work carried out for the Prior Art Candidate Search track of CLEF-IP 2011. In this search sce-nario, information need is expressed as a patent document (query topic). We compare two methods for estimating query model from the patent document to support summary-based query modeling and description-based query modeling. The former approach utilizes a known text sum-marization technique, called “TextTiling”, and is adopted for patent doc-uments. The latter approach uses the description section of a patent document for estimating the query model. With summary-based query modeling we aspire to capture the main topic of the document as well as the most important subtopics and discard subtopics, which are only marginally discussed in the patent document. We submitted four runs for the Prior Art Candidate Search task. According to recall@1000 our best run was ranked 3rd across 6 participants and 8th, across all 30 sub-mitted runs. In terms of MAP our best run achieved the 3rd rank across participants and 4th rank, across all runs.