Proximity 2 -aware Ranking for Textual, Temporal, and Geographic Queries (extended version)
Michael Gertz · 2013
Abstract. Temporal and geographic information needs are frequent and important but not well served by standard IR systems. There are nei-ther good ways to add temporal or geographic constraints to a normal text query, nor are geographic and temporal expressions in the docu-ments interpreted as such kind of information, i.e., their semantics is not exploited. Recent approaches address such needs by extracting and normalizing temporal and geographic expressions from documents. They calculate specific scores for the temporal and/or geographic parts of a query. However, all approaches assume independence between the differ-ent query parts. In this paper, we present a new model to rank documents according to combined textual, temporal, and geographic queries. In this model, the independence assumption between the query parts is eliminated by cal-culating different proximity scores. Thus, documents are regarded to be more relevant if terms and expressions satisfying the different query parts occur close to each other in a document. In addition, we present a second type of proximity feature addressing the problem of sparse results. For this, we determine the temporal and geographic distance between expres-sions in a document and the queried time interval and geographic region. This allows to take into account documents containing expressions close to the time interval or region of interest. As our evaluations based on the NTCIR-GeoTime data show, our proposed model outperforms baseline models that do not use either of proximity information. 1