Inferring Geographical Ontologies from Multiple Resources for Geographical Information Retrieval
Davide Buscaldi, Paolo Rosso, Piedachu Peris · 2006
Most of the information available in electronic format, such as in the World Wide Web or in digital libraries, involves some kind of spatial awareness. For instance, news usually describe an event and the place where this event occurred: “Earthquake in Turkey”, “Visit of the Pope in Valencia”. Currently, the Information Retrieval (IR) research commu-nity is increasing its efforts dedicated to the retrieval of ge-ographical information, as testified by the creation of the GeoCLEF 1 [5] evaluation exercise at the CLEF 2005, re-cently repeated in 2006, and the advances of the SPIRIT2 project [6]. These efforts are aimed to the solution of typical issues of the geographical IR task. In many cases, explicit geographical information is missing from the documents, for instance the indication of a broader geographical entity is omitted when it is supposed to be well-known to the read-ers (e.g. usually France is not named in a news related to Paris). Another common problem is the synonymy, when there are many ways to indicate a geographical entity. This is particularly true for foreign names, where spelling varia-tions are frequent. The solution to these problems has been generally individuated in the use of geographical-oriented ontologies [4, 6]. The manual construction of this kind of resources is usually a long, laborious process, and in many cases they are not freely available, such as the Getty The-saurus of Geographical Names3 (TGN). In order to over-come this issue, we made some attempts [2, 3] to use the geographical information included in WordNet, the well-known general domain ontology developed at the University of Princeton [7]. Unfortunately, the quantity of geographical information included in WordNet is quite small. Although it is quite difficult to calculate the number of geographi-