Mining Wikipedia Article Clusters for Geospatial Entities and Relationships.
Jeremy Witmer, Jugal Kumar Kalita · 2009
We present in this paper a method to extract geospatial en-tities and relationships from the unstructured text of the En-glish language Wikipedia. Using a novel approach that ap-plies SVMs trained from purely structural features of text strings, we extract candidate geospatial entities and relation-ships. Using a combination of further techniques, along with an external gazetteer, the candidate entities and relationships are disambiguated and the Wikipedia article pages are modi-fied to include the semantic information provided by the ex-traction process. We successfully extracted location entities with an F-measure of 81%, and location relations with an F-measure of 54%