Dealing with Geographic Information in Location-Based Search Engines
Mr Saeid Asadi · The University of Queensland · 2008
With the massive increase of information available on the World Wide Web,search engines have emerged as the main media to search and retrieve this infor-mation. Location-based Web search has become a popular field of study as manylocal services have become available through the Web and new technologies useWeb-based geographic information. The potential market and growing interestin location-based Web search has encouraged major search engine companies todevelop location-aware search tools. However, many of these facilities are notpopular and their performance remains weak for many reasons including difficul-ties in obtaining geographic information from Web pages, lack of gazetteers andthe ambiguous nature of location names.This thesis presents the goals, tasks and outcomes of research undertakenas part of a PhD program which aimed to deal with geographic information inlocation-based search engines. The main focus of this research was the detection,extraction and proper use of geographic information on the Web. Several algo-rithms and models were suggested to improve the quality of location-based searchby properly handling location names and addresses. Before focusing on the mainresearch tasks, the related work on location-based Web search was reviewed anda comparative study was run to compare current location-aware search enginesfrom a user's point-of-view. Four major tasks were then addressed and completedin this research: Firstly, search engine queries were analyzed to understand thepatterns of location-based queries. It was found that most of these location-basedqueries did not directly mention a reference location even though their topic natu-rally refers to a particular location. Secondly, a heuristic pattern-based approachwas used to extract all blocks of addresses in Web pages and assign Web pageswith proper locations found in their content. The experiment showed that purecontent-based geo-tagging approaches are not able to cover all Web resources.As a result, for the task three, target location, or the location of the visitors ofWeb pages, was introduced as a new dimension of location which can be assignedliterally to all resources on the Web. In the fourth and final task, improvementof search results was studied by suggesting a new geo-ranking model which worksbased on local popularity of Web resources. The experiments showed more accurate geo-ranking when local popularity of Web resources is considered.This research addressed some of the problems of dealing with geographic information in location-based Web search. Proper handling of addresses andlocation names is an essential task for geographic search engines. Therefore, itis expected that the findings of this research facilitate designing and developingmore effective location-based search engines.