Characterizing Web Resources for Improved Search
Luis Gravano · 2000
Web resources are extremely diverse, not only along every conceivable topical and non-topical dimension, but also in terms of the access interface that they present to users. Current search engines ignore crucial non-topical dimensions of web resources that could be used to improve the quality of query results. As an important initial step to exploit such dimensions for web search, we have focused on geographical relevance. Web sites containing information on restaurants or apartment rentals, for instance, are relevant primarily to web users in geographical proximity to these locations. In contrast, an on-line newspaper may be relevant to users across the United States. We have studied how to mine the web and automatically estimate the geographical scope of web resources by using web hyperlinks and the actual content of web pages. For example, we can map every web page to a location based on where its hosting site resides. Then, we can consider the location of all the pages that point to, say, the Stanford Daily home page. By examining the distribution of these pointers, we can conclude that the Stanford Daily is of interest mainly to residents of the Stanford area, while The Wall Street Journal is of nation-wide interest. Similar conclusions can be drawn for other resources by analyzing the geographical locations that are mentioned in their pages. We have developed and evaluated algorithms for computing geographical scopes, described in [2], and implemented a geographically-aware search engine for on-line newspapers, accessible