Focused Crawling Using Name Disambiguation on Search Engine Results
Nicolas Martin, Khaled Khelif · 2011
In this paper, we report our approach allowing source selection in order to support Web data collection and tracking of events and biographical facts about a targeted person. The choice of the sources is crucial to enhance the quality of information extraction tools and it is considered as the first step in the collect and tracking task. We designed a source selection process to filter out ones that are not relevant for the targeted person - because they refer to an homonym. In this process, the name of the targeted person is submitted to the system and each result (title, snippet and url)is represented in the vector space model and then clustered, so that each cluster represents all the results about the same entity. The experimental results show that our approach can achieve interesting disambiguation performance only considering the search results.