Person Name Disambiguation on the Web by Two-Stage Clustering

Masaki Ikeda, Shingo Ono, Issei Sato, Minoru W. Yoshida, Hiroshi Nakagawa · 2009

The more important web searching becomes, the more we have to focus on the “same name ” problem in web searches. In this paper, we report our algorithm for disambiguating person names in web search results. It is a document clustering algorithm based on hierarchical agglomerative clustering using named entities, compound keywords, and URLs as features for document similarity calculation. We propose a two-stage clustering algorithm to improve the low recall values, in which the clustering results of the first stage are used to extract features used in the second stage clustering. We participated in the WePS-2 evaluation with this algorithm. We explain the results and describe other experiments performed with the WePS-1 data sets.

Read the paper · More papers on PaperTik