Discovering Identities in Web Contexts with Unsupervised Clustering

Ted Pedersen · 2008

We describe the application of unsupervised clustering methodologies to the problem of discriminating among ambiguous names found in short passages of text that appear on Web pages. We show how to tailor these methods to handle the very noisy data that we typically find on the Web. We experiment with several variations in feature selection, two methods that automatically determine the number of clusters in the data, two different representations of the contexts to be discriminated, and with dimensionality reduction. Our evaluation is carried out using Web contexts for five different ambiguous names that were manually disambiguated to use as a gold standard. 1

Read the paper · More papers on PaperTik