How many different John Smiths, and who are they?
Anagba Kulkarni, Ted Pedersen · 2006
In this work we propose three unsupervised measures to automatically identify the number of distinct entities a given ambiguous name refers to in a corpus. We exper-iment with 22 artificially created name conflations and observe that the measure (PK2) formulated as the ra-tio of two successive clustering criterion function values outperforms the other two measures. We also describe a method to assign a unique label to each discovered clus-ter so as to identify the underlying entity that it refers to.