Which Who are They? People Attribute Extraction and Disambiguationin Web Search Results
Man Lan, Yu Zhe Zhang, Yue Lu, Jian Su, Chew-Lim Tan · 2009
People name search often returns a lot of Web pages containing the strings of personal names. Due to namesake, extracting target person attributes (such as birthday, occupation, affiliation, nationality, contact information, etc.) is expected to be helpful to differentiate documents related to different people and thus group documents related to the same person. This paper presents the methodology for the two tasks of Web people disambiguation: target person Attribute Extraction (AE) and people Clustering. Specifically, in this paper we address three questions: (1) How to effectively extract target person attribute information from raw Web pages? (2) Is the information of extracted attributes able to lead to better performance than the information of raw Web pages for Web people clustering? (3) Which is important for Web people clustering, feature representation or clustering algorithms? To solve them, we first present an effective method to extract different types of target person attributes from raw Web pages by using deep Web page cleaning and processing pipelines with multiple techniques including traditional named entities recognition (NER), regular expression patterns, gazetteer-based matching and so on. Then we explore the methodology for Web people clustering from two aspects, i.e., feature representations (tokens from raw Web page, information of extracted attributes) and clustering strategy. The comparative experimental results showed that deep Web page cleaning contributes significantly to performance improvements for target person attribute extraction task. For people clustering task, the clustering algorithm contributes more to performance improvement than feature representations.