Disambiguation ofPeople inWebSearch Usinga Knowledge Base

Tomonari Masada, Atsuhiro Takasu · 2007

Results ofqueries bypersonal namesoften containinnewspapers, authors inpublications, people insocial net- documents related toseveral people because ofthenamesakework)havebeenresearched andseveral methods like vector problem. Inorder todifferentiate documents related todifferent space modelnameentity extraction graph modelhavebeen people, aneffective methodisneededtomeasuredocument'spac model , nam e e extron gramodelhave bee similarities andtofinddocuments related tothesameperson.proposed However, these methods aredifficult toapplyfor Someprevious researchers haveusedthevector space model documents inthewebaswebdocuments havedifferent char- orhavetried toextract commonnamedentities formeasuringacteristics fromdocuments having beenresearched. Compar- similarities. Wepropose anewmethodthatusesWebdirectories ingwithnewspapers' articles orpublications, webdocuments asaknowledge basetofindshared contexts indocument pairsareofmorevarious templates andcontain muchnoise. andusesthemeasurement ofsharedcontexts todetermine similarities between document pairs. Experimental results show Inordertosolvethenamedisambiguation problem in thatourproposed methodoutperforms thevector space model theweb,we propose a new methodthatcan effectively method andthenamedentity recognition method. measure similarities ofwebdocuments. Aswebdocuments often contain noisy data, tofind outatopic ofawebpageis

Read the paper · More papers on PaperTik