Entity Linking and Name Disambiguation in Chinese Micro-Blogs
Li Li, Yunlong Guo, Yu Xiang, Xiao Shuang Xu, WeiGang Zeng · 2014
The amounts of data on social networks have been increasing sharply with the development of Web 2.0. Extracting social media content for the construction and extension of the knowledge base, mainly through remove ambiguities of entities from microblogs, has attracted attention from both academia and industry. Understanding Chinese microblogs is challenging because of the inherent features of Chinese language, the in- formal usage of the language and the wide variety of content it covers. In this paper, we focus on entity disambiguation in Chinese microblogs. A Web crawler is first developed to collect relevant information from both Baidu Encyclopedia and Chinese Wikipedia. The creation of the entity dictionary is based on the exiting mapping rules obtained from Baidu search engine by the newly developed crawler. A novel disambiguation strategy including a clustering algorithm based on Newman fast algorithm is proposed along with a label disambiguation algorithm. We then evaluate our approach on the Chinese microblog data set. The experimental result achieved 89.34% in terms of accuracy, which is 4.35% better than the best result of 84.99% (of all participating teams). Our approach is promising in identifying entity links and discovering the potential links in Chinese microblogs.