Quality-Aware Entity-Level Semantic Representations for Short Texts.
Wen Hua, Kai Zheng, Xiaofang Zhou · Rare & Special e-Zone (The Hong Kong University of Science and Technology) · 2016
Recent prevalence of Web search engines, microblogging services as well as instant messaging tools give rise to a large amount of short texts including queries, tweets and instant messages. A better understanding of the semantics embedded in short texts is indispensable for various Web applications. We adopt the entity-level semantic representation which interpretes a short text as a sequence of mention-enity pairs. A typical strategy consists of two steps: entity extraction to locate entity mentions, and entity linking to identify their corresponding entities. However, it is never a trivial task to achieve high quality (i.e., complete and accurate) interpretations for short texts. First, short texts are noisy, containing massive abbreviations, nicknames and misspellings. As a result, traditional entity extraction methods cannot detect every potential entity mentions. Second, entities are ambiguous, calling for entity linking methods to determine the most appropriate entity within certain context. However, short texts are length-limited, making it infeasible to disambiguate entities based on context similarity or topical coherence in a single short text. Furthermore, the platforms where short texts are generated are usually personalized. Therefore, it is necessary to consider user interest and its dynamics overtime when linking entities in short texts. In this paper, we summarize our work on quality-aware semantic representations for short texts. We construct a comprehensive dictionary and extend traditional dictionary-based entity extraction method to improve recall of entity extraction. Meanwhile, we combine three novel features, namely content feature, social feature and temporal feature, to guarantee precision of entity linking. Empirical results on real-life datasets verify the effectiveness of our proposals.