Finding domain-specific termsfrom search engine's query logs
Weijian Ni, Tong Liu, Qingtian Zeng · 2015
Automatic domain-specific terms recognition is a basic step for various natural language processing applications. As for most traditional approaches, a domain-specific corpus of high quality and coverage needs to be available in advance. This paper aims to find domain-specific termsfrom a type of general corpus, i.e., search engine's query logs, which is of higher availability, coverage and timeliness than domain-specific corpora. In the proposed approach, the problem of automatic domain-specific query recognition is formulated as a supervised learning task, where feature representations of every query candidate are derived according to inherent structure of query logs. In addition, an under-sampling technique is employed to solve class-imbalance problem in the supervised learning task. By experimental evaluationon real query logs from a commercial search engine, the result demonstrates the advantages of the proposed approach.