A coupled linguistics/statistical technique for query structure classification and its application to Query Expansion

Bhawani Selvaretnam, Mohammed Belkhatir, Chris H. Messom · 2013

The retrieval effectiveness of Query Expansion (QE) is very much dependent on the ability to accurately identify and expand core concepts which are truly representative of the intended search goal. Two characteristics of natural language queries which hinder the performance of query expansion for information retrieval are query length and structure. The varying lengths of a query translate to the number of core concepts that may exist and the possibility of there being multiple query intents embedded within a single query. On the other hand, the structure of queries reveals the linguistic properties which allows for the determination of whether they take the form of well-formed sentences or are simply bags-of-words which in the strictest sense are a series of words with no obvious relations amongst them. Whilst query lengths are easily assessed, we propose a two-level automated classification technique consisting of linguistics based and statistical processing for query structure classification. The proposed method has revealed high levels of classification accuracy on TREC ad hoc test queries.

Read the paper · More papers on PaperTik