Why Not Use Query Logs As Corpora

Olena Medelyan · 2004

Abstract. Generally, every Web search engine logs the user sessions. These records, called query logs, contain valuable information about the behaviour of Internet users and their language. There are only a few experiments on mining query logs, but they confirm that query logs are very useful for designing natural language applications in Web retrieval. This paper shows how lexical and semantic information can be extracted from query logs using statistical methods. I first summarize approaches in query log processing and mining for different purposes. After a short description of the used query logs, I present new domain- and language-independent methods for generating a compound dictionary and extracting semantically similar terms. The evaluation will shed light on the quality of proposed methods and show that the results are good enough to be directly integrated in query processing and improve information retrieval on the Web.

Read the paper · More papers on PaperTik