Time Dependent Approach for Query and URL Recommendations U sing Search Engine Query Logs
R. Umagandhi, A. V. Senthil Kumar · 2013
Search engine retrieves the significant and essential information from the web, based on the query keyword given by the user. Query log file is a repository contains every query request and its navigation in the search engine and maintained either in the system desktop or in the proxy server. This paper has proposed an algorithm for query and URL recommendation which is based on the user's search histories and a click through data. A new framework is constructed here based on the navigation time. First, the proposed algorithm identifies frequently accessed Queries and URLs from the log file using the frequent pattern generation algorithms. Next hub and authority weights are calculated for the frequent items. The similar queries are clustered; it uses the temporal characteristics of historical click-through data. The intuition is to reveal that more accurate semantic similarity of queries can be obtained by considering the timestamps of the log data. The cluster generated in this approach is used to provide query and URL recommendations to the user. Finally the method has been evaluated using real data set from the search engine query log. The tremendous growth of World Wide Web has paved the way for getting the required information from the web. Search engines are used to retrieve the result from the web in terms of web snippets for the query given by the user. The retrieved result may not be relevant all the time. At times irrelevant and redundant results are also retrieved by the search engine because of the short and ambiguous query keywords (1). The user scans the search result from the top to the bottom according to Joachim's (2) and then decides whether the web snippet is either relevant or irrelevant. A study done by C. Silverstein (3) on Alta Vista Query Log has shown that more than 85% of the queries contain less than three terms and the average length of the query is 2.35 terms. So the shorter length query does not provide any meaningful, relevant and needed information to the users. Table I shows some of the examples for ambiguous query keywords. (4) Reported that up to 23.6% of web search queries are ambiguous, this causes poor retrieval results.