Analysing users www search behaviour
Clare Bradford, Ian Marshall · Lancaster EPrints (Lancaster University) · 1999
In a recent study [1], Internet users ranked search as their most important activity, awarding it a 9.1 on a 10-point scale.The next most important activity ranked only 6.3.Internet search engines are continually updating their indexes, and scaling up their parallel processors to keep up with the growth of the WWW. It is estimated that there are 800 million indexable pages in the WWW [2], and the number is growing at a rate of a million pages per day [3].Existing search engines cover around 15% of these pages, with the six largest public search engines collectively covering only 60%.The coverage is decreasing rapidly.User experience also shows that the frequency of out of date and broken links is increasing.In addition the traffic generated by indexing spiders and metasearch agents is adding significantly to congestion and delay.Perhaps most importantly the users valuation of search engines is declining, since the engines supply too much information.Previous work [4] has indicated that a cache hierarchy could provide an improved data-set for WWW search engines.A cache hierarchy would cover around 50%[5] of the publicly indexable pages, without generating spider traffic, and avoiding the need for the multiple requests generated by metasearch tools.This proposal therefore solves some of the operational problems, but does not address the users need for more effective filtering.Currently a typical simple query will generate in excess of 105 responses, and users can take many attempts to make their query more specific before successfully reducing the response to a more manageable level.Experience at AltaVista [6] has shown that few users request pages after the first results are listed.They either refine their search or leave the search site.We are attempting to address this problem on two fronts.Firstly we have suggested [7] that cataloguing the contents of the cache will enable users to choose the directory most closely matching their topics before firing their search.The categories are chosen by monitoring the users activity, and adapting to improve the response.An alternative version of the same type of approach [8] allows users to create their own catalogues and federate them.Secondly since the search engines are hindered by their inability to scale [9], we are attempting to reduce the scale by removing duplication.