URL as starting point for www document categorisation
Vojtěch Svátek, Petr Berka · 2000
Information about the category (type) of a WWW page can be helpful for the user within search, filtering, as well as navigation tasks. We propose a multidimensional categorisation scheme, with bibliographic dimension as the primary one. We examine the possibilities and limits of performing such categorisation based on information extracted from URL, which is particularly useful for certain on-line applications such as meta-search or navigation support. In addition, we describe the problem of ambiguity of URL terms, and suggest a method for its partial overcoming by means of machine learning. As a side--effect, we show that general purpose WWW search engines can be used for providing input data for both human and computational analysis of the web. 1 Introduction The task of document categorisation is common within web applications, in particular for navigational, search and filtering systems, which give access to large amounts of documents. The aim of the categorisation may be . to e...