Modeling the Internet and the Web: Probabilistic Methods and Algorithms

Pierre Baldi, Paolo Frasconi, Padhraic Smyth · 2002

Having focused in earlier chapters on the general structure of the Web, in this chapter we will discuss in some detail techniques for analyzing the textual content of individual Web pages. The techniques presented here have been developed within the fields of information retrieval (IR) and machine learning and include indexing, scoring, and categorization of textual documents. The focus of IR is that of accessing as efficiently as possible and as accurately as possible a small subset of documents that is maximally related to some user interest. User interest can be expressed for example by a query specified by the user. Retrieval includes two separate subproblems: indexing the collection of documents in order to improve the computational efficiency of access, and ranking documents according to some importance criterion in order to improve accuracy. Categorization or classification of documents is another useful technique, somewhat related to information retrieval, that consists of assigning a document to one or more predefined categories. A classifier can be used, for example, to distinguish between relevant and irrelevant documents (where the relevance can be personalized for a particular user or group of users), or to help in the semiautomatic construction of large Webbased knowledge bases or hierarchical directories of topics like the Open Directory

Read the paper · More papers on PaperTik