VARIOUS APPROACHES TO WEB INFORMATION PROCESSING
Kristína Machová, Peter Bednár, Marián Mach · 2007
The paper focuses on the field of automatic extraction of information from texts and text document categorisation including pre-processing of text docu- ments, which can be found on the Internet. In the frame of the presented work, we have devoted our attention to the following issues related to text categorisation: in- creasing the precision of categorisation algorithm results with the aid of a boosting method; searching a minimum number of decision trees, which enables the improve- ment of the categorisation; the influence of unlabeled data with predicted categories on categorisation precision; shortening click streams needed to access a given web document; and generation of key words related with a web document. The pa- per presents also results of experiments, which were carried out using the 20 News Groups and Reuters-21578 collections of documents and a collection of documents from an Internet portal of the Markiza broadcasting company.