Measuring Effectiveness of Text-Decorated HTML Tags in Web Document Clustering.
Mark P. Sinka, David Corne · International Conference WWW/Internet · 2004
Web document analysis, and its associated research, underpins much of what is referred to as web intelligence and the envisaged 'semantic web'. A key issue in this field is how to encode a web document from the raft of potential document features without losing salient information. Current research almost always uses word-based feature vectors such as term frequency of specific words (TF) and/or variants such as normalised term frequency and TF*IDF. We explore the question of whether existing word-based term vectors can be usefully augmented by using text-decorated words delimited by the H1 HTML tag. We measure the effectiveness of a feature vector by encoding documents from a benchmark set in terms of this feature vector, and then measuring the accuracy of an unsupervised clustering task using this encoding. A thorough investigation is performed using a variety of parameter values, to explore whether any increase in accuracy is achieved over vectors constructed just from the plain document text. Tests on the BankSearch dataset showed 9 different parameter combinations (using the text-decorating tag words) that had an improved accuracy over the vectors obtained via the plain document text.