Classifying Web corpora into domain and genre using automatic feature identification

Serge Sharoff · 2007

Texts in representative corpora are typically classified into their domain and genre. However, it is not clear if existing domain and genre typologies can be applied at all to unlabeled data collected from the Web, for instance, to results of crawling. This study attempts to establish the most suitable categories for describing domains and genres of arbitrary web texts and to estimate the accuracy of their automatic classification using machine learning methods, such as Support Vector Machine (SVM) and clustering (repeated bisections and graph clustering). We also discuss methods for inducing the most discriminative features to perform this classification. The method has been designed to work with few or no linguistic resources and has been validated on a variety of languages: English, German, Chinese and Russian.

Read the paper · More papers on PaperTik