Feature Selection and Negative Evidence in Automated Text Categorization
Luigi Galavotti, Fabrizio Sebastiani, Maria Simi · 2000
We tackle two dierent problems of text categorization, namely feature selection (FS) and classier induction. We propose a new FS technique, based on a simplied version of the 2 statistics and a novel variant, based on the exploitation of negative evidence, of the well-known k-NN method. We report the results of systematic experimentation of these two methods performed on the Reuters-21578 benchmark. 1. INTRODUCTION Text categorization denotes the activity of automatically building, by means of machine learning techniques, automatic text classiers (see e.g. [2]). Two key steps are document indexing and classier induction. Document indexing refers to the task of automatically constructing internal representations of the documents, able to synthetize the meaning of the documents. Usually, a text document is represented as a vector of weights d j = hw1j ; : : : ; wrj i, where r is the number of features (i.e. words) that occur at least once in at least one document of the coll...