The Method of Text Categorization on Imbalanced Datasets
LI Xin-fu, Yu Yan, Yin Peng · 2009
In practical applications, datasets are usually imbalanced, but traditional approaches usually lead a low recognition rate. To address this problem, in this paper, over-sampling of the minority class has been proposed to increase the number of minority class, so as to achieve balance, thereby enhancing recognition rate of minority class. The experiments show that this approach achieved satisfactory results.