Data Imbalance Problem in Text Classification

Yanling Li, Guoshe Sun, Yehang Zhu · 2010

Aimming at the ever-present problem of imbalanced data in text classification, the authors study on several forms of imbalanced data, such as text number, class size, subclass and class fold. Some useful conclusions are gotten from a series of correlative experiments: first, when the text of two class is almost the same number, the difference of word number become major factor to affect the accuracy of the classification, second, to improve the accuracy of the classification through increasing the small class size is limited, third, in the case of unbalanced data, the same words which are appeared in two class often carry strong class information, that is, class overlap will not affect the classification accuracy.

Read the paper · More papers on PaperTik