Dealing with highly imbalanced textual data gathered into similar classes

Jean-Charles Lamirel · 2013

This paper deals with a new feature selection and feature contrasting approach for classification of highly imbalanced textual data with a high degree of similarity between associated classes. An example of such classification context is illustrated by the task of classifying bibliographic references into a patent classification scheme. This task represents one of the domains of investigation of the QUAERO project, with the final goal of helping experts to evaluate upcoming patents through the use of related research.

Read the paper · More papers on PaperTik