Text categorization study case: Patents' application documents

Neide de Oliveira Gomes, Emmanuel Passos · 2011

This paper presents computational methods aiming to patent's text categorization in Portuguese language, involving techniques from machine learning and computational linguistics. The algorithm used was the k-Nearest Neighbor method (k-NN) modified which showed good results, although it requires much computational time in the training stage. For the pre-processing step, it was implemented, with modifications, the stemming method called StemmerPortuguese that includes the removal of suffixes, besides the removal of stopwords and treatment of compound terms.

Read the paper · More papers on PaperTik