Approaches to samples selection for machine learning based classification of textual data
František Dařena, Ján Žižka · RePEc: Research Papers in Economics · 2011
The paper focuses on retrieval of relevant documents written in a natural language based on availability of several candidate examples which are used as the basis for the automatic selection of only items that are similar to these predefined patterns. Presented approach should face problems related to processing user created content in natural language that include a poor control over the topic and the structure of the content and often also huge computational complexity. Three methods of selecting the best samples from a large set of candidate samples are presented - random selection, manual selection and a new approach called automatic biased sample selection, and measures based on Euclidean distance and cosine similarity are used for classification. The experiments are carried out with real world data consisting of customer reviews downloaded from amazon.com, converted to different representations based on bag-of-words procedure. The experiments and the results of the presented approach provided satisfactory values and can lead to an alternative approach to manual selection and evaluation of textual samples.