Concept-based knowledge discovery in texts extracted from the Web

Stanley Loh, Leandro Krug Wives, José Palazzo Moreira de Oliveira · ACM SIGKDD Explorations Newsletter · 2000

This paper presents an approach for knowledge discovery in texts extracted from the Web.Instead of analyzing words or attribute values, the approach is based on concepts, which are extracted from texts to be used as characteristics in the mining process.Statistical techniques are applied on concepts in order to find interesting patterns in concept distributions or associations.In this way, users can perform discovery in a high level, since concepts describe real world events, objects, thoughts, etc.For identifying concepts in texts, a categorization algorithm is used associated to a previous classification task for concept definitions.Two experiments are presented: one for political analysis and other for competitive intelligence.At the end, the approach is discussed, examining its problems and advantages in the Web context.

Read the paper · More papers on PaperTik