Knowledge Discovery in Documents by Extracting Frequent Word Sequences

Helena Ahonen-Myka · Illinois Digital Environment for Access to Learning and Scholarship (University of Illinois at Urbana-Champaign) · 1999

INFORMATION needs caused by the increasing amount of available digital data, the notion of knowledge discovery has been developed.Knowledge discovery methods typically attempt to reveal general patterns and regularities in data instead of specific facts, the kind of information that is hardly possible for any human being to find.In this article, a method for extracting iiiuximtil ft-~quent sequences in a set of documents is presented.A maximal frequent sequence is a sequence ofwords that is frequent in the document collection and, moreover, that is not contained in any other longer frequent sequence.A sequence is considered to be frequent if it appears in at least n documents when n is the frequency threshold given.Frequent maximal sequences can be used, for instance, as content descriptors for documeiit ment is represented as a set of sequences, which can then be used to discover other regularities in the document collection.As the sequences are frequent, their combination of words is not accidental.Moreover, a sequence has exactly the same form in many documents, providing a possibility to do similarity mappings for information retrieval, hypertext linking, clustering, and discovery of frequent co-occurrences.A set of sequences, particularly the longer ones, as such may also give a concise sum- mary of the topic of the document.

Read the paper · More papers on PaperTik