Using AWK to extract patterns from CBC radio news
Jacques Gélinas · 1997
More than 10000 news bulletins from the English Radio service of the Canadian Broadcasting Corporation were assembled to form a five megabyte textual document database. Random errors were introduced to simulate the errors resulting from scanning paper documents. The news items are represented by a vector of well chosen bigram, trigram and 4-gram frequencies after deleting the common words and stemming each remaining term. An experimental software using the text processing language AWK computes a similarity measure between a given query and each news bulletin. When plotted against time, these similarity measures reveal the rise and fall of the frequency of news items related to a given subject. The result is a visual representation of patterns for activities, thus extracting knowledge contained in the CBC Radio news bulletins. The robustness of the N-gram representation against typographical errors can also be clearly shown, and is surprising. Finally, possible modifications to the software are indicated. The software runs with GNU tools on Linux and under Windows95, and is available from the author.