Scalable Phrase Mining for Ad-hoc Text Analytics
Srikanta Bedathur, Klaus Berberich, Jens Dittrich, Nikos Mamoulis, Gerhard Weikum, Srikanta Bedathur, Klaus Berberich, Jens Dittrich, Nikos Mamoulis, Gerhard Weikum · MPG.PuRe (Max Planck Society) · 2009
Large text corpora with news, customer mail and reports, or Web 2.0 contributions offer a great potential for enhancing business-intelligence applications.We propose a framework for performing text analytics on such data in a versatile, efficient, and scalable manner.While much of the prior literature has emphasized mining keywords or tags in blogs or social-tagging communities, we emphasize the analysis of interesting phrases.These include named entities, important quotations, market slogans, and other multi-word phrases that are prominent in a dynamically derived ad-hoc subset of the corpus, e.g., being frequent in the subset but relatively infrequent in the overall corpus.The ad-hoc subset may be derived by means of a keyword query against the corpus, or by focusing on a particular time period.We investigate alternative definitions of phrase interestingness, based on the probability of phrase occurrences.We develop preprocessing and indexing methods for phrases, paired with new search techniques for the top-k most interesting phrases on ad-hoc subsets of the corpus.Our framework is evaluated using a large-scale real-world corpus of New York Times news articles.