Parallelism speeds data mining
Sara Reese Hedberg · IEEE Parallel & Distributed Technology Systems & Applications · 1995
Increasingly, scientists and business people are mining mountains of stored data for valuable nuggets of insight using computerized data mining (also called knowledge discovery in databases, or KDD) techniques to learn more about customer behavior, discover new quasars, and catch crooks. Data mining recognizes patterns in data and predicts patterns from the data. Today's data-mining techniques, sometimes called 'siftware', range from OLAP (online application processing) tools that query multidimensional databases, to various statistical techniques, to advanced artificial intelligence techniques like machine learning, neural networks, rule-based systems and genetic algorithms. To efficiently mine gigabytes and terabytes of data in a timely fashion, a growing number of organizations are turning to parallelism to speed up the processing. The US Department of Treasury, for example, has fielded a data-mining application that sifts through all large cash transactions reported by banks and casinos to detect potential money laundering. The application runs on a 6-processor Sun server.>