NU-MineBench: Understanding the Performance and Scalability Characteristics of Data Mining Algorithms
Jayaprakash Pisharath, Ying Liu, Wei‐keng Liao, Gokhan Memik, Alok Choudhary, Pradeep Kumar Dubey · 2004
Data mining has become one of the most essential tools for various businesses as well as researchers in diverse fields. The surge in the operational speed of computing systems, and also the emergence of compact, low-cost, high-performance parallel and distributed systems have provided abundant venues for improving the performance of data mining algorithms. However, in recent years, there has also been a tremendous increase in the size of data that is collected and also the complexity of data mining algorithms themselves. The rate of this growth exceeds the rate of performance improvements in computing systems, thus widening the performance gap between data mining systems and algorithms. In this paper, our goal is to narrow this gap by enabling designers to build systems that are tuned in accordance with the requirements and developments of data mining algorithms. We achieve this by performing a detailed characterization of a set of representative data mining programs from both the hardware and software perspectives. We first study several widely-used data mining algorithms from multiple categories and, then, use them to design NU-MineBench, a benchmarking suite containing representative data mining applications. MineBench suite includes two classification, two association rule mining, and four clustering applications. We evaluate the NU-MineBench applications on an 8-way shared memory