An efficient association mining implementation on cluster of SMPs
Ruoming Jin, Gagan Agrawal · 2005
Introduction Association mining is one of the most important data mining tasks. Over the last 6 years, several sequential and parallel algorithms have been developed for association mining. An excellent survey is available from Zaki [30]. Parallel association mining algorithms have been developed for both distributed memory and shared memory architectures. Some recent efforts have also targeted hierarchical systems like cluster of SMP workstations, which have both distributed memory and shared memory parallelism [29]. We believe that two of the important challenges still remaining in developing parallel implementations of data mining tasks are: ffl Developing efficient implementations that process disk resident datasets. The amount of data available for analysis can be huge, and can easily exceed the aggregate main memory available on today's small and medium sized parallel machines. ffl Implementing parallel data mining algorithm without ver