Mining Molecular Datasets on Symmetric Multiprocessor Systems
Thorsten Meinl, Marc Wörlein, Ingrid Fischer, Michæl Philippsen · 2006
Although in the last few years about a dozen sophisticated algorithms for mining frequent fragments in molecular databases have been proposed, searching big databases with 100,000 compounds and more is still a time-consuming process. Even the currently fastest algorithms like gSpan, FFSM, Gaston, or MoFa require hours to complete their tasks. This paper presents thread-based parallel versions of MoFa [5] and gSpan [26] that achieve speedups up to 11 on a shared-memory SMP system using 12 processors. We discuss the design space of the parallelization, the results, and the obstacles that are caused by the irregular search space and by the current state of Java technology.