Sapphire: experiences in scientific data mining
Chandrika Kamath · Journal of Physics Conference Series · 2008
The size and the complexity of the data from scientific simulations, observations, and experiments are becoming a major impediment to their analysis. To enable scientists to address this problem of data overload and benefit from their improved data collecting abilities, the Sapphire project team has been involved in the research, development, and application of scientific data mining techniques for nearly a decade. In this paper, I first describe the Sapphire system architecture that was motivated by the needs of a diverse set of applications. Then, using examples from different domains, I discuss our experiences in mining science data and some of the challenges we faced in analyzing data ranging from a megabyte to several terabytes.