An Intelligent Archive Testbed Incorporating Data Mining Lessons and Observations
Hampapuram K. Ramapriyan, D. Isaac, Wenjing Yang, Brian Bonnlander, David J. Danks · 2006
The advances of the last two decades in remote sensing instruments, computational, storage and commu- nications hardware, and launches of a series of Earth ob- serving satellites by U.S. and international agencies, have created a data rich environment for scientific research and applications. NASA's Earth Observing System (EOS) Data and Information System (EOSDIS) has now been operational for over 11 years, and has been effectively cap- turing, processing, archiving and distributing a few tera- bytes of standard data products each day to a diverse and globally distributed user community and, along with other NASA sponsored data system activities, forming a value chain for users to obtain valuable data. Visions for the future include a highly distributed system of data and service providers and users' being able to lo- cate, fuse and utilize data with location transparency and high degree of interoperability, and being able to convert data to information and usable knowledge in an efficient, convenient manner, aided significantly by automation. We can look upon the distributed provider environment with capabilities to convert data to information and to knowl- edge as an Intelligent Archive in the Context of a Knowl- edge Building system (IA/KBS). There have been several research investigations into intelligent data understanding including data mining and knowledge discovery. However, these investigations typically perform proofs of concept on a relatively small scale. Before their contributions can be implemented on a large scale commensurate with today's Earth science data archives, it is necessary to test them in a pseudo-operational environment. The purpose of this pa- per is to describe a testbed that serves this purpose and discuss some of the observations and lessons learned from its implementation.