Idact transformation manager
Michael J. Oudshoorn, Brian Hay · 2006
As scientific models and analysis tools become increasingly complex, they allow researchers to manipulate larger, richer, and more finely-grained datasets, often gathered from diverse sources. These complex models provide scientists with the opportunity to investigate phenomena in much greater depth, but this additional power is not without cost. Often this cost is expressed in the time required on the part of the researcher to find, gather, and transform the data necessary to satisfy the appetites of their data-hungry computation tools, and this is time that could clearly be better spent analyzing results. In many cases, even determining if a data source holds any relevant data requires a time-consuming manual search, followed by a conversion process in order to view the data. It is also commonly the case that as researchers spend time searching for and reformatting data, the expensive data processing hardware remains underutilized. Clearly this is not an efficient use of either research time or hardware. This research effort addresses this problem by providing a Transformation Manager which leverages the knowledge of a group of users, ensuring that once a method for a data transformation has been defined, either through automated processes or by any single member of the group, it can be automatically applied to similar future datasets by all members of the group. In addition to the use of existing transformation tools for well known data formats, methods by which new transformation can be automatically generated for previously unencountered data formats are developed and discussed. A major deliverable for this work is a prototype Transformation Manger, which implements these methods to allow data consumers to focus on their primary research mission, as opposed to the mechanics of dataset identification, acquisition, and transformation.