Domain-independent scientific function finding
Cullen Schaffer, Casimir A. Kulikowski · 1990
Programs such as Bacon, Abacus, Coper, Kepler and others are designed to find functional relationships of scientific significance in quantitative data without relying on the deep domain knowledge scientists normally bring to bear in analytic work. Whether these systems actually perform as intended is an open question, however. To date, they have been supported only by anecdotal evidence--reports that a desirable answer has been found in one or more selected and often artificial cases. In this dissertation, I thus attempt to develop, not only new approaches to domain-independent scientific function finding, but, equally, a rigorous methodology under which research into such methods can be conducted. A fundamental problem with previous work is that it has investigated scientific data analysis in the abstract--without referring to actual scientific data. By contrast, the work reported here is founded on a collection of 352 real scientific data sets. This empirical base supports a number of strong conclusions. First, while researchers working with artificial data have targeted complex multivariate relations, real data provides powerful evidence that even the simplest bivariate relationships are difficult to identify reliably. Second, despite its ubiquitous presence in previous work, the notion of heuristic search of a potentially explosive space of formulas appears to help very little with the problem of reliably identifying basic bivariate relationships. Instead, third, substantial performance improvement results from viewing function finding as a decision problem, the problem of classifying data sets reliably within a fixed--and quite limited--system of functional categories. This dissertation presents what I believe to be the strongest domain-independent scientific function-finding algorithm currently in existence and, certainly, the only one which has been rigorously demonstrated. At the same time, it suggests fundamental limitations in the power of such algorithms. Through an extensive empirical case study and critical analysis of previous work, I arrive at a very different view of function finding than the one which has dominated work in this area for more than a decade. This new view narrows our focus of attention significantly while broadening the potential for productive research.