Learning to reason about data: The spring school on intelligent data analysis
Rosaria Silipo, Giuseppe Di Fatta · Intelligent Data Analysis · 2001
With the advent of faster and cheaper storage units, over the last decade it has become more and more common practice to collect large amounts of data. Market research firms, environmental monitoring stations, control systems for production chains, health care providers, bioinformatic companies, web page archives and other similar systems may easily record millions of observations/variables per day. This generates exponentially growing data sets [1, Chapter 1]. Because of their size many of such large data sets will likely contain precious information about the originating process in terms of its evolution and hidden patterns. Each feature of the system indeed may be more or less important to explore depending on the goal of the analysis and more or less easy to describe depending on the data. Appropriate data analysis tools are then required for each problem and for each set of data. Many data analysis techniques are available nowadays, each one with a different flavor. For example, Statistics, because of it’s branching from mathematics, is usually considered the most rigorous; while Neural Networks have gained the field as a “black-box” method. Some of the data analysis techniques were developed in physics especially for time series analysis; others simulate the evolutionary behavior of populations; some others concentrate on non-crisp decision processes. But all these strategies can be very effective and, in principle at least, there is a great potential for synergy. For each given problem, how do we choose the appropriate analysis method? And if different methods yield different results, may we use them in cooperation as to show different aspects of the same system? That is, are they complementary? Can we help the analysis with some of the background knowledge? In addition many of the previously cited methods might not be designed for large data sets. Is it then possible to reduce the data set dimensionality without loosing informative contents? Is there any pre-processing technique that can enhance the aspects of the system we want to investigate? How do we deal with sets of heterogeneous data? Those are some of the questions researchers deal with when starting a new investigation. Intelligent Data Analysis (IDA) represents an attempt to investigate and integrate different techniques, including Artificial Intelligence and Statistics, to enable an efficient analysis of large amounts of data.