Data Mining for Space Applications
Karen Zita Haigh, Wendy Foslien, Valerie Guralnik · Space OPS 2004 Conference · 2004
algorithms are usually based on engineering knowledge of the problem and hand generated.They cannot identify problems outside prescoped knowledge.Data mining algorithms can sort through massive amounts of data to locate key features, identify key data segments, or directly train models for event detection.Data mining techniques can better use the vast archive of operational data to formulate models for early event detection.Data mining is essentially the process of automatically extracting valid, useful, previously unknown, and comprehensible information from large databases and using it to make crucial decisions.Data mining techniques are drawn from machine learning and artificial intelligence, pattern recognition, statistics, database systems, and data visualization.They are geared toward applications where traditional techniques may not be suitable, for example because of the enormity of the data, the high dimensionality of the data, or the heterogeneous, distributed nature of the data.Data mining techniques examine the collected data for unusual situations and the results can be used as a basis for an EED algorithm to detect unexpected problems.Usually, a data mining technique will learn a model of the normal behavior of the system, and then the run-time algorithm will detect situations outside normal.Models can also be built for known, abnormal conditions.Key challenges for data mining algorithms in this domain include:Sampling rate: different parameters are sampled at different rates.The sampling rate is not always consistent.Noise: sometimes sensors will report incorrect values, or communications streams will be garbled.Lag time: often, there is a delay between the anomaly occurrence and when it is reported, especially if the data is sent from space to a ground station.Missing values: sensors may go off-line.Volume: this data generally has high dimensionality, much of it not relevant.Coverage: it is unlikely that all modes of all parameters can or will be collected. Processing the BGA dataThe data mining process includes five tasks: data selection, cleaning to make the data consistent and remove outliers, reduction and transformation to reduce the effective number of variables under consideration, analysis to detect patterns and extract knowledge, and interpretation to verify the results and possibly use them in context.A common rule of thumb is that data preparation (selection, cleaning, and reduction) will be about 60% (or two-thirds) of the total effort expended. Data Selection and VisualizationSince analysts are key players in selecting the right data and the right tool for the job [3,8,12], they must understand the data.There are many techniques for visualizing data.One example of a visualization method is Honeywell's Visual Query Language (VQL) [7], which we used to visualize the BGA data.VQL is a technology for locating time series patterns in historical or real-time data.VQL integrates user specifications of visual patterns with an automated search of a historical database for those patterns.With VQL, users can define qualitative patterns in the data graphically or select previously defined templates and then express the selections as search directives.The VQL search engine uses trend-oriented algorithms to find similar patterns in the data stream and returns a ranked list of matches.