Mining Aviation Safety Data: A Hybrid Approach
Eric Bloedorn · 2000
Data mining is broadly defined as the search for interesting patterns from large amounts of data. Techniques for performing data mining come from a wide variety of disciplines including traditional statistics, machine learning, and information retrieval. While this means that for any given application there is probably some “data mining” technique for finding interesting patterns, it also means there exists a confusing array of possible data mining tools and approaches for any given application. This problem is exacerbated when the available data contains both structured as well as unstructured (free-text) data. For example, the aviation safety data used in the reported experiments contains records which include both free text event descriptions as well as structured fields for phase-of-flight and location. Performing separate analysis on these different sources of data does not fully exploit the available information (e.g. clustering records without regard to narratives can match reports of total electrical failure with human factors problems). Unfortunately currently available tools provide little support. This paper describes one approach to combining the information available from all of these different types of data together to get a single ‘similarity’ score. The importance of picking tools appropriate to the types of data in hand is also stressed.