Exploring classifier attribute interactions and time series using constrained randomisations

Andreas Henelius · Aaltodoc (Aalto University) · 2017

Gaining insight into structures and properties in data is a central problem in data mining and knowledge discovery. This is essential when the data is to be used, e.g., in decision-making. In this thesis we consider investigating the structure of data in two cases: temporal structures in time series, and attribute interactions utilised by classifiers. Time series are ubiquitous and represent an important type of data. We investigate temporal structures in time series, focusing on interval sequences. We seek explanations for observed properties by constructing and evaluating null hypotheses describing the internal properties of the time series. We approach this as a hypothesis testing problem where observed time series are compared to randomly generated instances. The properties being investigated are modelled in terms of constraints on the randomisations, allowing complex relationships to be examined and explained. Furthermore, we apply computational methods in the analysis of a sleep study to explain the relationship between time series representing heart rate variability and performance on a psychomotor vigilance test. Classification has wide applicability in multiple domains, however, many high-performing classifiers are essentially opaque, black-box algorithms, making it difficult to gain insight into the basis for predictions. In classifier analysis we consider attribute interactions utilised by classifiers. An interaction means that two or more attributes jointly carry information with respect to, e.g., a class label. We study two different types of interactions. Firstly, we investigate relationships between attributes in a dataset and show how this is related to factorising the class-conditional joint data distribution, such that attributes in the same factor are interacting while attributes in different factors are independent, given the class. We devise a method for testing the hypothesis that a dataset originates from a generating distribution with a particular factorised form. Secondly, we investigate how classifiers exploit attribute interactions in making predictions and develop a novel framework based on constrained randomisations for partitioning the attributes of a dataset into groups based on how they are jointly exploited by the classifier. The methods developed here are useful in several data analysis applications, e.g., in enhancing the interpretability of opaque classifiers, detecting adverse drug interactions in pharmacovigilance, anonymising data and gaining insight into the structure of datasets.

Read the paper · More papers on PaperTik