COMEX: Identifying Mislabeled Human Behavioral Context Data Using Visual Analytics
Hamid Mansoor, Walter Gerych, Luke Buquicchio, Kavin Chandrasekaran, Emmanuel Agu, Elke Angelika Rundensteiner · 2019
Context-Aware (CA) systems that adapt their behavior based on their users' current context have broad applications in areas from healthcare to smart environments. Context Recognition (CR) is currently solved using machine and deep learning approaches that require realistic datasets with ground truth labels collected "in-the-wild" as users live their lives. In such studies, users periodically label their current context (current activity, body position, and location) while a mobile app continuously gathers sensor data. Unfortunately, users sometimes assign wrong or incomplete context labels, reducing the quality of the labeled dataset; and causing lower classification accuracy. We present COMEX, an interactive visual analytics tool that assists analysts in identifying instances of mislabeled context data to improve the quality of CA datasets. For this, we first provide a conceptual categorization of mislabeling error types. Thereafter we develop linked visualizations, augmented by anomaly scores indicating suspected labeling issues, which provide richer insights into the diverse characteristics of the target dataset. We validate our approach on an open source dataset that contains context information for 60 participants gathered over several days using smartphones. With the help of COMEX, participants of our case study identified numerous mislabelled instances in the dataset. We re-ran the classification task after excluding mislabelled data and saw improvements in classification accuracy.