DELFI: Mislabelled Human Context Detection Using Multi-Feature Similarity Linking
Hamid Mansoor, Walter Gerych, Luke Buquicchio, Kavin Chandrasekaran, Emmanuel Agu, Elke Angelika Rundensteiner · 2019
Context Aware (CA) systems that adapt to user behaviors have many real-world uses. CA systems require accurately labeled training data to learn models of users' context behavior. Unfortunately, it is difficult to gather sufficient realistic context data in controlled environments where reliable labels can be gathered. Therefore, recent works have used in-the-wild study designs, where data is gathered through passive sensing devices such as smartphones while users periodically supply corresponding context labels. However, labels gathered this way can be unreliable as users may provide incomplete or inaccurate labels which makes it difficult to build robust CA models. We propose DELFI (Detecting Erroneous Labels using Feature-linking Insights), a visual analytics approach to discover and clean unlabeled or mislabeled context data. Visualizations enable highlighting of similar data to find patterns and anomalies in behaviors. However, this is challenging when working with erroneous human-labelled data as linking similar context labels is flawed since the labels themselves are in question. DELFI identifies probably-mislabeled instances by color-coding them based on an anomaly score. Additionally, DELFI links similar instances based on a novel concept called Multi-Feature Similarity Linking, which facilitates the identification of probably true labels of mislabeled and unlabeled data. We demonstrate the utility of our approach with use cases and evaluation from domain experts.