Hybridization of Active Learning and Data Programming for Labeling Large Industrial Datasets

Mona Nashaat, Aindrila Ghosh, James Miller, Shaikh Quader, Chad Marston, Jean‐François Puget · 2018

Modern machine learning (ML) models are being used heavily in business domains to build effective decision support systems. As a primary requirement, supervised ML models need large labeled datasets. However, obtaining a high volume of labeled training data is both expensive and time-consuming. Researchers have proposed several labeling approaches to avoid manual labeling efforts. Active learning (AL) and Data Programming (DP) are two state-of-the-art techniques used to label datasets. Nevertheless, both approaches have their strengths and weaknesses. For example, AL is computationally expensive to apply on large industrial datasets; and labels generated by DP are often inaccurate and difficult to interpret. To address these challenges, in this paper, we propose a novel hybrid method that integrates the scalability of DP with the user engagement and accuracy of AL. The proposed approach aims at optimizing the labeling process by applying DP to generate initial noisy training data and then use AL to query the user to label only those points that maximize the accuracy of the final labels with a minimum annotation cost. To evaluate the proposed approach, we have used five open source datasets and a real-world business dataset of 1.5 million records. We use traditional active learning and data programming techniques as baselines to compare the performance and annotation cost of our proposed approach. The results show that the proposed method can achieve higher labeling accuracy than data programming. It also can minimize the labeling cost in real-world business scenarios, while delivering a comparable level of performance (accuracy) with active learning.

Read the paper · More papers on PaperTik