DocTable: Table-Oriented Interactive Machine Learning for Text Corpora
Sriram Yarlagadda, David J. Scroggins, Fang Cao, Yeshwanth Devabhaktuni, Franklin Buitron, Eli T. Brown · 2021
People working today with text data in any domain must develop understanding based on more documents than they can directly read. Tools exist that take advantage of visual analytics for a variety of text data tasks, including for technical experts, particularly in certain domains like intelligence and law. There is a missed opportunity to apply human-in-the-loop (HIL) machine learning to assist a general audience with text analysis tasks over large corpora that are difficult with visualization alone. In this paper, we propose an alternate approach to efficient sensemaking over document corpora, designed for users without technical expertise or training. We use a table-based interface, where the primary means of providing feedback to the machine is interactions with a data table view. This format is familiar to many, and augments the ability to tag documents into user-defined categories with machine learning that automatically predicts categories for additional documents. Marking only a few documents enables the machine learner to suggest automatic labels for the rest (with uncertainty scores for the predictions), and to reorder the table to reflect an integrated mix of the category models. Finally, interactive, force-directed layouts of topics and documents based on the models assist in sensemaking for workflow that scales beyond the practical limits of a table. To validate the technique, we present a prototype human-in-the-loop machine learning system, DocTable. We evaluate this technique with machine learning experiments that demonstrate the quality of the backend’s response to expected user inputs, a usage scenario to demonstrate its use on real world data, and a case study of expert feedback from our journalist collaborators.