Exploring Collections of research publications with Human Steerable AI
Alberto González Martínez, Billy Troy Wooton, Nurit Kirshenbaum, Dylan Kobayashi, Jason Leigh · Practice and Experience in Advanced Research Computing · 2020
Understanding highly-dimensional data sets is a complex task. Traditionally, this problem has been tackled with linear pipelines that rely on mathematical models and algorithms to summarize relationships and structure, producing a visual representation of the data in a collapsed, low-dimensional form. The main issue with these traditional pipelines is that they are driven solely by algorithms or models, and without a human in the loop, they can potentially limit sense-making by masking expected or known structure in the data. Textual data, such as that contained in research publications, is one example of unstructured highly dimensional data, wherein the raw data must be converted to an abstract numeric representation that is highly dimensional.