An infrastructure for curating, querying, and augmenting document data :

Eswaran Subrahmanian, Guillaume Sousa Amaral, Talapady N. Bhat, Kevin G. Brady, Mary C. Brady, Jacob Collard, Sarah Chouder, Philippe Dessauw, Alden A. Dima, John T. Elliott, Walid Keyrouz, Benjamin Long, Rachael Sexton, Nicolas J Lelouche, Ram Duvvuru Sriram · 2023

With the advent of the COVID-19 pandemic, there was the hope that data science approaches could help discover means for understanding, mitigating, and treating the disease. This manifested itself in the creation of the COVID-19 Open Research Dataset (CORD-19) which aggregated COVID-19- related scientific literature for use by the data mining community. As a group of interdisciplinary informatics researchers at NIST, we embarked on an effort to use our experience and previously developed systems to explore whether we could enhance the CORD-19 data set and facilitate its use. This effort produced a prototype scientific informatics system that extended data curation, data repository, resource registry, term extraction and indexing systems and resulted in a repackaging of CORD-19 as a Python data package. This paper documents our efforts, provides lessons learned, and proposes a general architecture for these types of systems.

Read the paper · More papers on PaperTik