Extracting Domain Information using Deep Learning

Amit Kr. Gupta, Weijia Xu, Pankaj Jaiswal, Crispin B. Taylor, Jennifer D. Regala · Proceedings of the Practice and Experience in Advanced Research Computing on Rise of the Machines (learning) · 2019

Across various scientific domains, digital publication of technical documents, often in the form of conference/journal article submissions, are the first accessible instance of new human knowledge in these respective fields. Synthesizing and curating this information is a slow and difficult process and often requires non-trivial human expertise. Given the ever increasing rate of these publications and the natural limitations of manual approaches, a computational solution to this problem is the paramount need of the hour. One of the central tasks is to extract important phrases and terminologies from scientific article. Although many tools are available to extracting keywords and named entities from a document, a key challenge is to determine how important they are with regard the context of the entire document. In some cases, important entities to an article might be a new vocabulary that haven't, or rarely, appeared from previous data. In other cases, there are also many other existing entities that are less important for the particular article but weighted more significantly from models based on prior knowledge. In this paper, we investigate how deep learning methods may be used to address this issue.

Read the paper · More papers on PaperTik