A novel Data Extraction Framework Using Natural Language Processing (DEFNLP) techniques

Tayyaba Hussain, Muhammad Usman Akram, Anum Abdul Salam · Natural Language Processing Journal · 2025

Evidence through data is critical if government has to address threats faced by the nation, such as pandemics or climate change. Yet several facts about data necessary to inform evidence and science are locked inside publications. We used scientific literature dataset, Coleridge Initiative - Show US the Data, to discover how the data can be used for the public good. In this research, we demonstrate a general Data Extraction Framework Using Natural Language Processing (DEFNLP) Techniques which challenge data scientists to show how publicly funded data has been used to serve science and society. The proposed framework uses NLP libraries and techniques like SpaCy and NER respectively and different huggingface Question Answering (QA) models to predict the datasets used in publications. DEFNLP findings can assist the government in immediate decisions making, accountability, transparent public investments, economic and public health benefits. Until now such an issue having large dataset which belongs to numerous research areas has not been addressed. This approach is domain independent and therefore can be applied to all kind of case studies and scenarios which require data extraction. Our methodology sets the state-of-the-art on Coleridge Initiative dataset, reaching the highest score of 0.554 using salti bert QA model with the less runtime i.e. 417.4 and output of 819 bytes than other QA models e.g., Longformer (runtime: 2710.2, output: 1780 bytes) and BigBird (runtime: 839.4, output: 177020 bytes) with 0.444 and 0.387 score respectively which impressively raised the leaderboard score with an outcome of 0.711. Its computation time to answer each query on CPU is far less i.e. 0.0696s (than 0.3556s and 0.8967s) and has suitable hyperparameters for our dataset as maximum answer length is 64, greater batch size as well as learning rate. In terms of timing and performance, each epoch took around 5 min on average on a computer with output size of 3.27kB which is again far better than other frameworks.

Read the paper · More papers on PaperTik