Investigating the Impact of Syntax-Enriched Transformers on Quantity Extraction in Scientific Texts
Necva Bölücü, Maciej Rybinski, Stephen Wan · 2023
Measurement extraction is an information extraction subtask focused on extracting quantities and their dependent entities within a given scientific text.Quantity extraction is the first and most important step in measurement extraction.Most existing approaches model the problem as a sequence-labeling task using pre-trained language models (PLMs).However, none of the existing systems have utilised explicit syntactic knowledge to extend the PLM-based modeling.We propose a syntax-enriched extension by integrating dependency tree representations as syntactic knowledge into transformer-based language models to address the task of quantity extraction.We apply our approach to a range of established transformer-based models to evaluate our approach and analyze its impact in experiments on scientific literature datasets.Our experimental results and in-depth analysis show that our approach, syntax-enriched RoBERTa, outperforms the other models, even in situations with scarce training data in the scientific domain.The results demonstrate the adaptability of the proposed model to the tasks, especially useful in low-resource scenarios.1