Active Learning with Pre-trained Language Models for Named Entity Recognition in Requirements Engineering
Michael Riesener, Maximilian Kühn, S. Schümmelfeder, D. Xiao, Johannes J. Norheim, Eric Rebentisch, Günther Schuh · Procedia CIRP · 2024
The ever-increasing number of textual requirements and the handling of these consumes significant time in product design. Therefore, automating requirements analysis with natural language processing is critical to increase efficiency. Pre-trained language models have proven useful for requirements analysis. However, to deploy pre-trained language models, a fully labeled dataset must first be created to finetune language models towards the requirements domain. The creation of a fully labeled dataset, especially for named entity recognition (NER) applications, poses a major challenge as it entails high manual labeling effort by requirements engineers. Furthermore, in an ever-changing requirements environment, the manual labeling process must be repeated to avoid model drift, adding additional workload. Recent advances of pre-trained language models combined with active learning can significantly reduce the manual labeling effort. Thus, we present a novel approach for leveraging active learning with pre-trained language models for NER. We demonstrate the efficacy of this combination in an experiment, where we are able to reduce the manual requirements labeling effort by 74% while even improving model performance. We argue that by reducing the manual effort, active learning can shorten the ramp-up time of language model deployment for requirements engineering and enable industrial adoption.