Requirements extraction from engineering standards – systematic evaluation of extraction techniques

Janosch Luttmer, Vitalijs Prihodko, Dominik Ehring, Arun Nagarajah · Procedia CIRP · 2023

Working with information and knowledge is highly important in engineering design. Especially due to the continuous growth of information embedded in documents, the extraction process needs to be automated and accelerated. Hereby, engineering standards are an important source of knowledge in product development. However, the current way of working with these documents is inefficient and requires a lot of manual steps. This research paper focuses on the automatic extraction of requirements from standards documents and aims to systematically evaluate different extraction techniques in order to identify the best suitable technique for the given problem. For this, a dataset with approximately 10,000 entries is generated from existing standards. It is used to evaluate different extraction techniques whereas rule-based as well as supervised and unsupervised machine-learning techniques are implemented. It is shown that requirements are extracted efficiently, i.e. the F1-score is mostly greater than 80%. This proves the overall suitability for automatic requirements extraction from standards documents. The best results are achieved by Support Vector Machine (F1-scoreSVM=89.6%), followed by the fine-tuned BERT model (F1-scoreBERT=87.9%). Additionally, the rule-based extraction shows promising results and should be considered in future due to its high transparency. Future work will focus on the optimization and combination of different techniques as well as the automatic integration into requirements engineering tools.

Read the paper · More papers on PaperTik