W&G-BERT: A Pretrained Language Model for Automotive Warranty and Goodwill Text
Lukas Jonathan Weber, Krishnan Jothi Ramalingam, Matthias Beyer, Alice Kirchheim, Axel P. Zimmermann · 2023
The demand for accurate information extraction architectures for industrial-based automotive warranty and goodwill (W&G) text data is steadily increasing. The analytical competence of current information extraction architectures is based on the Transformer architecture. Labeled datasets are necessary to finetune transformer-based language models like BERT on a domain-specific information extraction downstream task. We use customer related human-annotated W&G datasets to train BERT for Warranty and Goodwill Automotive Entity Recognition (W&G-BERT). As far as we know, we are the first authors publishing an automotive W&G-BERT model. Hence, we define the state-of-the-art in identifying the respective failure location and failure type in W&G feedback texts. Furthermore, the experimental results show that the pretraining of BERT on 2 million unlabeled W&G text data improves the performance by 1.5% (2.0%) f1-score (ensemble). We make the weights of the pretrained and finetuned models publicly available via the following link: lukasweber/WG_BERT · Hugging Face.