Fine-Tuning BERT-based Language Models for Duplicate Trouble Report Retrieval

Nathan S. Bosch, Serveh Shalmashi, Forough Yaghoubi, Henrik Holm, Fitsum Gaim, Amir H. Payberah · 2022 IEEE International Conference on Big Data (Big Data) · 2022

In large software-intensive organizations, trouble reports (TRs) are heavily involved in reporting, analyzing, and resolving faults. Due to the scale of modern organizations and products, multiple people independently often identify faults, leading to duplicate TRs. To mitigate the additional manual effort to identify and resolve these duplicate TRs, prior work at Ericsson focused on developing a 2-stage BERT-based retrieval system for identifying similar TRs when provided a new fault observation. This approach, although powerful, struggled to generalize to out-of-domain TRs. In this paper, we evaluate several fine-tuning strategies to integrate domain knowledge further, notably telecommunications knowledge, into the BERT-based TR retrieval models to (i) attain better performance on duplicate TR retrieval/identification and (ii) improve model generalizability to out-of-domain TR data. We find that adding domain-specific data into the fine-tuning models led to improved results on both overall model performance and model generalizability.

Read the paper · More papers on PaperTik