Error Detection of BERT-Based Bibliographic Information Extraction from Reference Strings
S. Nakayama, Teruhito Kanazawa, Fumito Uwano, Manabu Ohta · 2025
Extracting bibliographic information, such as titles and authors, from reference strings is indispensable for generating inter-document links in digital libraries. Thus, we have developed a BERT-based bibliographic information extraction method for reference strings. Although the extraction accuracy was competitive, labor-intensive manual error correction is in-evitable after the extraction. Therefore, we propose a method to detect reference strings that are highly likely to contain extraction errors in order to reduce the manual post-editing. We also use multiple natural language processing models other than BERT to extract bibliographic information from reference strings and detect extraction errors. In our experiments, we present the bibliographic information extraction accuracy of multiple models as well as the accuracy of detecting erroneous reference strings. Specifically, in experiments involving reference strings from an English journal, XLM-RoBERTa achieved a bibliographic information extraction accuracy of 0.958 and detected 89.0% of reference strings containing extraction errors. These detected reference strings represented 13.7% of the total. If all the detected reference strings are assumed to be manually corrected, the accuracy of bibliographic information improves to 0.995.