Research on Enhancing the Efficiency of Computer-Aided Translation Memory Retrieval Based on BERT Pre-trained Model
Yan Zhu · 2025
This paper presents a innovative composite framework for translation memory retrieval in computer-aided translation. This framework achieved an optimism by fusing technologies between semantic representation based on BERT and hierarchical navigable small world(HNSW) graph indexing. Based on the architecture of XLM-RoBERTa, multi-scale encoding mechanism has been developed and meanwhile an adaptive term weighting scheme is innovated for the first time. This theme achieved an 12.4% improvement of F1 score when parameter configuration α=0.2、β=0.05. This scheme combined textual surface features for editing distance metrics with deep semantic matching creatively to form a dual-channel ranking algorithm. This scheme has a great improvement in comprehensive evaluation in WMT2017 English-Chinese data collection. Compared with Trados Studio 2022 basic system, improvement of 18.7%(0.62→0.73) for MRR@10 index has been obtained in bio-medical document index retrieval and 32 milliseconds of computational latency has been reduced at the same time. A segmented correlation pattern between semantic matching thresholds and system effectiveness has been identified after deep analysis in which shows the obtaining of the optimal recall-precision balance occurring when θ values reside within [0.72,0.78]. This framework can be configured with terminology in different domains. And 91.4% concept coverage can be achieved in law related text evaluation which has been improved by 37.5% compared with traditional keyword-based approaches. These methodology breakthrough has established a transferable paradigm for building a high-performance translation memory system with modern adaptive neural machine translation framework.