Eliciting analogical reasoning from language models in retrieval-augmented translation under low-resource scenarios
Liyan Wang, Bartholomäus Wloka, Yves Lepage · Neurocomputing · 2025
Retrieval-Augmented Neural Machine Translation (RANMT), which augments translation models with relevant examples fetched by a similarity retriever, is proficient in well-resourced translations. However, its inherent advantages are not fully realized in low-resource contexts, where the sparsity of data can often result in less relevant or less useful information for translation. Recent literature indicates that nearest-neighbor examples from small training data have the unfortunate effect of impairing RANMT performance. Our examination of 16 low-resource tasks reveals a sharp deterioration in performance of a multilingual language model when conditioned on retrieved examples compared to direct translation. To address this problem, we explore a framework based on analogical reasoning, aiming to enhance the capacity of language models to infer translations from parallel examples in limited data settings. This framework mimics a cognitive process of human translation by structuring examples in analogy patterns. We propose a multi-objective learning strategy that augments vanilla training for conditional translation to learn latent knowledge from examples. We also investigate different retrieval methods for selecting translation examples based on lexical similarity, semantic relatedness, or a combination of both. The results show that our approach is effective in optimizing RANMT in low-resource settings, delivering notable improvements across all retrieval settings. In particular, augmented training akin to reasoning with analogies in two directions, contributes significantly to deriving benefits from examples, even when their relevance is limited. Moreover, our approach demonstrates superior performance in low-resource translation tasks compared to prompting large language models in few-shot contexts. It also proves to be competitive with models that have been extensively trained using substantial amounts of supervised data.