Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval

Kidist Amde Mekonnen, Yosef Worku Alemneh, Maarten de Rijke · 2025

Neural retrieval methods using transformerbased pre-trained language models have advanced multilingual and cross-lingual retrieval.However, their effectiveness for lowresource, morphologically rich languages such as Amharic remains underexplored due to data scarcity and suboptimal tokenization.We address this gap by introducing Amharic-specific dense retrieval models based on pre-trained Amharic BERT and RoBERTa backbones.Our proposed RoBERTa-Base-Amharic-Embed model (110M parameters) achieves a 17.6% relative improvement in MRR@10 and a 9.86% gain in Recall@10 over the strongest multilingual baseline, Arctic Embed 2.0 (568M parameters).More compact variants, such as RoBERTa-Medium-Amharic-Embed (42M), remain competitive while being over 13× smaller.Additionally, we train a ColBERT-based late interaction retrieval model that achieves the highest MRR@10 score (0.843) among all evaluated models.We benchmark our proposed models against both sparse and dense retrieval baselines to systematically assess retrieval effectiveness in Amharic.Our analysis highlights key challenges in low-resource settings and underscores the importance of language-specific adaptation.To foster future research in lowresource IR, we publicly release our dataset, codebase, and trained models. 1 * Equal contribution.

Read the paper · More papers on PaperTik