A Triple-Form Chinese-Mongolian Bilingual Dictionary Dataset
Yingli Shen, Liang-Chih Yu, Xiaoke Qi, Xiaobing Zhao · Data Intelligence · 2025
While bilingual dictionaries between source and target languages can signiffcantly enhance translation performance in large language models (LLMs) by providing critical lexical alignment signals. Notably, no publicly available word-level triples exist for Chinese, Mongolian, and theirlatinized counterparts. To address this accessibility challenge, we create a new dataset that contains 30, 000 ⟨Chinese, Mongolian, Mongolian(latin)⟩ word-level triple data. More speciffcally,We ffrst leverage cross-lingual word induction to extract a Chinese-Mongolian bilingual dictionary from a large-scale and readily available Chinese-Mongolian comparable corpus. Then,with the help of ffve Mongolian language experts, we conduct quality assessment, manual proofreading, and latin transcription of the extracted dictionary to generate the triple dataset. Theopen-source release of this dataset advances Chinese-Mongolian cross-lingual information processing research while establishing a replicable framework for constructing bilingual dictionariesin other low-resource language contexts.