Interactive Pretraining-Based Framework for Hindi-Chinese Parallel Sentence Pairs Extraction

Lianxi Wang, Junyang Zhong, Songxi Xu, Xuming Li, Zhuolin Chen · 2025

Information retrieval and machine translation performance is often unsatisfactory in low-resource languages, primarily due to the limited availability of bilingual parallel corpora for training models. Hindi, as a typical low-resource language, faces an acute scarcity of bilingual parallel data, especially in relation to Chinese, which hinders the representation capabilities of parallel sentence extraction models. To tackle this issue, this paper presents a framework that combines interactive and pretrained language models to extract parallel sentences. The framework incorporates architectural structure and training data generation, and is evaluated from closed and open perspectives using a constructed Hindi-Chinese parallel sentence corpus and the CCAligned web document-level parallel corpus. Experimental results demonstrate the exceptional performance of the constructed parallel sentence alignment model, enabling efficient and reliable extraction of bilingual parallel sentences from document-level corpora.

Read the paper · More papers on PaperTik