TibOCR-Bench: A Comprehensive Benchmark and Training Pipeline for Tibetan Multimodal OCR
LAMA Jie, Manla Cairang, Kuntharrgyal Khysru, Hua Guo Cai Rang, Jiahui Jiao, Yue Yingkai, Tan Qian · Data Intelligence · 2025
In recent years, large multimodal language models (such as GPT-4V and Gemini) have achieved significant advancements in natural language processing and vision-language tasks. However, their capabilities in text-related visual tasks for low-resource languages, particularly Tibetan, re- main insufficiently explored. To bridge this research gap, this paper proposes a systematic data construction pipeline specifically designed for Tibetan Optical Character Recognition (OCR) tasks, supporting comprehensive multi-scene and multi-task model training and evaluation. Based on this pipeline, we constructed two benchmark datasets: (1) TibOCR-Bench, the first eval- uation benchmark dedicated to Tibetan OCR, encompassing 30 diverse application scenarios, ensuring high-quality and uncontaminated data; and (2) TibOCR-Train, which specifically ad- dresses human preference-aligned training data requirements for Tibetan OCR. Considering the unique challenges of Tibetan OCR tasks, we developed a standardized evaluation framework, including clear inference protocols and quantitative metric calculations tailored to effectively as- sess model performance. Experimental results demonstrated that multimodal baseline models fine-tuned via instruction tuning and reinforcement learning-based human preference alignment strategies significantly outperformed their un-tuned counterparts, highlighting the effectiveness and practicality of the constructed datasets. Further analysis revealed these models’ strengths and weaknesses across various scenarios, including recognizing printed Tibetan text, handwrit- ten content, non-semantic text, and structured table extraction. To promote further research in Tibetan OCR, we have made publicly available the entire data construction pipeline, datasets, trained models, as well as the training and inference code. These open-source resources provide a robust foundation for future research and innovation in zero-shot multimodal techniques for low-resource languages.