A Comprehensive Theoretical and Empirical Framework for Fine-Tuning the CoRover's BharatGPT Transformer for Indic Languages

Ankush Sabharwal -, Vikas Tripathi -, Onkar Nath - · International Journal For Multidisciplinary Research · 2025

The widespread adoption of transformer-based models in natural language processing (NLP) has led to significant breakthroughs in numerous languages. However, models like BharatGPT (from CoRover.a) - though robust for high-resource languages - require specialized adaptation to effectively handle the rich morphological and syntactic diversity of Indic languages. In this paper, we propose a comprehensive framework for fine-tuning the BharatGPT transformer to support Indic languages. Our approach integrates tailored data preprocessing, script-specific embedding enhancements, and rigorous convergence analysis. We derive key theoretical properties of the fine-tuning algorithm, including a convergence theorem under Lipschitz continuity and bounded gradient variance assumptions, and we validate our approach with empirical evaluations using standard metrics such as perplexity, BLEU, and F1 score. The results demonstrate significant improvements across several Indic languages, thereby underscoring the effectiveness of our methodology.

Read the paper · More papers on PaperTik