A Hybrid Approach for Language Variety Prediction Using BERT and T5 Embeddings

Karunakar Kavuri, T. Raghunadha Reddy, Archana Gelli, B. Priyanka · 2025

A crucial component of author profiling is language variety prediction, which looks for stylistic and regional differences in language usage in written texts. Applications for this activity are substantial in fields including localization, forensics, and customized marketing. Conventional methods depend on manually created features or embeddings from separate models, which might not fully represent the nuanced linguistic differences. In order to improve language variety prediction performance, we presented a hybrid model in this work that integrates statistical features and embeddings from two sophisticated transformer models, such as T5 (Text-to-Text Transfer Transformer) and BERT (Bidirectional Encoder Representations from Transformers). Different statistical features identified in this work are Lexical features, Syntactic features, Orthographic and Phonological features, Stylistic features, Semantic features, Pragmatic and discourse-level features and Temporal features. The embeddings from both models are projected into a shared dimensions space, averaged, appended, and utilized as input to a downstream classifier in order to accomplish effective integration. This method makes use of T5's generative capabilities and BERT's contextual richness to better capture syntactic and semantic patterns in text. The English language diversity dataset from the PAN 2017 competition is used to test the suggested framework. According to experimental evaluations, the hybrid embedding technique achieves better accuracy and Fl-score for language variety prediction, outperforming individual transformer models by a wide margin. When compared to other models for language variety prediction, the XGBoost classifier with Statistical Features, BERT, and T5 embeddings performed better in terms of Fl-Score (93.4%) and Accuracy (93.5%). These results demonstrate the importance of including complimentary transformer embeddings to improve author profiling tasks' state-of-the-art performance.

Read the paper · More papers on PaperTik