Local and Global Context Feature Fusion for Effective Spoken Language Identification in Indian Linguistics

A R Ambili, Rajesh Cherian Roy · 2024

Spoken language identification (SPLID) in a mul-tilingual context, particularly within Indian languages, presents a significant challenge due to the diversity and complexity of phonetic and linguistic patterns. Integrating image classification features with mel-spectrograms in SPLID tasks offers a powerful approach to leveraging pre-trained models, cross-modal knowledge transfer, and enhancing the discriminative capacity of the SPLID system. This paper introduces a novel approach using local and global contextual information in the audio to improve SPLID performance. We propose a hybrid model that integrates features extracted from a Vision Transformer (ViT) and VggNet to capture intricate temporal and spectral characteristics of audio signals. Extensive experiments were conducted on the IIIT-H Hyderabad dataset, encompassing a comprehensive collection of Indian languages. With our method obtaining 99.21%, up from 95.93% using classic SPLID procedures, our results show a considerable improvement in accuracy. This substantial increase highlights the effectiveness of integrating local and global contexts, leading to a more robust and accurate spoken language identification.

Read the paper · More papers on PaperTik