Language Diarization Model for bilingual Code-Switched Speech Analysis
G. Mohan Dhanush, J. Gopichand, B. Balasaigayatri · 2025
Around 7,000 different languages are spoken in this world; therefore, many countries, such as Singapore, Malaysia, and the Netherlands, have more than one official language. The country itself forms part of the large category of people whose literacy includes at least two different languages, so it should minimize language problems for communication as with language diarization. There are many effective language diarization systems available for the Asian languages, which include Mandarin and Cantonese for Asia, as well as for Europe, that are Frisian and Dutch. However, very few of these systems rely on Indian languages such as Telugu, Kannada, or Gujarati. So, this is highly significant to develop a diarization system especially dedicated to Indian languages. In this experiment, we carry out the diarization analysis on bilingual and multilingual data since the testing database used for the algorithms proposed in the previous section has no samples of multilingual recordings. For testing the language diarization algorithm, we have used a self-recorded database with bilingual and trilingual recordings of English, Kannada, and Telugu languages. The proposed voice diarization system includes voice activity detection, speech-to-text conversion, word tokenization, and applications in translation and summarization of an audio sample. In the experimental result, the system proposed here shows a word accuracy rate of 73.957%, which further supports its effectiveness in handling multilingual audio recordings involving Indian languages.