Mobile‐Based Bilingual Speech Corpus
Nivedita Palia, Deepali Kamthania, Ashish Pahwa · 2021
In the context of speech processing, there is a massive significance of corpus which provides the credibility of all the experimental conclusions carried out with it. Bilingual speech corpus is a multispeaker speech corpus that is primarily oriented in Hindi and the English language. The present study describes the methodology and experimental setup used for the generation and development of bilingual speech corpus. Continuous and spontaneous (discrete) speech samples from both languages are collected on different mobile phones in a real-time environment, unlike studio environments. A brief comparative study of 18 readily available multilingual speech corpus developed for Indian languages is made against the proposed corpus. An annotation scheme is discussed which helps us later to carry out further study how the recognition rate varies on the basis of language, device, text, and utterance.