ASHI: A Database of Assamese Accented Hindi
Joyshree Chakraborty, Rohit Sinha, Priyankoo Sarmah · 2023
In this paper, a spontaneous, non-native Hindi speech database called ASsamese accented HIndi (ASHI) is described. This database has been collected from the upper, lower, and central Assam regions. It consists of 7.5 hours of speech data from 78 speakers. The speech data collection was done remotely via mobile/internet telephonic calls. The conversations with participants were recorded at our end using the default recording option available on the mobile phone, Skype, or Zoom. The participants spoke on the chosen topics in a spontaneous manner which involved frequent Hindi-English code-switching. Thus, 39% of the total vocabulary of the database is English words. The metadata collected from the speakers include information about their linguistic backgrounds. The manually corrected transcriptions are also provided, and the collected speech data are analyzed in terms of the number of unique words and breath groups per speaker. Finally, the salient metadata are visualized through tSNE plots.