The Development of a Thai Telephone Conversational Speech Corpus
Sumonmas Thatphithakkul, Kwanchiva Thangthai, Sahatsawat Sriphol, Vataya Chunwijitra · 2023
To be the crucial language resource for acoustic model and language model training of the Automatic Speech Recognition system over telephone channel, a Thai telephone conversational speech corpus is created. This paper described the design and development of a Thai telephone conversational speech corpus. We obtained 37 hours and 28 minutes in 8 kilohertz of recorded real telephone conversational speech data from a private company leading in chemical products. This conversation related to a customers’ interview to finish a company’s satisfaction survey by agent. The speech data contains 26 hours and 3 minutes from 6 agents and 11 hours and 25 minutes from 4,383 customers. A few linguists were asked to transcribe all speech data to Thai text transcription and annotated the characteristics of speech and non-speech by the tags designed by research team. The statistical analysis of the corpus in terms of the number of utterances, tokens, vocabulary and the occurrences of sound loss tokens and English tokens are reported. The total of utterances found in the corpus are 36,122 utterances. There are 309,046 tokens which can be extracted to 5,522 vocabulary size. the percentage of sound loss tokens and English tokens found is 3.73% and 5.45% of all tokens in the corpus respectively. The top 5 words occurrences consist of Thai final particles showed politeness of male and female speakers and the adjective expressed the satisfy feeling. In conclusion, we examine the efficacy of models trained using our corpus in terms of recognition performance. According to the results of the experiment, the models we propose outperform the baseline model.