Thai-Dialect: Low Resource Thai Dialectal Speech to Text Corpora
Artit Suwanbandit, Jaturong Chitiyaphol, Sutthinan Chuenchom, Kanyarat Kwiecien, Husen Sawal, Ruslan Uthai, Orathai Sangpetch, Ekapol Chuangsuwanich · 2023
We release a speech-to-text benchmark dataset containing 10 Thai dialects that cover different regions of Thailand. Our corpora consists of the standard dialect, Thai-central (THA); the northern dialects (Khummuang (NOD), Nan (KHB) and Yno (YNO)); the northeastern dialects (Korat (TTS), Khmer (KXM) and Laos (TTS)); and the southern dialects (Krabi (SOU), Pattani (MFA) and Phangnga (SOU)). All transcriptions are based on the Thai writing standard. We constructed baseline models by fine-tuning from self-supervised pre-trained models. Results show that multilingual/multidialectal systems outperform monolingual ones, and different dialect combinations can affect the performance of multilingual/multidialectal training.