TSynC-3miti: Audiovisual Speech Synthesis Database from Found Data
Ausdang Thangthai, Sumonmas Thatphithakkul, Kwanchiva Thangthai, Arnon Namsanit · 2020
Building audiovisual speech synthesis database is a crucial factor in the applications of audiovisual speech synthesis systems. Typically, most databases captured on soundproof studio and hired a professional voice talent who speak clearly articulation and able to act and control their voice to read prepared scripts. However, the major drawbacks of conventional audiovisual speech databases are small, costly and time-consuming. Hence, this paper tackles these drawbacks and focuses on building a large audiovisual speech synthesis database using freely available noisy found data on the Web instead of recording clean data. This database, called TSynC-3miti, is the first Thai audiovisual speech synthesis database which are designed for audiovisual speech synthesis use, such as HMM/DNN-based Speech Synthesis System (HTS). Tons of video data have been collected from the `3mitinews' channel on YouTube, which was broadcasted between 04-Jan-2017 and 23-Jan-2020. This paper introduces a procedure of data preparation from scratch, including face detection, text transcription, phoneme labels, and audiovisual data cleaning and feature extraction. The total video contains approximately 19 hours and also producing in audio, images, text transcriptions and phonetic labels.