An efficient algorithm to select phonetically balanced scripts for constructing a speech corpus
Min-Siong Liang, Ren-Yuan Lyu, Yuang-Chin Chiang · 2004
Here, we describe an efficient algorithm to select phonetically balanced scripts for collecting a large-scale multilingual speech corpus. It is expected to collect a multilingual speech corpus covering three most frequently used languages in Taiwan, including Taiwanese (Min-nan), Hakka, and Mandarin Chinese. To achieve the objective, the first step is to construct a multilingual phonetic alphabet, namely Formosa phonetic alphabet (ForPA). In addition, the multilingual lexicons (Fomosa lexicons) are also important parts for building the corpus. Until now, this corpus containing 600 speaker's speech of Taiwanese (Min-nan) and Mandarin Chinese has been finished and ready to release. There contains about 40 hours of speech in 247 thousand utterances in this release.