Phoneme alignment of filipino speech corpus
R.G. Sagum, R.A. Ensomo, Emerson M. Tan, Rowena Cristina L. Guevara · 2004
Segmentation and transcription of a speech corpus is a prerequisite in the development of an automatic speech recognition (ASR) system. In this paper, we develop a method for automatically segmenting and transcribing the Filipino speech corpus that is being developed at the DSP laboratory. A multi-layer perceptron (MLP) will take speech feature inputs, multiply them by weights computed from a training set of labeled speech. The system is based on a multi-layer perceptron and start synchronous decoder. The corpus was divided into three subcorpora, the paragraphs and sentences sub-corpus (par+sen), the words sub-corpus and the syllables sub-corpus. For the par+sen sub-corpus, we obtained a 62.64% phoneme recognition rate with 75.68% of labels within 20 ms of hand-labeled transcriptions; for the words-subcorpus, 63.93% phoneme recognition rate with 72/38% within 20 ms of hand-labeled transcriptions; and for the syllables sub-corpus, 72.60% phoneme recognition rate with 75.69% within 20 ms of hand-labeled transcriptions.