A knowledge-based temporal alignment system involving stochastic methods
Harouna Kabré, J. F. Malet, Jean-Marie Pécatte, Guy Pérennou, Nadine Vigouroux · The Journal of the Acoustical Society of America · 1990
In this paper the principle behind an automatic signal alignment system, used for speech database labeling, is described. This system is based on both a phonetic knowledge subsystem and a stochastic training procedure. It operates in two passes. During the first one, for each frame, temporal cues are worked out, similar to those handled by the hand labeler. The frames are next recombined into phonetic events that are assigned broad labels—e.g., vocalic, fricative, liquid, sonorant, occlusive, etc. The second pass aligns phonetic events onto a phonetic transcription—provided by an expert. It resorts to a stochastic model entailing (i) a phonetic knowledge base and (ii) a channel modeling first-level decoding errors. The training of the phonetic knowledge base consists in fine tuning the assessment of probability weights. This is done on hand-labeled corpora. Results, on both French and English corpora, will be presented, attention being paid to the problem of system multilingual adaptability.