A connectionist model for classifying speech into silence, glottal source, burst friction, or mixed categories
Steven J. Sadoff · The Journal of the Acoustical Society of America · 1989
An algorithm for classifying speech into four classes (silence, only glottal source, only burst friction, or mixed) is being developed. This scheme primarily differs from the standard silence/aperiodic/periodic classification in that the defining characteristics of glottal source and burst friction sounds do not depend on periodicity distinctions, but on the locus of the energy concentrations. One male speaker reciting the Rainbow passage has been recorded and analyzed. Utilizing a strictly layered backpropagation network, the automated learning procedure is trained using the first 30 s of the passage; the final 75 s are used for testing. Quantitative results, along with several examples illustrating this model, will be presented. Additionally, the performance and classification strategy of the network will be compared to that of human observers. [Work supported by AFOSR.]