Speech Processing using Artificial Neural Networks

Roberto B. Togneri, Mike Alder, Yianni Attikiouzel · 1990

A three layer perceptron network is used to classify the /i/ sound using isolated words from different speakers. A classification accuracy of 97% has been achieved. A map of phonemes is used to trace trajectories of utterances using the self-organising neural network. A crinkle factor is proposed which allows using the self-organising map to determine the inherent dimensionality of a set of points. By this technique speech data has been shown to possess an inherent dimensionality of at least four. A projection of the map and the speech data shows how the self-organising map fits the speech space. INTRODUCTION Neural networks have existed for a long time [1, 2] and have recently enjoyed a resurgence of interest, in particular their application to speech processing. If a set of utterances can be labelled (i.e. each frame is associated with a phoneme) then a supervised neural network like the multilayer perceptron [3, 4] can be used. Without making any assumption concerning the labelling ...

Read the paper · More papers on PaperTik