Speaker identification based on a modified Kohonen network

Karane Vieira, Bogdan M. Wilamowski, Robert F. Kubichek · Proceedings of International Conference on Neural Networks (ICNN'97) · 2002

A human information processing system is composed of neurons switching at speeds about a million times slower than computer gates. Yet humans are more efficient than computers at computing complex tasks such as speech and visual interpretation. A neural network (NN) method was developed to reproduce one of the abilities and power of the human brain: speaker recognition. To realize this method, the input patterns used for this network were the reflection coefficients (k/sub p/) of the speech signal. The speech was analyzed using an autoregressive LPC technique and the k/sub p/s were found by applying the Levinson recursion algorithm. Next, a simple transformation of these input patterns onto a hypersphere in augmented space was made using a multilayer perceptron (MLP) neural model. The Kohonen approach is commonly used for computing a distance in multidimensional input space so that the input patterns are projected on a hypersphere with unit radius. Although this technique is very efficient for clustering of patterns, it has one significant drawback. By normalizing the input pattern, the information about its magnitude is lost. The proposed modified network has a relatively simple architecture but is shown to be very effective in performing speaker recognition.

Read the paper · More papers on PaperTik