Dysphonic voice classification using wavelet packet transform and artificial neural network
Adalberto Schuck, Letícia Vieira Guimarães, J.O. Wisbeck · 2004
In [Schuck Jr. and Parraga, 2002] work, it was demonstrated the viability of the wavelet packet transform (WP) and the best basis algorithm (BBA) as a feature extractor (FE) for a dysphonic voice classification systems. It was shown the better choices of wavelet and cost functions were Symlet 5 and Shannon entropy. Also, a linear discriminator between normal and dysphonic voices was performed. This work present the use an artificial neural network (ANN) in addition to WP and BBA to perform a non-linear discriminator. The WP with 5 dilatation levels and BBA of the sustained vowel /a/ of 13 normal and 51 dysphonic previously diagnosed subjects were performed. Then the entropy values of each best tree's nodes were used for the classification. A ANN was designed with 3 layers ( 4 neurones in the hidden layer and 2 in the last layer). The non-linear function was hyperbolic tangent. The ANN was trained using backpropagation with a group of 6 normal and 21 dysphonic subjects chose at random from the database. Then, all the 61 subjects were classified. The system had a success rate 84.3%, with 4.6% of false negatives and 10.9% false positives.