Learning and production of speech pattern using multilayer neural networks
Mitsuo Komura, Akio Tanaka · Systems and Computers in Japan · 1991
Abstract A new neural network and its learning algorithm are proposed. The neural network consists of four layers—input, hidden, output and final output layers. The hidden and output layers are multiple. The proposed learning algorithm is called the SICL (spread pattern information and cooperative learning) method and has the following features: (1) the singular points of back propagation errors in the BP method are removed; (2) a spread pattern information (SI) algorithm is proposed; and (3) a cooperative learning (CL) algorithm is proposed. Using the SICL method, it is possible to learn analog data accurately and to obtain a stable output. Using this neural network, the authors developed a speech production system consisting of a phonemic symbol production subsystem and a speech parameter production subsystem. The system is applied to speech data and the system performance is examined. Especially, for the speech parameter production subsystem, the learning capacity and the optimal area for learning constants are studied. As a result, it is shown that it is possible to learn and produce speech data with high accuracy using the proposed neural network.