Study on efficient prior control for realizing practical systems of speech recognition
哲 小橋川 · Institutional Repositories DataBase (IRDB) · 2013
Due to the continual improvements in computer resources on the cloud and smart devices, applications based on speech recognition technologies are becoming more widely used.However, recognition accuracy degraded significantly if the speech is noisy, captured in real environments, or spontaneous, containing ambiguous utterances; both problems are barriers to the practical application of speech recognition.This paper provides useful countermeasures in the form of efficient prior control schemes by leveraging the attributes of practical scenarios.This research assumes two practical usage targets, a) noise robust speech recognition on tablet devices, and b) spontaneous speech recognition for contact centers and parliament bodies.This work deals with three situations of speech application; i) speech interface for tablet devices, ii) speech mining in contact centers, and iii) speech transcription for parliamentary meetings.We develop, for the first situation, i.e. tablet devices, 1) acoustic model adaptation and normalization using pre-observed noise.For the second situation, i.e. contact center, we develop 2) a fast unsupervised adaptation technique using frame-independent confidence scores, 3) a data selection technique using prior confidence, 4) a recognition time stabilization technique using prior beam width control.We also develop 5) fast acoustic pre-processing for the transcription of parliament meetings, third situation.To tackle the variation in S/N and convolutional noise expected with tablet devices, we develop acoustic model adaptation and normalization using pre-observed noise.This technique assumes that the background noise is relatively stationary and can be captured.It offers robust speech recognition under a wide range of S/N and convolutional noise; the noise captured prior to speech recognition allows noise reduction through the techniques of spectral subtraction (SS), additive noise adaptation using parallel model combination (PMC), and convolutional noise normalization using cepstral mean normalization (CMN).To improve accuracy under the constraint of a recognition time limit, often seen in contact centers, we develop a fast unsupervised adaptation technique based on frame-independent confidence scores.This technique leverages the property that the target is stored speech.It uses a limited number of Gaussian mixture models (GMMs) for the target speech in a preliminary step before speech recognition, and then improves accuracy rapidly by gender selection and of the application of maximum likelihood linear regression (MLLR).For contact centers, we develop a data selection technique with prior confidence estimation to reduce the cost of processing by dropping low confidence speech data; if such data is processed it is likely to disrupt the subsequent text mining functions and will eventually be rejected.The property of this technique is that massive volumes of data are stored and low confidence recognition results i ii ABSTRACT are unnecessary.It estimates prior confidence scores rapidly by using a limited number of GMMs and selects only high prior confidence data for speech recognition.Furthermore, for contact centers, we develop a prior beam width control technique to reduce the time wasted in processing low quality speech data that should be rejected.This technique also assumes that the target is massive volumes of stored speech.It rapidly estimates prior scores by using a limited number of GMMs, and stabilizes the computation time by controlling the search space spread in decoding.In addition, we also develop a fast acoustic pre-processing technique that can well handle changes in the recording environment and speaker to realize a parliamentary meeting transcription system.We can leverage the property that pre-processing is available since incoming parliament speech is segmented and temporarily stored in caches.This technique achieves high accuracy, even if computation time constraints are imposed, by combining four fast acoustic pre-processing methods of channel selection, speaker indexing, feature parameter normalization, and unsupervised adaptation.All five proposed techniques are based on the prior control approach and leverage the properties of practical speech recognition applications; they provide significant benefits over conventional speech recognition schemes. ABSTRACT IN JAPANESE