Classification of speech under stress based on physical characteristics of vocal folds vibration
Xiao Lan Yao · Institutional Repositories DataBase (IRDB) · 2013
The performance of automatic speech recognition systems (ASR) is degraded by stressinducing environments, such as those where noisy backgrounds, multi-tasking, fatigue, emotional situations, adverse physical conditions, and high workload stress are present. The study of speech under stress, also referred to as stressed speech, can help maintain and improve the robustness of speech recognition systems in these situations. The primary objective of this dissertation is to perform the classification of speech under stress based on a physical model. Presence of stress will cause speakers to change their physiological system to react and adapt himself to the stressed condition. Changes in physiological characteristics can result in variations in aerodynamics in the glottis and the vocal tract, and the stressed speech is produced. Therefore, a physical model is necessary to model airflow patterns in the physiological system in order to represent the process of speech production. An investigation on how physical model is used for classification of speech under stress (stress classification) is explored. This dissertation includes following objectives: I. Explore the physical characteristics of the vocal folds during speech under stress and estimate the physical parameters of the vocal folds for stress classification (the first study); II. Explore the physical characteristics of both the vocal folds and the vocal tract for stressed speech and propose the effective physical parameters (the second study); III. Model the aerodynamics in the laryngeal ventricle and the false vocal folds, and estimate a parameter for classification of stressed speech (the third study); IV. Use different classifiers to perform classification and compare their performance (the