A Diffusion Model Based on Hilbert–Huang Reconstruction for Pathological Voice Generation

Yuyang Jiang, Mingxuan Yan, Xiaojun Zhang, Zhi Tao · IEEE Transactions on Audio Speech and Language Processing · 2025

Most existing speech generation models require substantial amounts of learning data, significantly limiting their effectiveness when working with limited pathological voice samples. In this study, we propose a Diffusion Transformer model based on Hilbert-Huang reconstruction tailored for pathological vowels with limited data. Our model takes the pitch of pathological voice samples as input to generate new samples with authentic pathological acoustic characteristics. We employ a diffusion model as the generation framework and use a Transformer model as the denoising network within the diffusion process, enhancing the quality of generated speech. To capture the nonlinear and non-stationary characteristics of pathological voices in generated samples, we introduce a Hilbert reconstruction Decoder. This decoder incorporates the Hilbert-Huang transform to learn and retain pathological features, enhancing the presence of nonlinear and non-stationary components in the generated samples, making them closely resemble real pathological voices. Additionally, we propose a Pitch Learning Strategy that decomposes an entire pathological voice sample into segments based on its pitch period, effectively augmenting the model’s learning data and addressing the scarcity of pathological voice samples. We conducted experiments on three international pathological voice datasets, and under various acoustic evaluation metrics, the generated data closely resembled the characteristics of real data. In the classification experiments with data augmentation, our method achieves an average accuracy that is 4.4 p.p. (percentage points) higher than the best-performing models among the baselines.

Read the paper · More papers on PaperTik