Non-linear filtering for multilingual speech: A physics-inspired transformer framework for joint denoising and recognition

Omkar Vilas Sawant, Anirban Bhowmick · AIP Advances · 2025

Environmental noise severely degrades speech intelligibility and downstream processing. This paper presents a physics-inspired, transformer-based deep neural network for robust speech denoising. The model leverages complementary perceptually motivated acoustic features—including gammatone frequency cepstral coefficients, power-normalized cepstral coefficients, RelAtive SpecTrAl-Perceptual Linear Prediction (RASTA-PLP), modulation spectra, and cepstral frequency cepstral coefficients—to capture essential speech cues while suppressing noise. Analysis shows that different noise types (e.g., pink, babble, and transient) corrupt distinct spectrotemporal regions. This insight informed the model’s design, particularly its non-linear attention mechanism, which dynamically emphasizes clean speech components and suppresses localized noise distortions. We evaluate denoising effectiveness using multilingual Spoken Language Recognition (SLR) as a proxy for intelligibility. Experiments on a noisy Indian language corpus (−10 to −15 dB signal-to-noise ratio) and the Common Voice dataset demonstrate significant superiority over classical methods (e.g., Wiener filtering) and other deep neural network approaches. The proposed transformer model achieved the highest SLR accuracy, notably 97.18% on the Indian corpus, confirming its ability to preserve spectral and temporal speech integrity. Results consistently highlight the generalizability and robustness of this physics-guided, attention-based non-linear filtering approach across diverse multilingual speech.

Read the paper · More papers on PaperTik