The role of temporal information in automatic speech and voice recognition

Erdem Baha Topbas, Srikanth R. Madikeri, Elisa Pellegrino, Volker Dellwo · Universitätsbibliothek Trier · 2026

Speech signals contain three types of temporal information—amplitude envelope (ENV), temporal fine structure (TFS), and periodicity (PER)—which differ in their dominant fluctuation rates and in how they encode prosodic and segmental contrasts (Rosen, 1992). The periodicity captures the zero-crossing information contained in the waveform, the envelope captures signal intensity across the time domain, and the temporal fine structure focuses on the fluctuations of high rates. The role these cues play in speech intelligibility has been partially researched (Licklider & Pollack, 1948; Shannon, et al, 1995). These works do not explore speaker recognition and automatic speech recognition. This project conducts a systematic analysis to address this gap in the literature.We utilize systematically distorted signals comprising eight conditions: three versions retaining only one temporal property (ENV, TFS, or PER), three retaining two properties (ENV+TFS, ENV+PER, PER+TFS), and conditions in which all properties were either fully preserved or fully removed (creating a square wave). Signal manipulations included Hilbert transformation, infinite-peak clipping, and a novel method termed isochronous time warping, which alters signal periodicity.The speaker verification and classification results were explored by training a ResNet18 model and a series of Gaussian Mixture Models (GMMs) respectively. They were trained on clean and undistorted speech from the training subset of the TIMIT corpus (Garofolo et al., 1993) and tested on the test subset with the conditions outlined above. ResNet18 performance was assessed using equal error rate (EER). In the speaker verification task, performance was highest when all features were present and lowest when all were removed. Notably, removing periodicity alone produced the only statistically significant degradation, with EER rising from 17.2% to 48.3% (p = 0.0006). Analysis further revealed that while removing ENV or TFS individually had no significant effect, removing both together caused a meaningful drop in performance, with EER increasing from 6.3% in the baseline to 29%.These results can be seen in Figure 1. The GMM classification task yielded a more categorical result: any removal of temporal information—regardless of which property or combination was affected—produced a significant drop in classification accuracy. Speech intelligibility results were conducted using a pre-trained OpenAI Whisper Automatic Speech Recognition (ASR) model (Radford et al., 2022). This model was given utterances from the Audio MNIST corpus (Becker et al., 2024), distorted in the same condition as above. The corpus contains digits 0-9 being read by a variety of speakers. The transcripts acquired by Whisper were either directly or fuzzily matched to the digit options to get accuracy metrics. The fuzzy matching was done by applying the metaphone algorithm (Philips, 1980) to both the transcripts and the possible outputs to account for homophones and were matched by picking the candidate with the smallest Levenshtein distance (Levenshtein, 1965). The highest performance was obtained with the unmodified utterances, with a transcription accuracy of 96.5%. The ENV, TFS, and ENV+TFS conditions all performed at chance level (~10%). Taking the ENV+PER+TFS case as baseline, the mixed linear model results indicated statistically significant differences from the baseline in all distorted cases.Ranking the conditions from best to worst in terms of EER and transcription accuracy shows a similar ordering, with an intraclass correlation coefficient (2,1) score of 85%. This demonstrates their agreement on PER being the most important type of temporal information, with important implications in the design of robust models in both voice and speech recognition, by identifying possible data augmentation methods, and adjustments model architecture. These results show that PER is the most reliable feature, hinting that the models might be over-reliant on them. Therefore, training that involves reducing the emphasis on PER can increase accuracy forensic uses.

Read the paper · More papers on PaperTik