Performance of current models of speech recognition and resulting challenges
Wiebke Schubotz · Carl von Ossiezky University of Oldenburg · 2015
Speech is usually perceived in background noise (masker) that can severely hamper its recognition. Nevertheless, there are mechanisms that enable speech recognition even in difficult listening conditions. Some of them, such as e.g., the combination of across-frequency information or binaural cues, are studied in this dissertation. Moreover, masking aspects such as energetic, amplitude modulation or informational masking are considered. Speech recognition in complex maskers is investigated that systematically vary in their spectro-temporal properties and address all aspects listed above. Outcomes of current models of speech recognition are compared to the data observed in the listening experiments. This allows to assess how well the different models account for the observed speech reception thresholds, as each model incorporates different signal analysis strategies. The studies designate the limits of the current model approaches, and thus constitute a benchmark for speech recognition models which might be useful for improving our current state of the art in modelling speech recognition.