Evaluating Speech Intelligibility for Cochlear Implants Using Automatic Speech Recognition
Hengzhi Zhou, Mingyue Shi, Qinglin Meng · 2024
Evaluating speech intelligibility is a crucial stage in the development of sound processing algorithms for cochlear implants (CI). While the cost of recruiting actual CI patients and carrying out subjective experiments is high, objective measurements provide a cost-effective alternative with the added benefit of enabling large-scale testing. Most objective measurements were designed to evaluate different font-end noise reduction algorithms, but another critical aspect in CI development is coding strategy. In this study, the automatic speech recognition (ASR) network OpenAI Whisper was used to simulate a human perception model listening to CI-simulated speech. The speech was generated by a vocoder simulating the most used strategy, the Advanced Combination Encoder (ACE) strategy. An objective ASR experiment was conducted to compare with a previous subjective study that explored the relationship between two key parameters, i.e., the number of maxima (nmax) and electrical dynamic range (EDR), in the ACE strategy with both normal hearing and CI listeners. Although ASR performed poorly in recognizing words under severely distorted conditions, it exhibited a similar result trend to that of normal-hearing and CI users. This pilot study demonstrates the potential of using well-trained ASR and vocoder systems for objective CI speech intelligibility evaluation. Future works on fine-tuning ASR and comparison to traditional objective metrics are warranted.