Perception-based Objective Estimators of Speech Quality
Stephen, Connie Sholl · 1995
Four proposed perception-based techniques for objectively estimating speech quality and three traditional estimators are applied to coded speech samples. Agreement between objective estimates and corresponding subjective test scores is reported. Several observations on key elements of perception-based estimators are offered. 1 Background Speech coding often involves a four-way compromise among complexity, delay, bit-rate, and the perceived quality of decoded speech. The most critical perceived quality measurements will always rely on formal subjective tests. However, the costs associated with formal subjective tests are not justified in some situations. Specifically, much coder development work relies on objective estimators of perceived speech quality, al ong with “informal listening tests.” For example, of the 30 coders described at the 1993 IEEE Workshop on Speech Coding for Telecommunications, only nine had been tested in formal subjective tests, while several different objective quality estimators were employed [1]. Segmental SNR (SNRseg) was applied in four cases, spectral distortion measures were used in three cases, while SNR, perceptually-weighted SNRseg (PWSNRseg), Bark Spectral Distortion (BSD), and Cepstral Distance (CD) were each used once.