Improving Speech Intelligibility in Noise Using a Binary Mask That Is Based on Magnitude Spectrum Constraints
Gibak Kim, Philipos C. Loizou · IEEE Signal Processing Letters · 2010
A new binary mask is introduced for improving speech intelligibility based on magnitude spectrum constraints. The proposed binary mask is designed to retain time-frequency (T-F) units of the mixture signal satisfying a magnitude constraint while discarding T-F units violating the constraint. Motivated by prior intelligibility studies of speech synthesized using the ideal binary mask, an algorithm is proposed that decomposes the input signal into T-F units and makes binary decisions, based on a Bayesian classifier, as to whether each T-F unit satisfies the magnitude constraint or not. Speech corrupted at low signal-to-noise (SNR) levels (-5 and 0 dB) using different types of maskers is synthesized by this algorithm and presented to normal-hearing listeners for identification. Results indicated substantial improvements in intelligibility over that attained by human listeners with unprocessed stimuli.