Bootstrap masker generation method for speech masking systems
Yosuke Kobayashi, Kazuhiro Kondo · 2014
Currently, speech masking systems make use of pre-recorded speech signals to generate maskers, which we call offline generated maskers. The offline masker includes no feedback from the speech environment, and only the overall averaged masking effect of the speech is gained, not the environment. In our prior study, a speaker-dependent masker that is generated by mixing pre-recorded speech of the speaker to be masked has been found to be the most effective, and can mask at low sound levels. Accordingly, the real time acquisition of the masked speaker speech signal is required to create the masker, which we shall call online generated maskers. In the online masker generation, there is a problem that sufficient amount of sound material may not be available from the cached memory in real time. Therefore, we have applied the bootstrap method used in machine learning techniques, and generated a masker as if many speech samples from a small amount of speech is available. We tested the proposed masker using two subjective indexes, i.e., annoyance and listening difficulty. We used sentence speech signals recorded with a dummy head. We compared the bootstrap type online masker (BS), the ring buffer addition type online masker (RA), and 3 types of offline maskers. These maskers were played at three signal noise ratio (5, 0, -5 dB). As a result, the annoyance scores of the online maskers were about the same as the offline maskers. However, the listening difficulty scores improved, and the BS type online masker was the most effective masker when the SNR is higher, at 5 and 0 dB. In addition, the masking effect of the speaker dependent condition (masker created from the target speaker speech) using the BS-Human Speech-Like Noise (HSLN) was found to be significantly higher than others, especially at the Target to Masker Ratio (TMR) of 5 dB.