Speech enhancement for forensic applications
Andrew John Fisher · QUT ePrints (Queensland University of Technology) · 1995
Law enforcement agencies often engage in surveillance operations which involve the recording of spoken conversations. As is often the case, these recordings are made with a single microphone under covert conditions. Under this non-ideal situation, the speech signal is highly susceptible to be severely corrupted by various forms of noise, the most common of which is broadband in nature. This thesis presents a study conducted to investigate the enhancement of speech recordings for forensic applications. A new speech enhancement scheme has been proposed here, to provide noise reduction without compromising the intelligibility of the speech. The scheme implements a hybrid approach combining both spectral and root-cepstral subtraction. Extensive testing using both subjective and objective based intelligibility and acceptability assessment schemes, indicate that the system is successful in providing intelligibility improvement and superior signal-to-noise ratio with minimal spectral distortion. In addition, the proposed system was also tested in the capacity as a preprocessing stage to other speech applications such as speech recognition, speaker recognition and speech coding. The system proved to be beneficial for speech coding, while application to the recognition techniques was limited despite showing positive potential. Finally the system was implemented in real-time and was found additionally successful when applied to enhancement of speech transmitted over High Frequency communication channels.