MeetingLogger

Rohit Prasad, Long Nguyen, Richard M. Schwartz, John I. Makhoul · 2002

In this paper we describe our on-going effort in developing a speech recognition system for transcribing courtroom hearings. Court hearings are a rich source of naturally occurring speech data, much of which is in public domain. The presence of multiple microphones coupled with presence of noise and reverberation makes the problem simultaneously rich and challenging. We have exploited the availability of multiple channels to mitigate, to some extent, the severe noise problem prevalent in courtroom speech. By using a novel technique for channel change detection, domain-specific language modeling, and unsupervised channel adaptation we have been able to achieve a word error rate (WER) of 36% with an acoustic model trained on 150 hours of broadcast news data. We also report on our preliminary acoustic modeling experiments with the "legal" transcripts provided with 120 hours of courtroom speech training data.

Read the paper · More papers on PaperTik