Searching Recorded Speech Based on the Temporal Extent of Topic Labels
Douglas W. Oard, Anton Leuski · 2003
Recorded speech poses unusual challenges for the design of interactive end-user search systems. Automatic speech recognition is sufficiently accurate to support the automated components of interactive search systems in some applications. Recognizing useful recordings among those nominated by the system is difficult, however, because listening to audio is time consuming and because recognition errors and speech disfluencies make it difficult to mitigate this time factor by skimming automatic transcripts. Support for the browsing process based on supervised learning for automatic classification has shown promise, however, and a segment-then-label framework has emerged as the dominant paradigm for applying that technique to news broadcasts. This paper argues for a more general framework, which we call activation matrices, that provide a flexible representation for the mapping between labels and time. Three approaches to the generation of activation matrices could be generated are briefly described, with the main focus of the paper being the use of activation matrices to support search and selection in interactive systems.