Speech Retrieval

Ciprian I. Chelba, Timothy J. Hazen, Bhuvana Ramabhadran, Murat Saraçlar · 2011

Parts of this chapter have been previously published in Chelba et al. (2008) [ c○2008 IEEE]. The authors thank IEEE for granting permission to reproduce some paragraphs, tables and figures in this chapter. In this chapter we discuss the retrieval and browsing of spoken audio documents. We focus primarily on the application of document search where a user provides a query and the system returns a set of audio documents that best match the query. The primary technical challenges of speech retrieval lie in the retrieval system’s ability to deal with imperfect speech recognition technology that produces errorful output due to misrecognitions cause by inadequate statistical models or out-of-vocabulary words. This chapter provides an overview of the common tasks and data sets, evaluation metrics, and algorithms most commonly used in this growing area of research. 15.1 Task Description 15.1.1 Spoken Document Retrieval Speech retrieval refers to the task of retrieving the specific pieces of spoken audio data from a large collection that pertain to a query requested by a user. Before discussing methods for speech retrieval, it is important to define the different types of speech retrieval tasks and the methods in which potential solutions to these tasks will be evaluated. When discussing speech information retrieval applications, the basic scenario assumes that a user will provide a query and the system will return a list of rank-ordered documents. The query is generally assumed to be in the form of a string of text-based words (though spoken queries may be used instead of text in some applications). The returned documents are audio files purported by

Read the paper · More papers on PaperTik