Cross-language speech retrieval

Páraic Sheridan, Martin Wechsler, Peter Scháuble · 1997

We present here the realisation of cross-language speech retrieval system which retrieves German speech documents in response to user queries specified as French text. This has been achieved through the integration of two existing modules of the SPIDER information retrieval system, namely the query pseudo-translation module and the speech retrieval module. In the process of integrating these modules for cross-language speech retrieval, we have also made improvements to each in turn. We have progressed our research on aligning comparable corpora, an important resource for training the similarity thesauri used in our query pseudo-translation module, and we have now encorporated an automatic refinement loop into our speech retrieval module, using a probabilistic matching approach that has been proven effective in retrieving error-prone documents. We have also introduced the use of automatic relevance feedback for query expansion and tested this both as a refinement to automatically pseudo...

Read the paper · More papers on PaperTik