Brown at CL-SR’07: Retrieving conversational speech in English and Czech

Matthew Lease, Eugene Charniak · 2007

Brown’s entry to the Cross-Language Speech Retrieval (CL-SR) track at the 2007 Cross Language Evaluation Forum (CLEF) 1 was based on the language model (LM) paradigm for retrieval [17]. For English, our system introduced two minor enhancements to the basic unigram: we extended Dirichlet smoothing (popular with unigram modeling) to bigrams, and we smoothed the collection LM to compensate for the small collection size. For Czech, time-constraints restricted us to using a basic unigram model, though we did apply Czech-specific stemming. While our English system performed well in the evaluation and showed the utility of our enhancements, several aspects of it were rushed and need to be addressed in future work. Our Czech system did not perform competitively but did provide us with a useful first experience in non-English retrieval. 1

Read the paper · More papers on PaperTik