Integrated transcription and identification of named entities in broadcast speech

Steve J. Renals, Yoshihiko Gotoh · 1999

This paper presents an approach to integrating functions for both transcription and named entity (NE) identification into a large vocabulary continuous speech recognition system. It builds on NE tagged language modelling approach, which was recently applied for development of the statistical NE annotation system. We also present results for proper name identification experiment using the Hub-4 evaluation data. 1. INTRODUCTION The accurate identification of proper names and other named entities (NEs) has a useful role to play in spoken language processing, as component in speech understanding systems, and as a way of structuring recogniser output (e.g., as a cue to punctuation and capitalisation). Recently trainable hidden Markov model systems for NE identification have been reported with a precision /recall performance similar to that of the best grammar based systems and only a small amount of degradation when applied to speech recogniser output [1, 2]. We have previously presented...

Read the paper · More papers on PaperTik