A stream-based audio segmentation, classification and clustering pre-processing system for broadcast news using ANN models

Hugo Meinedo, João Paulo da Silva Neto · 2005

This paper describes our work on the development of a low latency stream-based audio pre-processing system for broadcast news using model-based techniques. It performs speech/nonspeech classification, speaker segmentation, speaker clustering, gender and background conditions classification. As a way to increase the modelling accuracy our algorithms make extensive use of Artificial Neural Networks (ANN) thus avoiding the rough assumptions normally made about the audio signal distribution. Experiments were conducted on the COST278 multilingual TV broadcast news database and compared with current state of the art algorithms using standard evaluation tools. Additionally we investigated the impact of automatic audio preprocessing system within the recognition using a large broadcast news test database for the European Portuguese. These testsshow a smalldegradation in recognition performance when compared with hand labelled audio segmentation. Our system is part of a prototype close-captioning system that is daily processing the main news show of two Portuguese Broadcasters.

Read the paper · More papers on PaperTik