Topic Recognition Using TED-LIUM Release 3
Jun Yao · 2022
The widespread of mobile digital cameras and phones facilitate creation of multimedia including audio and video records. Categorizing the spoken documents by themes is a key task in processing the huge collection of audio files. Topic Recognition (TR) is a widely used method in categorizing text corpus since the 1980s. Automatic speech recognition (ASR) is a system that converts speech into text. Typically, topic recognition task for spoken documents is consisting of two phases, ASR phase and text topic classification (TTC) phase. In this paper, we will create two models, one is a model based on text and acoustic features by modifying the pretrained LM (Language Model) + triphone HMM (Hidden Markov Model) model using modified example script in Kaldi, another model is a three-layer neural network based on text features, we will combine these two models to complete the topic recognition task. The models will be trained and tested on the open-source TED-LIUM release-3 dataset. The first model is our topic predicting model, it will utilize the official transcripts provided in the dataset. The second model is created to decode text from audios. We then use the first model to predict topics of the decoded text from the second model. Since the second model will propagate error via the decoded text (i.e., transcripts), we would expect the prediction accuracy based on the decoded text to be moderate. The original dataset is lack of topic labels, we will create and publish an open-source topic labels dataset as the golden standard for future use. Our experiment shows that for multiple topics recognition task, random guess prediction accuracy is merely 0.11, test accuracy of the first model using the official transcripts provided by TED-LIUM r3 is 0.40, test accuracy of the first model using the decoded text from the second model is 0.28. As a reference, human prediction accuracy by the author is 0.53. Multiple topic recognition task is a hard problem, this paper provides a starting point based on the TED-LIUM r3 dataset.