Active learning with semi-automatic annotation for extractive speech summarization

Justin Jian Zhang, Pascale Fung · ACM Transactions on Speech and Language Processing · 2012

We propose using active learning for extractive speech summarization in order to reduce human effort in generating reference summaries. Active learning chooses a selective set of samples to be labeled. We propose a combination of informativeness and representativeness criteria for selection. We further propose a semi-automatic method to generate reference summaries for presentation speech by using Relaxed Dynamic Time Warping (RDTW) alignment between presentation speech and its accompanied slides. Our summarization results show that the amount of labeled data needed for a given summarization accuracy can be reduced by more than 23% compared to random sampling.

Read the paper · More papers on PaperTik