Towards speech-driven lesson summary generation

Phillip Blunt, Bertram Haskins · Procedia Computer Science · 2025

This study explores the development and evaluation of a speech-driven lesson summary generation prototype aimed at enhancing educational accessibility and efficiency through Automatic Speech Recognition (ASR) technology. The prototype addresses the crucial issue of noise interference in educational settings by using Large Vocabulary Continuous Speech Recognition (LVCSR) systems to transcribe and summarise classroom lessons using both Google Cloud Speech-to-Text and CMU Sphinx. To increase ASR accuracy in noisy classrooms, the project looks into integrating feature extraction, noise management strategies, and machine learning innovations. The performance of the prototype is carefully evaluated in several test cases to compare the efficacy of speech enhancement methods and the overall robustness of the system against noise. The results show that Google Cloud Speech-to-Text regularly performs better than CMU Sphinx, including in noisy conditions. This suggests that by ofering accurate transcription services, such technologies could significantly improve educational procedures. The study makes the case for more investigation into the practical use of such prototypes in various educational contexts in order to obtain in-depth feedback and improve system functionality.

Read the paper · More papers on PaperTik