Towards Key Point Identification (KPI) for Lecture Videos: Approaches and Performance Evaluation

Jiaqi Wang, R. Kwok, Edith C.‐H. Ngai · ACM Transactions on Multimedia Computing Communications and Applications · 2025

To maximize the utility of lecture videos, in today’s fast-paced society with dwindling attention spans, various e-learning technologies are introduced, e.g., non-linear learning, bite-sized learning, and personalized lecture video fragment recommendation. In this article, we conduct a detailed performance study on a key enabler for aforementioned technologies: Lecture Video Fragmentation by Key Point Identification in lecture videos. We begin with a taxonomy of existing methods, which are classified into two categories: boundary-based methods, where the fragmentation is achieved using specific methods depending on the modality, and representation-based methods, where the fragmentation task is formulated as a boundary prediction task based on representations of smaller video chunks. Various configurations of these methods are also examined in detail. To conduct an extensive, comprehensive, and objective comparison study, we address the limitations of existing datasets by introducing a new lecture video fragmentation dataset, MITFLD, without any synthetic videos. We also propose a unified framework kpi , which includes the implementation of datasets, metrics, and compared methods to facilitate the experiments and future research on lecture video fragmentation. The experiments cover different configurations of existing methods on two large datasets (AVLecture and MITFLD). Further experiments are also conducted for ablation studies, such as the effect of feature combinations and the influence of lecture modes. Through the experiments, the representation-based method BiLSTM with self-supervised learning representations is found to exhibit promising performance. Key insights and potential future directions are also discussed.

Read the paper · More papers on PaperTik