Automated Video Title Generation for Mobile Learning Resources: A Deep Learning Approach with Educational Context Awareness
正哉 奥宮 · Advances in Mobile Learning Educational Research · 2025
With the rapid growth of mobile learning platforms, short educational videos have emerged as a critical resource for learners. However, manually generating concise and pedagogically meaningful titles for these videos remains a time-consuming challenge. To address this issue, this study proposes a deep learning framework designed for automated video title generation in educational contexts. The framework integrates Convolutional Neural Networks (CNNs), Long Short-Term Memory (LSTM) networks, and natural language processing (NLP) techniques, with explicit awareness of pedagogical relevance. The proposed approach operates in three stages: 1) extracting key frames from input videos using an optimized shot detection algorithm, 2) analyzing these frames with CNN models to derive semantic representations of visual content, and 3) processing the representations through an LSTM network to generate descriptive text. The output is further refined using the TextRank algorithm to ensure conciseness and contextual coherence. Experimental results demonstrate that our framework effectively generates high-quality video titles that are both educationally informative and contextually engaging, outperforming baseline methods in alignment with curriculum standards and learner-centric search intent.