AI-Driven framework for generating informative summaries from YouTube videos using GPT-3

Nitin Raut, Amit Purushottam Pimpalkar, Nilesh M Shelke, Vishal Tiwari, Tabassum H Khan, Vikrant Chole · 2025

With the exponential growth of video content on platforms like YouTube, extracting key insights efficiently has become a pressing need. The boom of video content in today&s;s age of information overload places increasing demands on end-users to digest more and better-quality information than ever before. However, this challenge calls for a new tactic, artificial intelligence (AI), to automate video summarization. The research represents a pioneering response to the contemporary challenge of shifting through the vast expanse of online video content for meaningful insights. Through the strategic application of AI techniques, this research endeavours to simplify the process of information consumption by autonomously generating concise and informative summaries of videos sourced from platforms such as YouTube. This work presents an AI-driven approach using Large Language Models (LLMs), specifically Generative Pre-training Transformer (GPT) 3, to generate concise and informative summaries from YouTube videos. Our system extracts subtitles, converts them into text, and processes them through GPT-based models to create user-personalized summaries of varying lengths (short, medium, large, and pointwise). Its operational framework involves the extraction of video transcripts, followed by utilizing AI models to condense the content effectively. In cases where transcripts are unavailable, the system seamlessly transitions to downloading the video, extracting its audio, transcribing it, and then synthesizing a summary. It extracts cognitive insights from text-to-text transcribing videos by integrating linguistic subtleties into an environmental context. By incorporating advanced technologies like Natural Language Processing (NLP) and machine learning, the research equips users with the tools to extract pivotal insights from videos swiftly. We evaluate summarization effectiveness using ROUGE and BLEU scores, comparing Google Gemini, Hugging Face, and baseline models. Results demonstrate that LLM-based models outperform traditional baseline methods, achieving a 193 ROUGE-1 score of 0.88 and a BLEU-1 score of 0.85, better than the baseline models. The superior scores highlight enhanced linguistic coherence, contextual understanding, and information retention of LLM models in summarization tasks.

Read the paper · More papers on PaperTik