Text-Guided Sports Highlights: A CLIP-Based Framework for Automatic Video Summarization

Marcos Rodrigo, Carlos Cuevas, Narciso N. Garcia · IEEE Access · 2025

We present a text-guided framework for automatic sports video summarization that leverages Contrastive Language–Image Pretraining to classify frames as highlight or non-highlight from natural-language descriptions. The pipeline generates multiple highlight/non-highlight sentence pairs with a large language model, removes weak candidates via distribution- and area-based filters, averages the remaining framewise predictions, and refines events through lightweight post-processing. We further introduceSport-CLIP, a multi-sport benchmark (diving, long jump, pole vault, tumbling) designed to probe robustness across distinct motion patterns and highlight durations. Comprehensive experiments on MATDAT (tricking) and SportCLIP showthat the method achieves consistent, high-quality summarization: an overall average F-score of 84.08% across sports using a single parameter set, with 91.05% and 87.98% on pole vault and long jump, respectively. Parameter sweeps indicate stable performance over a broad range of smoothing windows and filtering thresholds, underscoring resilience to moderate setting changes. Because the approach is training-free and prompt-driven, it integrates readily into consumer-electronics pipelines (e.g., cloud or mobile media platforms) and adapts to new sports and user intents by editing text prompts rather than retraining models. Code and data are publicly available at www.gti.ssr.upm.es/data.

Read the paper · More papers on PaperTik