Beyond Audio: Enhancing SoccerNet-Echoes with Multimodal Event Extraction Using LLMs

Mehdi Houshmand Sarkhoosh, Sushant Gautam, Cise Midoglu, Thu Nguyen, Jan Held, Anthony Cioppa, Silvio Giancola, Vajira Lasantha Thambawita, Michael A. Riegler, Pål Halvorsen · International Journal of Semantic Computing · 2025

In this paper, we present SoccerNet-Echoes, an extension of the SoccerNet dataset which has been curated by augmenting the 550 games in the original dataset with multilingual audio commentary transcriptions, with a pipeline utilizing OpenAI’s Whisper models for transcription and Google Translate for translation to English. We demonstrate the potential of SoccerNet-Echoes through several applications. Our experiments reveal that incorporating ASR-generated transcripts as a third modality alongside audio and video can improve the performance of multimodal event detection, with our audio–video-text model achieving a top F1-score of 0.7175. We also introduce a novel framework that leverages Large Language Models (LLMs) to extract both predefined, official events, as well as unscripted, unofficial events directly from the commentary. Our evaluation shows that the Gemini-1.5-Pro model effectively identifies official events from text alone, and that LLM-generated game summaries are more descriptive and accurate when using SoccerNet-Echoes compared to using only structured event data. Surprisingly, our experiments also show that feeding powerful LLMs like Gemini-1.5-Pro with visual data may not improve results compared to their text-only counterpart, but rather degrade performance, for which we analyze the potential reasons. By releasing SoccerNet-Echoes, we provide a resource for the scientific community and offer benchmarks that highlight the current capabilities and limitations of ASR and LLM technologies in the domain of multimodal sports analysis.

Read the paper · More papers on PaperTik