BAPO: A Large-Scale Multimodal Corpus for Ball Possession Prediction in American Football Games
Ziruo Yi, Eduardo Blanco, Heng Fan, Mark V. Albert · 2022
We present a new task for multimodal information extraction: identify the players or teams that possess the ball in each play of an American football game. We also introduce BAPO, a large-scale corpus that consists of 100 games totaling around 200 hours of video broadcasts along with ball possession information for 15,132 plays. This corpus is rich and diverse because it involves a great number of players and different scenarios. BAPO poses a new challenge to build multimodal models that take into account language (what the broadcasters say), audio (broadcasters' tone, sounds from the audience, etc.) and vision (frames from the video broadcast). We further propose a baseline model, BAPOTer, and conduct comprehensive experiments. Results demonstrate that language is key to solve this task and that leveraging all three modalities is beneficial. The corpus is available at https://github.com/Isabella1118/BAPO.