End-to-end let's play commentary generation using multi-modal video representations
Chengxi Li, Sagar Gandhi, Brent E. Harrison · 2019
In this paper, we explore how multi-modal video representations can be applied in an end-to-end fashion for automatically generating game commentary based on Let's Play videos using deep learning. We introduce a comprehensive pipeline that involves directly taking videos from YouTube and then using a sequence-to-sequence strategy to learn how to generate appropriate commentary. We evaluate our framework using Let's Play commentaries for the game Getting Over It with Bennet Foddy. To test the quality of the commentary generation, we apply perplexity to evaluate our language models using different input video representations to highlight different aspects of gameplay that might influence commentary.