Leveraging multimodal content for podcast summarization
Lorenzo Vaiani, Moreno La Quatra, Luca Cagliero, Paolo Garza · Proceedings of the 37th ACM/SIGAPP Symposium on Applied Computing · 2022
Podcasts are becoming an increasingly popular way to share streaming audio content. Podcast summarization aims at improving the accessibility of podcast content by automatically generating a concise summary consisting of text/audio extracts. Existing approaches either extract short audio snippets by means of speech summarization techniques or produce abstractive summaries of the speech transcription disregarding the podcast audio. To leverage the multimodal information hidden in podcast episodes we propose an end-to-end architecture for extractive summarization that encodes both acoustic and textual contents. It learns how to attend relevant multimodal features using an ad hoc, deep feature fusion network.