A new multimodal deep-learning model to video scene segmentation

Tiago Henrique Trojahn, Rodrigo Mitsuo Kishi, Rudinei Goularte · 2018

The recent development of deep learning techniques, like convolutional networks, shed a new light over the video (story) scene segmentation problem, bringing the potential to outperform state-of-the-art non-deep learning multimodal approaches. However, one important aspect of the multimodality still needs investigation in the context of deep learning: the multimodal fusion. Often, features are directly fed to a network, which may be an inadequate approach to perform the underlying multimodal fusion.

Read the paper · More papers on PaperTik