A Multimodal Hierarchical Variational Autoencoder for Saliency Detection

Zhengyang Yu, Jing Zhang, Nick M. Barnes · 2023

Existing multimodal Salient Object Detection (SOD) methods do not generalize well for more complex and scalable multimodal learning scenarios. In this paper, we propose a multimodal hierarchical variational auto-encoder for generalized multimodal SOD. By introducing joint inference factorization methods in multimodal VAEs, our model is scalable to partially missing modality data. A latent hierarchy is proposed which enhances expressiveness in latent space and allows multi-level interaction between features across modalities. By further exploring the latent hierarchy, we provide intuitive uncertainty visualizations and observe that the main source of uncertainty in SOD derives from the lower-level features. Based on this, we propose a simple yet effective sampling-based confidence estimation method, that brings robustness when encountering untrustworthy modality data for inference. Extensive experimental analysis illustrate that our model can satisfy crucial properties that make a desirable multimodal SOD framework.

Read the paper · More papers on PaperTik