FedBayesMamba: Uncertainty-aware federated learning for multimodal and audio–visual sequential modeling with selective state space models

Xianxun Zhu, E Xiaosong, Michele Nappi, Imad Rida, Hui CHEN · Pattern Recognition · 2026

Federated learning has emerged as an effective paradigm for training machine learning models across distributed clients without sharing raw data. In many real-world applications, sequential data are inherently multimodal, involving heterogeneous streams such as audio, visual, and temporal signals. However, most existing federated approaches rely on deterministic neural networks, which often struggle to capture predictive uncertainty under heterogeneous data distributions, cross-modal inconsistencies, and dynamic client participation. In this paper, we propose FedBayesMamba , a Bayesian federated learning framework for multimodal sequential data modeling based on selective state space models. The proposed approach introduces Bayesian parameterization into the Mamba architecture to enable uncertainty-aware sequence modeling while preserving the computational efficiency of state space models. To effectively integrate uncertainty across distributed clients, we further develop a posterior aggregation strategy that combines client-level posterior distributions in a principled probabilistic manner. Extensive experiments on multiple benchmark datasets demonstrate that the proposed framework achieves competitive predictive performance and improved uncertainty estimation under Non-IID federated settings. The results also indicate that FedBayesMamba exhibits strong robustness and stability in challenging federated scenarios. These findings highlight the potential of combining Bayesian learning with state space models for multimodal temporal modeling, particularly in audio-visual perception and cross-modal sequence understanding tasks.

Read the paper · More papers on PaperTik