SegMam: Multi-Scale Deformable Cross-Mamba Attention With Pseudo-LiDAR From Multi-Camera Depth for BEV Segmentation
Junhan Kim, HanKyeol Yu, Kyungjae Ahn, Su Yeon Kim, Yeonsik Kang · IEEE Access · 2026
Bird’s Eye View (BEV) semantic segmentation serves as a critical perception technology for spatially coherent understanding of complex urban environments from multi-camera inputs in autonomous driving and intelligent transportation systems. While recent Transformer-based methods have demonstrated excellent results, they face challenges in handling high-resolution BEV representations, particularly for mid- to long-range perception and robust operation under adverse weather conditions. To address these challenges, this paper proposes a novel framework utilizing multi-scale deformable Cross-Mamba attention for multi-camera BEV semantic segmentation. Our approach replaces conventional cross-attention mechanisms with selective state space models, enabling efficient spatial integration during inference while effectively preserving extended contextual dependencies, thereby supporting object and drivable area segmentation. To improve geometric distance estimation without relying on LiDAR or Radar sensors, monocular depth maps extracted from RGB imagery are converted into Pseudo-LiDAR point clouds that serve as the value stream of the Cross-Mamba module, where they are aggregated with multi-scale RGB BEV queries lifted from hierarchical visual features. This dual-stream design facilitates direct interaction between three-dimensional geometric information and multi-scale visual features, yielding consistent improvements under nighttime and rainy conditions and for mid-range vehicle perception, while remaining competitive in inference latency with state-of-the-art camera-only baselines. To foster future research, we provide the monocular depth parameter files, along with our code and trained model checkpoints at https://github.com/KimJunHan/SegMam.git