ViMPose: Human Pose Estimation Based on Vision Mamba
Bo Yang, Wenyuan Cun, Gang Peng, Jingjing Guo, Chaoshun Li, Jiong Xin Zhao · 2024
Recently, state space models (SSM) based on efficient hardware-aware design, such as the Mamba deep learning model, have demonstrated exceptional efficacy in visual feature recognition functions. However, few studies have explored the potential of this novel architecture for pose estimation tasks. In this paper, we propose ViMPose, a baseline model for human pose estimation based on Vision Mamba. We demonstrate the model's excellent performance in pose estimation from multiple aspects, including model simplicity, inference speed, and lightweight parameters. Specifically, ViMPose employs a new backbone with bidirectional Mamba blocks to extract features from given human instances and uses a lightweight decoder for human pose estimation. Employing the scalable capacity and lightweight nature of Vision Mamba, ViMPose achieves high recognition accuracy with fewer parameters, striking a new balance between real-time efficiency and performance. Furthermore, ViMPose exhibits lower memory usage when processing high-resolution image inputs. Findings of the COCO dataset experiments highlight the ViMPose model's effectiveness and considerable promise for human pose estimation tasks.