MCFNet: Multi-scale Cross Fusion Network for 3D Human Pose Estimation
Dazhong Wang, Rui Zhou Zhenzhen Liu, Pengfei Yi, Jing Dong, Dongsheng Zhou · 2024
3D human pose estimation has attracted increasing attention due to it frequently used in fields such as human-computer interaction. Due to the structure of the human skeleton is a kind of undirected topology, Graph Convolutional Network(GCN) is a suitable method to study 3D human pose estimation. However, traditional GCNs overlook the distinctions between joints and the implicit connections among nonadjacent joints. Additionally, the topological structure of the human body leads to significant error accumulation in pose estimation, which make the joints have higher error levels in highly flexible regions. We present Multi-scale Cross Fusion Network(MCFNet), a network that sequentially learns the human skeleton structure from joint, limb, and body semantic levels by segmenting original features. The MCFNet can combine the High-order MGCN(Modulated Graph Convolutional Network) module for joint scale capture and the Conv-SE module for limb scale feature fusion, thus achieving excellent joint feature extraction and reduction of the error accumulation. Experimental results on datasets like Human3.6M demonstrate that our proposed method surpasses the performance of state-of-the-art networks and achieves the highest accuracy.