MVMSFN: A Multi-View and Multi-Scale Fusion Network for Online Detection of Heterogeneous Gestures
Hao Long, Song Wang, Hao Hu, Yanru Wang, Hao Xu, Hesong Wang · 2024
Gesture is considered as an important approach to implement natural human-machine interaction, but the existing online gesture detection methods still suffer from the problem of unsatisfactory detection rate and false trigger rate that may lead to failed interactions. In this paper, we propose an improved Multi-View and Multi-Task(MVMT) model called Multi-View and Multi-Scale Fusion Network(MVMSFN) for the problem of inaccurate gesture segmentation from the background data. Firstly, for the adaptive feature fusion, we construct Multi-View Fusion Model(MVFM), which introduces spatial attention weight matrix to learn to enhance information interaction among multi-view branches, and Multi-Scale Fusion Model(MSFM), which extracts correlations among features from different layers with different scales to mitigate the effect of inconsistent gesture lengths. Then we design Decoupled Gesture Progress Module (DGPM), which sets up auxiliary tasks that enable the network to predict the execution progression of static and dynamic gestures respectively, to address the network performance degradation problem caused by applying uniform gesture modeling in the heterogeneous gesture environment(including dynamic and static gestures) where there is great difference in representation between the two main gesture categories. The experimental results on the latest benchmark show that MVMSFN achieves the best performance in metrics, in which the false positive score outperforms the state-of-the-art approach(OO-dMVMT). In summary, MVMSFN provides higher accuracy of gesture segmentation and lower false trigger rate while maintaining relatively low computational time for real-time applications. The code is available at https://github.com/miixcc/MVMSFN-Gesture.