121Real-Time Fall Detection via Spatio-Temporal Collaborative Attention and Multimodal Feature Fusion Based on Deep Learning
Y.H. Chen · 2025
Existing fall detection methods face three critical challenges in complex dynamic environments: high false alarm rates, insufficient modeling of long-duration action sequences, and privacy risks from data leakage. Traditional unimodal models struggle to capture the multi-stage causal evolution of fall motions comprehensively. To address these limitations, we propose a Spatio-Temporal Collaborative Attention Network (STCANet) that integrates bidirectional spatio-temporal attention mechanisms with optimized multimodal feature fusion, significantly enhancing detection accuracy and efficiency. The architecture employs a dual-path Transformer framework, jointly modeling spatial joint correlations and temporal causal chains through space$\rightarrow$time and time$\rightarrow$space pathways. Additionally, a fusion framework combining kinematic features (centroid velocity/joint angular velocity) with geometric features (silhouette deformation/aspect ratio) is designed to strengthen the model's discriminative power for fall recognition. Furthermore, a lightweight skeleton-based data anonymization model is developed to ensure privacy security while achieving synergistic optimization of both privacy and computational efficiency. Experimental results on the Human3.6M dataset demonstrate a detection precision of 94.2% and recall rate of 93.8%, with false alarms reduced by 74.7% compared to state-of-the-art methods. The model requires only 1.2M parameters and achieves real-time inference at 167 FPS.