Multimodal Machine Learning Framework for Fall Risk Assessment Utilizing Sensor, Video, and Contextual Insights
Jason Yang, Ruoqi Li · 2025
Falls represent a significant cause of injury and disability among the elderly, yet time and cost constraints are major barriers to effective fall risk screening. Utilizing machine learning models into fall risk assessment offers an approach to increase accuracy, efficiency, and accessibility by providing objective and automated solutions. This paper proposes the first multimodal model developed that simultaneously analyzes video, sensor, and contextual data based on the Timed Up and Go (TUG) test-an optimized and efficient fall risk assessment framework that captures real-world scenarios, such as getting up from a chair and turning around. Through conducting testing, trunk swing, step width variability, and arm separation emerged as critical gait parameters to extract for distinguishing fallers from non-fallers. Also, velocity data extracted by the sensor proved to correlate more with a participant's fall risk than acceleration data. Important fall-related contextual data, such as age, sex, height, weight, and BMI, also played a substantial role in increasing predictive accuracy. Compared to previous approaches, this paper fully combines clinical data from multiple reliable sources into one and offers a solid perspective for better decision-making and predictive power. The proposed multimodal model, a four-hidden-layer neural network, had an accuracy, precision, recall, and F1 scores of$\mathbf{9 2 \%, 9 3 \%, 9 2 \%, ~ a n d ~} \mathbf{9 1 \%}$, respectively. This outperformed two independent models' performance that were trained on publicly available datasets in addition to some other pre-existing models by a wide margin. The results demonstrate the potential of the multimodal machine learning framework as a rather accurate, efficient, and readily accessible alternative in estimating the risk of fall.