Analyse et prédiction du comportement humain dans des séquences temporelles non contrôlées

Benjamin Szczapa · theses.fr (ABES) · 2022

Human behavior understanding has been an important research topic in the past decades. Indeed, the development of machines that work and help humans in their daily lives has never been more important than it is today. It is important to develop appropriate methods to better understand human behavior. In this sense, recent breakthroughs in computer science and computer vision have made the development of such methods possible. Understanding body and facial movements can be done by detecting 2D or 3D landmarks from different sources like a video or the feed of a camera. Performing this acquisition process over time makes it possible to construct temporal sequences of landmark configurations that can be processed to address different tasks, including the recognition of actions and emotions. However, deformations can be observed during the analysis, due to view variations, inaccurate landmark detection or tracking, especially in uncontrolled situations. In this thesis, we propose two space-time approaches of body joint and facial landmark sequences, while tackling different problems in understanding of human behavior. Firstly, we propose a representation based on trajectories of Gram matrices computed from body joints or facial landmarks. The Gram matrices representation defines positive semi-definite matrices of fixed rank that lay on a non-linear Riemannian manifold, where traditional computations and machine learning techniques could not be applied. To overcome this issue, the trajectories defined by sequences of Gram matrices on the manifold of SemiPositive definite matrices are analyzed by considering metric properties induced by the Riemannian geometry of the manifold. The proposed approach was evaluated in several applications related to body movements and action recognition from skeletons using 2D and 3D body joints as well as facial expression analysis to estimate the level of pain directly from 2D facial landmarks. Secondly, we propose a neural network architecture based on a Graph Convolutional Network and a Transformer model that combines the computation of attention at spatial and temporal level of 2D facial landmark sequences. We evaluate this second approach in the estimation of pain level at sequence level. The results obtained by applying the two proposed approaches on widely used data are competitive with respect to recent state-of-the-art methods proposed in the literature.

Read the paper · More papers on PaperTik