Enhancing action recognition by leveraging the hierarchical structure of actions and textual context
Manuel Benavent-Lledó, David Mulero-Pérez, David Ortiz-Pérez, Jose Garcia-Rodriguez, Antonis Argyros · Computer Vision and Image Understanding · 2025
We propose a novel approach to improve action recognition by exploiting the hierarchical organization of actions and by incorporating contextualized textual information, including location and previous actions, to reflect the action’s temporal context . To achieve this, we introduce a transformer architecture tailored for action recognition that employs both visual and textual features. Visual features are obtained from RGB and optical flow data, while text embeddings represent contextual information. Furthermore, we define a joint loss function to simultaneously train the model for both coarse- and fine-grained action recognition, effectively exploiting the hierarchical nature of actions. To demonstrate the effectiveness of our method, we extend the Toyota Smarthome Untrimmed (TSU) dataset by incorporating action hierarchies, resulting in the Hierarchical TSU dataset , a hierarchical dataset designed for monitoring activities of the elderly in home environments. An ablation study assesses the performance impact of different strategies for integrating contextual and hierarchical data. Experimental results demonstrate that the proposed method consistently outperforms SOTA methods on the Hierarchical TSU dataset, Assembly101 and IkeaASM, achieving over a 17% improvement in top-1 accuracy. • We propose a vision-language transformer model that improves action recognition using contextual information and action hierarchies, outperforming state-of-the-art visual-only models on the Hierarchical TSU, Assembly101 and IkeaASM action recognition benchmarks. • We introduce the Hierarchical TSU dataset for contextual action recognition with structured action hierarchies for activities of daily living. • We conduct an extensive ablation study on integrating contextual and hierarchical data to enhance action recognition performance.