Multi-modal action recognition in car cabin environments via cross-modal complementary fusion
Zhengwei Fang · DR-NTU (Nanyang Technological University) · 2026
Action recognition within car cabin environments is essential for intelligent driver monitoring and advanced driver assistance systems (ADAS). Although RGB-based approaches have demonstrated strong performance under favorable conditions, they are highly susceptible to degradation in illumination-volatile scenarios such as night-time driving or tunnel environments. To address this challenge, this dissertation presents a multi-modal action recognition framework that integrates RGB and Near-Infrared (NIR) sensing to enhance recognition reliability across varying environmental conditions. Built upon the UniFormerV2 video transformer backbone, we conduct a systematic investigation of multiple fusion paradigms, including early-, late-, and intermediate level strategies. In particular, we propose a Cross-Modal Complementary Fusion (CMCF) mechanism that explicitly accounts for modality-dependent reliability through localized feature enhancement and global adaptive weighting. This design enables selective exploitation of complementary cues while suppressing unreliable modality signals. Extensive experiments on the Drive&Act benchmark demonstrate that multi modal fusion consistently outperforms single-modality baselines. The proposed CMCF framework achieves favorable performance, effectively mitigating modality-specific noise and enhancing feature complementarity in challenging in-cabin environments.