Multi-stream CNNs with Orientation-Magnitude Response Maps and Weighted Inception Module for Human Action Recognition
Fatemeh Khezerlou, Aryaz Baradarani, Mohammad Ali Balafar, Roman Gr. Maev · 2023
In this paper, a multi-stream convolutional neural network (CNN) is proposed, which is integrated with multi-modal data acquired from a video camera, Kinect and wearable inertial sensors. We introduce the orientation-Magnitude Response (OMR) maps exploited from optical flow to represent motion information in a single 2D image. A weighted inception module is proposed to extract the effectiveness of adjacent and nonadjacent neighbors of OMR maps involved with different weights in extracting multi-scale spatial relationships of the current pixel. The local and global features of the 3D pose acquired from the Kinect sensor are utilized to locate person and ordering joint importance. For 3D pose features, a CNN with a channel attention module has been designed. The inertial signals are also fed into the separate CNN to learn high-level representation. The proposed model is evaluated on UTD-MHAD and MSR-Daily activity datasets to evaluate its effectiveness and efficiency.