Feature Fusion and Multi-Stream CNNs for ScaleAdaptive Multimodal Sign Language Recognition
Navya Singla, Manas Taneja, Neha Goyal, Rajni Jindal · 2023
Sign language can be defined as a visual-spatial language used by deaf and hard-of-hearing individuals for everyday communication. However, the use of sign language is often limited by the lack of technology capable of recognizing and interpreting it. Our research aims to create a method for recognizing static sign language, a crucial aspect of sign language communication. We propose using multi-stream Convolutional Neural Networks (CNNs), incorporating local scale awareness and multimodality, to recognize static Hand signs more effectively. The local scale awareness aspect is achieved through spatial pyramidal pooling (SPP), which allows CNN to obtain and merge features from different scales. The multimodal aspect is achieved using Color (RGB), depth information, and extracted image features as these modalities provide complementary information for recognizing SL. We also introduce a method for feature extraction based on the fusion of extracted features from the Gabor Filter and the Local Binary Pattern (LBP) method. The multi-stream CNN architecture is used to effectively combine the information from multiple sources, whereas the proposed feature fusion method provides extensive details about image features, resulting in improved recognition accuracy. We achieved a high accuracy rate in recognizing individual signs and our approach has the potential to improve accessibility for hard-of-hearing communities.