Lightweight Face Anti-spoofing Network Based on Multi-modal Feature Fusion

Yue Zhang, Hongjiao Li · 2025

Face anti-spoofing approaches have been receiving increasing attention from both academia and industry. However, the shift of face recognition technology to mobile and embedded devices driven by the growth of the Internet of Things highlights a notable trend. While lightweight face anti-spoofing models are urgently needed to ensure the reliability and security of facial recognition systems, unimodal information sources remain too homogeneous to extract effective features, resulting in low accuracy when distinguishing between real and fake faces. With the application of deep learning to face anti-spoofing (FAS), a large number of multimodal methods have been proven to be more effective than unimodal methods. In this work, we propose a lightweight FAS model with multimodal input data to address this problem. Specifically, the model processes patch-level images from multimodal inputs (YCbCr, Depth, and IR) through separate branches and extracts features using the lightweight Mobilenetv2 network. Finally, an attention-based feature fusion module is designed to effectively fuse the features from each branch, enabling accurate classification of real and fake faces. Extensive comparative experiments demonstrate that the proposed method significantly reduces parameters while maintaining high accuracy. The accuracy on the multimodal dataset (CASIA SURF) is 98.1269% (TPR@FPR10e-3) and 0.3546% ACER. furthermore, the backbone network has only 2.1M parameters and 80.5M FLOP.

Read the paper · More papers on PaperTik