A liveness detection protocol based on deep visual-linguistic alignment
Viet-Trung Tran, Van-Sang Tran, Xuan-Bang Nguyen, The-Trung Tran · 2022
Face anti-spoofing has become increasingly critical due to the widespread deployment of face recognition technology. Current approaches mostly focus on presentation attacks, where they rely on textual and spatio-temporal features in captured facial videos. However, in an environment where end-users manage their own devices, attackers can cheat by using virtual camera sensors and easily bypass sophisticated approaches for presentation attacks. In this paper, we propose a novel liveness detection protocol where users are required to read a random-generated sequence of words. Our proposed prediction model, LipBERT, a deep visual-linguistic alignment, is trained to detect if the captured facial stream conforms to the valid textual sequence. For the experiments, we introduce VNFaceTalking1, an extensive dataset of 188,561 samples (around 130 hours in total). Each sample is at most 3 seconds video of frontal face talking Vietnamese. Experiments on the VNFaceTalking dataset demonstrate promising results.1https://github.com/tranvansanghust/VNFaceTalking