Self-Supervised Transformer With Masked Latent Representations for Multimodal Biometric Authentication

Mohamed Benouis, Elisabeth André, Yekta Said Can · 2025

Recent advancements in non-invasive wearable devices have enabled the daily collection of vast amounts of physiological data, driving significant progress in health-related applications. These developments have also enhanced biometric technology, offering secure and user-friendly identity verification methods. However, despite growing interest in physiological signal-based biometric systems, their performance often falls short compared to traditional methods such as facial, vein, and fingerprint recognition. Their effectiveness is further hindered by challenges like poor data quality (e.g., single-modality or short-term data) and label shortage. To address these challenges, we propose a multimodal transformer-based framework trained in a self-supervised manner. Our approach begins by encoding each modality into a lower-dimensional representation, which is then processed through two specialized sub-branches. The first sub-branch utilizes a multi-modal transformer encoder to handle unmasked embeddings, while the second employs random masking and a transformer encoder-decoder to learn latent representations. By minimizing the distance between the outputs of these sub-branches, our method learns robust and generalized multimodal representations for identity-related tasks. We evaluated our framework on two public multimodal physiological datasets, testing its robustness under two main challenges: (1) missing one modality and (2) limited labeled data. Our results demonstrate the effectiveness of the proposed approach in both fine-tuned and frozen configurations, highlighting its superior performance compared to other baselines.

Read the paper · More papers on PaperTik