Privacy-aware detection of fake identity documents: methodology, benchmark, and improved algorithms (FakeIDet2)
Javier Muñoz-Haro, Rubén Tolosana, Julián Fiérrez, Rubén Vera-Rodríguez, Aythami Morales · Information Fusion · 2025
• Proposal of a privacy-aware framework for fake identity document (ID) detection. • Novel database with over 900K real/fake patches extracted from 2,000 ID images. • Novel fake ID detector based on learnable fusion from patch embeddings. • Release of a standard, reproducible benchmark using physical/synthetic ID attacks. Remote user verification in Internet-based applications is becoming increasingly important nowadays. A popular scenario for it consists of submitting a picture of the user’s Identity Document (ID) to a service platform, authenticating its veracity, and then granting access to the requested digital service. An ID is well-suited to verify the identity of an individual, since it is government issued, unique, and nontransferable. However, with recent advances in Artificial Intelligence (AI), attackers can surpass security measures in IDs and create very realistic physical and synthetic fake IDs. Researchers are now trying to develop methods to detect an ever-growing number of these AI-based fakes that are almost indistinguishable from authentic (bona fide) IDs. In this counterattack effort, researchers are faced with an important challenge: the difficulty in using real data to train fake ID detectors. This real data scarcity for research and development is originated by the sensitive nature of these documents, which are usually kept private by the ID owners (the users) and the ID holders (e.g., government, police, bank, etc.). The present study proposes a new privacy-aware methodology for research and development in fake ID detection that promotes collaboration between ID holders and AI researchers. In practice, the main contributions of our study are: 1) We present and discuss our proposed patch-based methodology to preserve privacy in fake ID detection research. 2) We provide a new public database, FakeIDet2-db, comprising over 900K real/fake ID patches extracted from 2,000 ID images, acquired using different smartphone sensors, illumination and height conditions, etc. In addition, three physical attacks are considered: print, screen, and composite. 3) We present a new privacy-aware fake ID detection method, FakeIDet2, which introduces two novel learnable modules: Patch Embedding Extractor and Patch Embedding Fusion. 4) We release a standard reproducible benchmark that considers physical and synthetic attacks from popular databases in the literature. The results achieved by our proposed FakeIDet2 in detecting very realistic fake IDs under unseen type of attacks are encouraging: 9.56 ± 0.83% and 14.47 ± 2.17% Equal Error Rate(EER) in the very challenging datasets DLC-2021 and KID34K, respectively. FakeIDet2-db and the accompanying benchmark are publicly available https://github.com/BiDAlab/FakeIDet2-db .