Bridging Modalities: Improving Thermal Face Recognition with Low-Quality Cross-Modality Synthesis
Chengxi Dong, Jiajin He, Yunqi Cai · 2024
Current deep learning models require a large amount of data, which is not affordable for tasks with limited data preparation. Take face recognition with thermal images as an example, there are only a few public datasets of thermal faces and the recording devices are far from diverse. Consequently, much less thermal face images are available compared to visible faces, leading to a significant performance gap. Visible-to-thermal cross-modality data augmentation is a vital tool to alleviate the data sparsity problem, however it often requires a strong synthesizer. In this paper, we found that a weak synthesizer, i.e., a CycleGAN in our study is sufficient to conduct effective cross-modality data augmentation. The basic idea is to build a CycleGAN model trained with limited cross-domain data that converts visible face images to thermal face images, so that speaker-related information associated with visible faces can be propagated to thermal faces, leading to improved generalizability. We found that although the synthesized thermal faces are neither perfect nor faithful, the performance of the thermal face recognition model can still be significantly improved.