Human Face Generation from text Description and Sketch Using GANs and Transformers

Iliass Igmoullan, Abdessamad Elboushaki, Hassan El Bahi, Rachida Hannane · 2025

We present a novel approach for generating realistic human face images from both text descriptions and sketches by combining Transformer-based language models with Generative Adversarial Networks (GANs). In our work, we use a diverse StackGAN framework conditioned on embeddings derived from textual captions and visual sketch inputs. Text semantics are encoded using a BERT transformer, while sketch features are extracted via a pre-trained VGG16 model. The multimodal embeddings are fused and fed into a two-stage GAN to progressively synthesize high-resolution facial images. Three loss functions are used to evaluate our model: the adversarial loss, KL divergence loss, and triplet margin loss, which simultaneously help in promoting realism, diversity, and identity preservation during training. Experimental results on CelebA show that our method provides higher-quality images that better align with both input modalities. The challenges we face include low quality of automated captioning and resolution of sketches. however, the results look hopeful and pave the way for future work in cross-modal face synthesis.

Read the paper · More papers on PaperTik