Culturally-Aware Multiclass Bangla Text-to-Image Generation via Fine-Tuned Stable Diffusion Turbo
Susmita Mondal Sristi, Moumita Sen Sharma · 2025
Text-to-Image (TTI) generation has seen remarkable advancements, but most existing models predominantly focus on English, limiting accessibility for languages like Bangla, spoken by over 250 million people. While some research efforts have explored Bangla TTI, these have often been limited to specific classes or domains (e.g., faces or birds). This paper presents the first approach for multiclass Bangla TTI generation, enabling fast and high-quality image synthesis with cultural relevance. The proposed system integrates BanglaBERT Large, a transformer-based Bangla language model for robust text encoding, with Stable Diffusion Turbo, an optimized diffusion model enhanced via Adversarial Diffusion Distillation (ADD), to achieve accelerated generation speed and superior image fidelity. Our model is evaluated on BNature, a dataset containing diverse images depicting Bangladeshi culture and landscapes. Comprehensive evaluations incorporating both qualitative and quantitative measures demonstrate substantial improvements over existing Bangla TTI approaches. Notably, our system achieves a significant enhancement in generation speed while maintaining high-fidelity image quality, as evidenced by a Fréchet Inception Distance (FID) of 117.42, Learned Perceptual Image Patch Similarity (LPIPS) of 0.74, and an Inception Score (IS) of 19.97-outperforming previous models such as Bangla AttnGAN and standard Stable Diffusion in multiclass generation tasks. This research facilitates applications in culturally adaptive content creation, education, and e-commerce, empowering Bangla speakers to visually express their creativity and identity.