Sequential Story Illustration Generation with Fine-Tuned Diffusion Model for Children Stories
P Matan, P Velvizhy · 2025
Illustrating children's stories is essential for nurturing creativity, imagination, and emotional engagement in young readers. High-quality visuals enhance the storytelling experience, yet traditional methods for creating illustrations demand significant artistic expertise, time, and financial resources, making them less accessible and scalable. Existing generative AI models, often lack finegrained adaptation to the domain-specific requirements of children's literature, such as stylized imagery and narrative coherence. To address this gap, this study finetuned the ShuttleAI/Shuttle-3-Diffusion model to generate illustrations that align closely with the narrative context of children's stories. The goal was to develop a system capable of producing stylistically consistent, high-resolution images from sequential textual inputs, ensuring continuity across storylines. The methodology involved preprocessing and fine-tuning the Norod78/Cartoon-Blip-Captions dataset to match the unique visual style needed for children's stories. A sequential image generation pipeline was implemented, where individual sentences were fed as prompts to the model, enabling it to maintain thematic and stylistic alignment across multiple illustrations. Training parameters were optimized, and rigorous evaluations were conducted to benchmark performance. The results revealed that the finetuned model achieved significant improvements in Fréchet Inception Distance (FID) and Inception Score (IS), outperforming baseline generative models. Qualitative analyses highlighted the model's ability to generate visually engaging and contextually accurate illustrations. However, challenges persist in handling complex narratives and expanding stylistic flexibility. This work offers substantial implications for the future of automated content creation in children's literature, paving the way for scalable, cost-effective, and personalized storytelling solutions. Future research could focus on leveraging larger, diverse datasets and real-time application scenarios, further advancing the integration of generative AI in creative and educational domains.