Sarid: Arabic Storyteller Using a Fine-Tuned LLM and Text-to-Image Generation
Maria Alabdulrahman, Renad Khayyat, Kawthar Almowallad, Zahra Alharz · 2024
We propose a novel approach to Arabic story generation by fine-tuning a pre-trained Large Language Model (LLM). Our pipeline includes two stages: text generation and image generation. By fine-tuning the davinci-003 LLM on a dataset of 527 Arabic stories, we tailor the generated stories based on user preferences. For image generation, we utilize the Midjourney model. The results demonstrate the efficacy of fine-tuning a pre-trained image generation model on a limited dataset, as measured by the ROUGE score. Sarid's contributions include addressing the lack of Arabic story generation models, providing a comprehensive dataset of Arabic stories, and integrating text and image generation for a cohesive story generation pipeline.