Challenges in Generating Accurate Text in Images: A Benchmark for Text-to-Image Models on Specialized Content

Zenab Bosheah, Vilmos Bilicki · Applied Sciences · 2025

Rapid advances in text-to-image (T2I) generative models have significantly enhanced visual content creation. However, evaluating these models remains challenging, particularly when assessing their ability to handle complex textual content. The primary aim of this research is to develop a systematic evaluation framework for assessing T2I models’ capabilities in generating specialized content, with emphasis on measuring text rendering accuracy and identifying model limitations across diverse domains. The framework utilizes carefully crafted prompts that require precise formatting, semantic alignment, and compositional reasoning to evaluate model performance. Our evaluation methodology encompasses a comprehensive assessment across many critical domains: mathematical equations, chemical diagrams, programming code, flowcharts, multi-line text, and paragraphs, with each domain tested through specifically designed challenge sets. GPT-4 serves as an automated evaluator, assessing outputs based on key metrics such as text accuracy, readability, formatting consistency, visual design, contextual relevance, and error recovery. Weighted scores generated by GPT-4 are compared with human evaluations to measure alignment and reliability. The results reveal that current T2I models face significant challenges with tasks requiring structural precision and domain-specific accuracy. Notable difficulties include symbol alignment in equations, bond angles in chemical diagrams, syntactical correctness in code, and the generation of coherent multi-line text and paragraphs. This study advances our understanding of fundamental limitations in T2I model architectures while establishing a novel framework for the systematic evaluation of text rendering capabilities. Despite these limitations, the proposed benchmark provides a clear pathway for evaluating and tracking improvements in T2I models, establishing a standardized framework for assessing their ability to generate accurate and reliable structured content for specialized applications.

Read the paper · More papers on PaperTik