How Readable Is LLM-Generated Code Snippets? A Comparison of ChatGPT, DeepSeek, and Gemini
Giovanna Fernandes, Marcelo de Almeida Maia, Carlos Eduardo C. Dantas · 2025
Developers often search for reusable code snippets on the Web. With the increasing adoption of Large Language Models (LLMs) to support programming tasks and the growing number of available models, this work proposes an evaluation of the readability of 981 code snippets generated by ChatGPT, DeepSeek, and Gemini (327 by each model) by analyzing the warnings detected by static analysis tools such as SonarLint. Additionally, we present a preliminary approach that combines SonarLint recommendations with LLMs to automatically refactor code snippets with the goal of improving code readability. The results show that ChatGPT produces code with fewer readability warnings according to SonarLint. All three LLMs were also able to remove readability warnings in more than 60% of the affected code snippets. However, challenges remain when combining LLMs with static analysis tools, particularly in understanding the context of certain rules and avoiding the removal of relevant code. The insights from this study reveal opportunities for deeper integration between static analysis tools and LLMs.