Evaluating Large Language Models using Arabic Prompts to Generate Python Codes
Nassir Jabir Al-khafaji, Basit Khalaf Majeed · 2024
Currently, the popularity of large language models (LLMs) for instance, ChatGPT from OpenAI and Gemini from Google is increasing greatly in our lives, due to their unparalleled performance in various applications. These models play a vital role in both academic and industrial fields. With this popularity, evaluating these models has become extremely important, especially when using the Arabic language. Despite the increasing popularity and performance of AI, there have been no empirical studies evaluating the use of LLMs for Arabic prompts in the field of code generation. However, the code generation in LLM can be strongly influenced by the choice of prompt. Evaluating the LLMs by Arabic prompts helps us better understand the strengths and weaknesses of these models. Therefore, we highlighted the evaluation of the most popular LLM programs (Chatgpt-3.5, ChatGPT-4 and Gemini) when generating Python codes based on Arabic prompts. In this study we employed CodeBLUE score and Flake8 as a metric to evaluate the LLMs capabilities for code generation via Arabic prompts. Our results indicate significant differences in performance across different LLMs and prompts levels. This study lays the foundation for further research into LLM capabilities based on Arabic prompts and suggests practical implications for using LLM in automated code generation and test-driven development tasks.