Language effects on GPT-based machine learning classifications

Golnoosh Babaei, Sharareh Nosrati, Sepideh Safari · Statistics · 2025

This study investigates how the language used in prompts affects the classification performance of ChatGPT in a credit risk prediction task. Using a structured dataset of loan applications, we evaluated ChatGPT's ability to classify loans as ‘accepted’ or ‘rejected’ in three languages: English, French, and Italian. The model was tested under various few-shot learning settings, using 0, 50, 100, and 150 training examples. Results showed that English and French consistently outperformed Italian across all key metrics, including accuracy, precision, recall, and F1 score. Performance improved significantly with 50 examples but declined at 100 due to increased data dispersion and class overlap, as confirmed by PCA and t-SNE visualizations. At 150 examples, model performance recovered in English and French, while Italian continued to show weaker results. These findings highlight the importance of both data structure and language in prompt-based learning, demonstrating that model performance depends not only on training size but also on the clarity and linguistic context of the input.

Read the paper · More papers on PaperTik