Evaluating the GPT-3.5 and GPT-4 Large Language Models for Zero-Shot Classification of South African Violent Event Data
Eduan Kotzé, Burgert Senekal · 2024
This study aims to empirically evaluate the performance of GPT -3.5 and GPT-4 language models in text classification tasks, specifically in the classification of violent events in a South African context. The study also compares the findings against a baseline model based on Support Vector Machines (SVM). A dataset of 7280 WhatsApp messages, in both English and Afrikaans, referring to violent events in South Africa was used for testing. The findings indicate that GPT-4 outperforms GPT -3.5 in most evaluation metrics for both English and Afrikaans input text. GPT-4 demonstrates higher Fl scores, indicating a better balance between precision and recall, leading to a more accurate classification of text in both languages. GPT-4 also shows higher accuracy, precision, and recall scores compared to GPT-3.5. However, both GPT models performed worse than the SVM baseline model in terms of accuracy, precision, and Fl scores, although GPT-4 outperformed the SVM baseline model in terms of recall. Overall, this study provides empirical evidence of the improved performance of GPT-4 over GPT -3.5 in text classification tasks, specifically in the context of violent events in South Africa. However, the study also shows the shortcomings of both GPT models in zero-shot text classification tasks. In addition, both models performed worse on Afrikaans text, which shows that further training is required for low-resource languages.