Determining the effectiveness of GPT-4.1-mini for multiclass text categorization

Yurii Voloshchuk, Олександр Володимирович Міца · Eastern-European Journal of Enterprise Technologies · 2025

The object of this study is the process of multiclass automatic categorization of user queries using large language models under the conditions of a language transition from English to Ukrainian. The scientific task relates to the fact that most modern large language models (LLMs) are optimized for English while their effectiveness for morphologically complex and low-resource languages, particularly Ukrainian, remains insufficiently studied. In this work, an experimental approach was devised and implemented to evaluate the transferability of the GPT-4.1-mini model from English to Ukrainian in the task to categorize 11,047 user queries spanning nine applied domains. The analysis employed conventional metrics (Recall, Precision, Weighted-F1, Macro-F1) alongside a novel indicator, the Uncertainty/Error Rate (U/E), which captures the proportion of model refusals and “hallucinations.” The findings demonstrate that the highest quality was achieved on the English dataset (Macro-F1 = 69.78%, U/E = 0.05%). When Ukrainian prompts were applied, Macro-F1 decreased to 63.73%; however, the U/E equaled 0%, indicating higher reliability of responses. Using English prompts with Ukrainian-language data preserved nearly the same level of accuracy (Macro-F1 = 69.66%), thereby revealing strong internal translation and generalization mechanisms. The novelty of this study is attributed to the use of a large multidomain parallel corpus, the systematic comparison of prompts in two languages, the application of the state-of-the-art model GPT-4.1-mini, and the introduction of the U/E metric as a reliability criterion. The proposed approach demonstrates the feasibility of applying GPT-4.1-mini to Ukrainian-language information services without additional training, particularly for automatic query routing in financial, medical, legal, and other domains.

Read the paper · More papers on PaperTik