Performance Evaluation of Text Embedding Models for Ambiguity Classification in Indonesian News Corpus: A Comparative Study of TF-IDF, Word2Vec, FastText BERT, and GPT

Sutriawan Sutriawan, Supriadi Rustad, Guruh Fajar Shidik, Pujiono Pujiono · Ingénierie des systèmes d information · 2025

Ambiguity in sentence classification is a major challenge in natural language processing (NLP), as it requires a deep understanding of complex semantic contexts.Although various text embedding models have been applied to text classification tasks, comprehensive evaluations of their effectiveness in detecting ambiguous sentences, particularly in Indonesian news corpora, remain limited.This study addresses that gap by comparing the performance of five text embedding models TF-IDF, Word2Vec, FastText, BERT, and GPT combined with five binary classification algorithms: Logistic Regression, Random Forest, bagging, Multinomial Naive Bayes, and Gaussian Naive Bayes.The dataset was derived from the XL-Sum Indonesian news corpus, with sentences automatically labeled as ambiguous or unambiguous using the Claude 3.5 language model.Experimental results show that the combination of Gaussian Naive Bayes with GPT embeddings achieved the best performance in ambiguous sentence classification, with a recall of 71% and an F1score of 60%.Meanwhile, the combination of TF-IDF with bagging yielded the highest accuracy of 83% for unambiguous sentence classification.These findings highlight the critical role of selecting appropriate embedding and classification models to enhance accuracy in semantically ambiguous sentence classification for the Indonesian language.

Read the paper · More papers on PaperTik