Language Anxiety Detection in English Texts Using BERT and SVM
Rahmat Sigit Hidayat, Kemas Muslim Lhaksmana, Iis Kurnia Nurhayati · 2025
Language anxiety—a mental state involving fear, self-doubt, or caution while communicating in a foreign language—can intrude on learning and limit one's ability to effectively communicate in academic and professional environments. In this research, whether the possibility of automatically identifying such anxiety by machine learning techniques on English written text is achieved is explored. The experiment compares two approaches: a Bidirectional Encoder Representations from Transformers (BERT) and a traditional Support Vector Machine (SVM) classifier. Obtained 323 written responses of students in English language classes at Telkom University. Such responses were obtained from open-ended questionnaire items and labeled as "anxiety" and "non-anxiety" posts. Data preparation entailed cleaning, contraction expansion, case folding, lemmatization, and elimination of unnecessary words and retention of necessary words with regard to mental health. Since the data was skewed—with far more posts related to anxiety—oversampling techniques were employed to balance the two classes while training. The deep learning architecture of BERT was fine-tuned with Optuna for hyperparameter adjustment like learning rate and batch size, whereas the SVM model employed TF-IDF for feature extraction and GridSearchCV for parameter search. Both models were tested using accuracy, precision, recall, and F1-score. Outcomes showed that BERT in general performed better, with 85 percent accuracy and improved on detecting anxious content. The SVM model, while less accurate at 78 percent , was more stable in certain cases and consumed less computational resources. These outcomes demonstrate that transformer models deliver strong performance in affectively intensive tasks like this, but that simpler models can still maintain pragmatic usability, especially in resource-constrained environments. Other improvement would be from refining the training data, merging models, or entering multilingual environments.