Optimizing Technical Telecommunication Assessment Evaluation with Large Language Models

Fatkhul Chorib, Gelar Budiman, Khilda Afifah · 2025

This research evaluates the effectiveness of Large Language Models (LLMs) in automated essay evaluation, focusing on ChatGPT 4.0, GPT-2, BERT, and RoBERTa. The models were fine-tuned using the Stanford Question Answering Dataset (SQuAD) and a custom essay scoring dataset, with performance assessed based on Mean Squared Error (MSE), Normalized Mean Squared Error (NMSE), and the R2 Score. Among these, ChatGPT 4.0 demonstrated superior performance, achieving an MSE of 4.59, an NMSE of 0.50, and an R2 Score of 0.72. The findings confirm the model's strong alignment with human scoring standards, providing both accuracy and consistency. This study offers a robust framework to enhance automated essay evaluation, reducing reliance on manual grading and promoting fairness. Future studies should focus on addressing limitations such as dataset diversity and computational efficiency to further improve model generalization and scalability. Overall, this research highlights the transformative potential of advanced LLMs in establishing scalable, accurate, and equitable educational assessment systems.

Read the paper · More papers on PaperTik