BERT, RoBERTa, and DistilBERT for Intent Classification: A CLINC-150 Evaluation with QA Applications

Vishist Singh Solanki, Er. Sudhanshu Sharma, Lakshay Sindhu, Leena Bisht, Hariom Singh · 2025

Intent classification is especially significant in NLP, and for contextual question-answering (QA) systems like virtual assistants. Thus, in this paper, we explore three transformer models (BERT, RoBERTa, and DistilBERT) on intent classification using the CLINC-150 dataset (23,850 utterances labeled with one of 151 intent classes) as a benchmark dataset for much of the work. All transformer models used were fine-tuned on the intent classification task for 3 epochs, with evaluation metrics of accuracy, weighted F1score, inference time, number of trainable parameters, and training/validation losses. The results suggested the RoBERTa model (0.888 accuracy, 0.882 F1-score, 0.223 validation loss) was the best performing overall, while DistilBERT was the most efficient model (0.0053 seconds inference time, 67.07 million parameters) with competitive accuracy and one extra decimal (0.889 accuracy). The BERT model acted as a baseline comparable model, obtaining moderate performance (0.875 accuracy, 0.866 F1-score). The loss curves also demonstrated convergence (all models) by epoch 3, with RoBERTa showing improvements in generalization. The results provided more evidence towards RoBERTa's applicability for performance in QA systems and evidence towards DistilBERT for QA systems in low-resourced settings or environments. Future work would assess performance on new datasets across the various QA domain datasets while considering hyper-parameter optimization actions as a way to carry efficiency into performance on multi-domain datasets.

Read the paper · More papers on PaperTik