Resource-Efficient Named Entity Recognition: Evaluating Classical, BiLSTM-CRF, and Distilled Transformer Architectures

K Deepthi Reddy, V.D.S Krishna, Priyanka Priyanka, M. R. Archana, V. N. V. L. S. Swathi · 2025

Named Entity Recognition (NER) is a critical task in the field of Natural Language Processing (NLP) which is aimed at recognizing and giving categories to the entities such as names, dates, and locations from raw text. While deep learning approaches, especially transformer models like BERT and RoBERTa, have set new benchmarks in performance, they often come with substantial computational costs, making them less feasible for real-time or resource-constrained environments. This research investigates both traditional and deep learning-based NER models, aiming to strike a balance between accuracy and computational efficiency. In particular, we compare conventional machine learning techniques, such as Conditional Random Fields (CRF) and Support Vector Machines (SVM), with more sophisticated neural network architectures, including Bidirectional Long Short-Term Memory (BiLSTM) combined with CRF, and compact transformer models like DistilBERT and ALBERT. Our evaluation utilizes the CoNLL-2003 dataset, a widely recognized benchmark featuring various entity categories across diverse contexts. To optimize the performance of classical models, we incorporate FastText embeddings for enhanced feature representation, while fine-tuning transformer models for domain-specific adaptations. The results indicate that BiLSTM-CRF with FastText embeddings achieves an F1-score of 88.5% on the CoNLL-2003 dataset, offering a promising solution for efficiency-oriented applications. On the other hand, DistilbERT shows a commendable F1-score of 91.2 % on CoNLL-2003 and 87.4% on OntoNotes 5.0, effectively balancing computational efficiency and high performance. Additionally, ALBERT, through its parameter-sharing strategy, attains an F1-score of 90.1 % on CoNLL-2003, with notably reduced memory consumption, further solidifying its potential for largescale, real-time NER tasks. This study highlights the trade-offs involved in selecting the right NER model for diverse practical scenarios, where both speed and accuracy are paramount.

Read the paper · More papers on PaperTik