Breaking the Speed–Accuracy Trade-Off: A Novel Embedding-Based Framework with Coarse Screening-Refined Verification for Zero-Shot Named Entity Recognition
Meng Yang, Shuo Wang, Hexin Yang, Ning Chen · Computers · 2026
Although fine-tuning pretrained language models has brought remarkable progress to zero-shot named entity recognition (NER), current generative approaches still suffer from inherent limitations. Their autoregressive decoding mechanism requires token-by-token generation, resulting in low inference efficiency, while the massive parameter scale leads to high computational and deployment costs. In contrast, span-based methods avoid autoregressive decoding but often face large candidate spaces and severe noise redundancy, which hinder efficient entity localization in long-text scenarios. To overcome these challenges, we propose an efficient Embedding-based NER framework that achieves an optimal balance between performance and efficiency. Specifically, the framework first introduces a lightweight dynamic feature matching module for coarse-grained entity localization, enabling rapid filtering of potential entity regions. Then, a hierarchical progressive entity filtering mechanism is applied for fine-grained recognition and noise suppression. Experimental results demonstrate that the proposed model, which is trained on a single RTX 5090 GPU for only 24 h, attains approximately 90% of the performance of the SOTA GNER-T5 11B model while using only one-seventh of its parameters. Moreover, by eliminating the redundancy of autoregressive decoding, the proposed framework achieves a 17× faster inference speed compared to GNER-T5 11B and significantly surpasses traditional span-based approaches in efficiency.