A Study of Errors in the Output of Large Language Models for Domain-Specific Few-Shot Named Entity Recognition

Elena Volkanovska · LDV-Forum/Journal for language technology and computational linguistics · 2025

This paper proposes an error classification framework for a comprehensive analysis of the output that large language models (LLMs) generate in a few-shot named entity recognition (NER) task in a specialised domain. The framework should be seen as an exploratory analysis complementary to established performance metrics for NER classifiers, such as F1 score, as it accounts for outcomes possible in a few-shot, LLMbased NER task. By categorising and assessing incorrect named entity predictions quantitatively, the paper shows how the proposed error classification could support a deeper cross-model and cross-prompt performance comparison, alongside a roadmap for a guided qualitative error analysis.

Read the paper · More papers on PaperTik