CodeNLP at SemEval-2023 Task 2: Data Augmentation for Named Entity Recognition by Combination of Sequence Generation Strategies
Michał Marcińczuk, Wiktor Walentynowicz · 2023
In the article, we present the CodeNLP submission to the SemEval-2023 Task 2: Multi-CoNER II Multilingual Complex Named Entity Recognition.Our approach is based on data augmentation by combining various strategies of sequence generation for training.We show that the extended procedure of fine-tuning a pre-trained language model can bring improvements compared to any single strategy.On the development subsets the improvements were 1.7 pp and 3.1 pp of F-measure, for English and multilingual datasets, respectively.On the test subsets our models achieved 63.51% and 73.22% of Macro F1, respectively.