Identifying Chinese Opinion Expressions with Extremely-Noisy Crowdsourcing Annotations

Xin Zhang, Guangwei Xu, Yueheng Sun, Meishan Zhang, Xiaobin Wang, Min Zhang · Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) · 2022

Recent works of opinion expression identification (OEI) rely heavily on the quality and scale of the manually-constructed training corpus, which could be extremely difficult to satisfy.Crowdsourcing is one practical solution for this problem, aiming to create a large-scale but quality-unguaranteed corpus.In this work, we investigate Chinese OEI with extremelynoisy crowdsourcing annotations, constructing a dataset at a very low cost.Following Zhang et al. (2021), we train the annotator-adapter model by regarding all annotations as goldstandard in terms of crowd annotators, and test the model by using a synthetic expert, which is a mixture of all annotators.As this annotatormixture for testing is never modeled explicitly in the training phase, we propose to generate synthetic training samples by a pertinent mixup strategy to make the training and testing highly consistent.The simulation experiments on our constructed dataset show that crowdsourcing is highly promising for OEI, and our proposed annotator-mixup can further enhance the crowdsourcing modeling.

Read the paper · More papers on PaperTik