Noise Improves Noise: Verification of Pre-Training Effect with Weakly Labeled Data on Social Media NER
Yejin Kim, Juae Kim, Jungyun Seo · 2020
Recently, we are living in a flood of social media data. As the cost of data-annotating is too expensive, there is not much data annotated for natural language processing tasks. The purpose of our experiment is to boost the performance with weakly labeled data without human efforts. In this paper, we present the experimental result following our two purposes. First, we verify the effectiveness of using the pre-training depending on the source domain. Second, we figure out how much improvement appeared on each embedding model. Empirically, the pre-training improves the performance of the target domain regardless of the source domain. Our best model using weakly labeled Wikipedia data outperforms previous models with 49.89% F1-score.