Multi-task Learning Model for Detecting Internet Slang Words with Two-Layer Annotation
Yohei Seki, Yihong Liu · 2022
The popularity of Internet slang words has become a trend. Some of these new words can be seen as out-of-vocabulary words, but some previously registered words have new uses. In the case of social media texts, Internet slang words with important meanings cannot be ignored by replacing them with unknown words, as is the usual practice. Definitions for Internet slang words can be divided into two main types: “new semantic words” and “new blended words,” which can then be subcategorized based on the richness of the word formations. By combining these definitions with the strong relevance of word formations, we propose a hierarchical, shared, multi-task learning method based on joint character and word embeddings to label sequentially the main types and subcategories of Internet slang words. In our experiments, our method demonstrated the second-best performance behind multi-task BERT in detecting the main types of Internet slang words, and performed best in detecting subcategories among the shared-language models tested. In addition, our model achieved an average 32.5% F1-score improvement in detecting main types and subcategories compared with single-task methods. We conclude that the use of two-layer annotation improves the performance of the models, and it facilitates better observations and detailed analyses of differences in detection models.