Lexical Normalization of Japanese Tweets Using Related Images
Zhelin Xu, Atsushi Matsumura, Tetsuji Satoh · 2021
Twitter is noisy and contains many nonstandard words. Furthermore, in Japanese tweets, many words have multiple variant notations. Therefore, the use of such noisy data may interfere with tasks such as identifying potential communities. In this paper, based on the assumption that words with the same meaning will have similar related images, we propose a method of normalization for nonstandard words and variant notations in Japanese tweets using related images. First, we collect images related to a word from Bing and use OpponentSIFT features to properly represent the content of those images. Next, we use clustering to narrow down the set of images to extract related images. Finally, we determine the similarity between words based on the similarity of the sets of related images.