Deep Neural Networks for Social Media Word Segmentation of Asian Languages

Ngoc Tan Le, Fatiha Sadat · 2018

Information extraction today faces new challenges with noisy, short, unstructured data. This is especially the case for social media messages, such as tweets, in which language can be erroneous or cryptic, and contains references to a great number of new entities. Traditional NLP systems are challenged and need to develop new strategies to handle with these data. With the emergence of the neural network-based approach, the research about the word segmentation has benefited from large-scale raw texts by leveraging them for pretrained character and word embeddings. To this end, we experimented the use of both character and word embeddings to provide extra features to input layer of our neural network-based system architecture. This system has been tested on both Chinese and Japanese social media datasets. With the help of rich pretrained embeddings, our model achieved the promising results both on Chinese and Japanese social media word segmentation task by comparing with the state-of-the-art NLP tools.

Read the paper · More papers on PaperTik