Vietnamese Diacritics Restoration Using Deep Learning Approach

Bùi Thanh Hùng · 2018

This paper presents a solution for the insertion of diacritics into a text where they are missing, especially for Vietnamese text. Missing diacritics on text is making it difficult for the recipient to understand the content being transmitted or wasting time re-reading the content. Due to language distinction, Vietnamese language has many diacritical marks and more particularly, white space is not often seen as the basis for determining the word. In this paper, we propose a deep learning mechanism for automatic insertion of diacritics. The proposed model is a character-based Long Short Term Memory network serving as a language model. The deep learning approach is powerful, and relatively robust to rare cases that may occur in the text, with little reduction in accuracy. Building the corresponding model through this approach, our experiments have shown that the achieved results and the ability to enhance proposed system are obviously promising.

Read the paper · More papers on PaperTik