Detection of Chinese Word Usage Errors for Non-Native Chinese Learners with Bidirectional LSTM
Yow-Ting Shiue, Hen‐Hsen Huang, Hsin‐Hsi Chen · 2017
Selecting appropriate words to compose a sentence is one common problem faced by non-native Chinese learners.In this paper, we propose (bidirectional) LSTM sequence labeling models and explore various features to detect word usage errors in Chinese sentences.By combining CWIN-DOW word embedding features and POS information, the best bidirectional LSTM model achieves accuracy 0.5138 and MRR 0.6789 on the HSK dataset.For 80.79% of the test data, the model ranks the groundtruth within the top two at position level.