Robust character based tagging with domain lexical features for Chinese spoken language understanding

Changchun Bao, Yali Li, Li Ta, Jielin Pan, Yonghong Yan · 2010 Sixth International Conference on Natural Computation · 2010

Word information is useful in natural language understanding. But in Chinese language processing, word information is not given natural. While word-segmentation works well for text in NLU, it deteriorates Chinese SLU because of the flexibility and distortion of spoken utterance plus ASR errors. This paper propose a novel approach, sub-word features, to take use word information and help understanding spoken utterance while retain the robustness of character-wise processing. By means of this approach, we can also effectively use named entity list to improve SLU performance. Experiments show that the sub-word features give an average of 0.7 improvement for ASR, and the usage of named list given an average of 4.7 improvement.

Read the paper · More papers on PaperTik