Boosting N-gram Coverage for Unsegmented Languages Using Multiple Text Segmentation Approach
Solomon Teferra Abate, Laurent Besacier, Sopheap Seng · 2010
Automatic word segmentation errors, for languages having a writing system without word boundaries, negatively affect the performance of language models. As a solution, the use of multiple, instead of unique, segmentation has recently been proposed. This approach boosts N-gram counts and generates new N-grams. However, it also produces bad N-grams that affect the language models ' performance. In this paper, we study more deeply the contribution of our multiple segmentation approach and experiment on an efficient solution to minimize the effect of adding bad N-grams. 1