Efficient dictionary and language model compression for input method editors

Taku Kudo, Toshiyuki Hanaoka, Jun Mukai, Yusuke Tabata, Hiroyuki Komatsu · Workshop on Advances in Text Input Methods · 2011

Reducing size of dictionary and language model is critical when applying them to real world applications including machine translation and input method editors (IME). Especially for IME, we have to drastically compress them without sacrificing lookup speed, since IMEs need to be executed on local computers. This paper presents novel lossless compression algorithms for both dictionary and language model based on succinct data structures. Proposed two data structures are used in our product “Google Japanese Input” 1 , and its open-source version “Mozc” 2 .

Read the paper · More papers on PaperTik