Smoothlm: A language model compression library

Ahmet Afşın Akın, Cemil Demir · 2014

In this paper we will present SmoothLm, a language model compression and random access library. Like some other previous work, this library uses Minimal Perfect Hash Functions (MPHF) to reach high compression rates. We improved a previous MPHF algorithm in terms of generation and query speed and named it Multi Level MPHF. We also present a mechanism that use this MPHF structure on very large data sets quickly with limited memory usage. SmoothLm's generates lossy models and it provides a quantization mechanism for probability values for extra compression. We use SmoothLm in our in house speech recognition engine and our experiments showed that with correct parameters, being a lossy model or applying quantization does not hurt performance. Library is proper for applications developed in Java and source code is available with a free license.

Read the paper · More papers on PaperTik