Character-based Language Model
Vít Baisa · RASLAN · 2014
Abstract. Language modelling and also other natural language process-ing tasks are usually based on words. I present here a more general yetsimpler approach to language modelling using much smaller units of textdata: character-based language model (CBLM). 1 In this paper I describethe underlying data structure of the model, evaluate the model using stan-dard measures (entropy, perplexity). As a proof-of-concept and an extrin-sic evaluation I present also a random sentence generator based on thismodel. Keywords: language model, suffix array, LCP, trie, character-based, ran-dom text generator, corpus 1 Introduction Current approaches to language modelling are based almost utterly on words.To work with words, the input data needs to be tokenized which might bequite tricky for some languages. The tokenization might cause errors whichare propagated to following processing steps. But even if the tokenizationwas 100% reliable, another problem emerges: word-based language modelstreat similar words as completely unrelated. Consider two words platypusand platypuses. The former is contained in the latter yet they will be treatedcompletely independently. This issue can be sorted out partially by usingfactored language models [1] where lemmas and morphological information(here singular vs. plural number of the same lemma) are treated simultaneouslywith the word forms.In most systems, word-based language models are based on n-grams(usually 3–4) and on Markov chain of the corresponding order where only afinite and fixed number of previous words is taken into account. I propose amodel which tackles with the above-mentioned problems. The tokenizationis removed from the process of building the model since the model usessequences of characters (or bytes) from the input data. Words (byte sequences)which share prefix of characters (bytes) are stored on the same place in the