Variable-length sequence modeling: multigrams
Frédéric Bimbot, Roberto Pieraccini, Esther Levin, Bishnu Saroop Atal · IEEE Signal Processing Letters · 1995
The conventional n-gram language model exploits dependencies between words and their fixed-length past. This letter presents a model that represents sentences as a concatenation of variable-length sequences of units and describes an algorithm for unsupervised estimation of the model parameters. The approach is illustrated for the segmentation of sequences of letters into subword-like units. It is evaluated as a language model on a corpus of transcribed spoken sentences. Multigrams can provide a significantly lower test set perplexity than n-gram models.>