A Compound Poisson Model for Word Occurrences in DNA Sequences

Stéphane Robin · Journal of the Royal Statistical Society Series C (Applied Statistics) · 2002

Summary We present a compound Poisson model describing the occurrence process of a set of words in a random sequence of letters. The model takes into account the frequency of the words and their overlapping structure. The model is compared with a Markov chain model in terms of fit and parsimony. Special attention is given to the detection of poor or rich regions. Several applications of the model are presented and a combination of the Markov and compound Poisson models is proposed.

Read the paper · More papers on PaperTik