Universal models with memory for genomic sequence analysis
Ioan Tăbuş, Yinghua Yang, Jaakko T. Astola · 2008
In this paper we discuss the use of universal models for solving several genomic sequence analysis problems. A number of typical genomic problems, e.g., approximate matching, segmentation, and clustering, can be phrased as specific modeling problems involving discrete variables, for which discrete regression models need to be estimated based on rather short data segments. Universal models are known to possess appealing optimality properties, not only asymptotically, but also for short samples. We briefly review universal models with memory, which have been shown recently to perform well for the compression of full genomes. Two new applications of universal models with memory for genomic sequence analysis are shown here, the first one is the segmentation of DNA sequences for uncovering gene duplications and the second one is haplotype segmentation.