Efficient Compression of Genomic Sequences

Diogo Pratas, Armando J. Pinho, Paulo J. S. G. Ferreira · 2016

The number of genomic sequences is growing substantially. Besides discarding part of the data, the only efficient possibility for coping with this trend is data compression. We present an efficient compressor for genomic sequences, allowing both reference-free and referential compression. This compressor uses a mixture of context models of several orders, according to two model classes: reference and target. A new type of context model, which is capable of tolerating substitution errors, is introduced. For ensuring flexibility regarding hardware specifications, the compressor uses cache-hashes in high order models. The results show additional compression gains over several specific top tools in different levels of redundancy. The implementation is available at http://bioinformatics.ua.pt/software/geco/.

Read the paper · More papers on PaperTik