Pattern matching in compressed genomic sequence data

A. S. Keerthy, S. Manju Priya · 2017

Compressing Genomic data for efficiency of storage, transmission and retrieval has been a challenge for biologists as well as computer scientists across the globe for the past decade. The researchers and scientists have concluded on many measures to compress and store the genomic data. The present challenge faced by the research and scientific community is the analysis of these compressed data. It is always possible to decompress the data and do the analysis. Researchers do not consider it as an efficient method as it nullifies the advantages of compressing the genomic data. Analyzing the genomic data involves identifying the presence of microsatellites, tandem repeats, genes, etc. Pattern matching is the efficient way to detect the presence of a known pattern within a sequence. Compressed pattern matching, attempts to detect the presence of known patterns, in a compressed sequence without decompressing it. Compressed pattern matching has been successfully implemented for textual data. This paper attempts to explore a pattern matching technique in compressed genomic data without uncompressing it.

Read the paper · More papers on PaperTik