Towards Aggregation Based I/O Optimization for Scaling Bioinformatics Applications

Jack Stratton, Michael Albert, Quentin Jensen, Max Ismailov, Filip Jagodzinski, Tanzima Zerin Islam · 2020

Bioinformatics software often integrates multiple off-the-shelf programs into a single compute pipeline. Each stand-alone program generates output, that is frequently saved into a plain-text file, which is then processed as input by the program that is responsible for the next stage of the computation. Modern advances in genome sequencing and protein structure resolution methods have yielded large amounts of data, that when processed by bioinformatics compute pipelines, results in vast numbers of file reads and writes. Since computation capabilities of large-scale distributed systems grow much faster than their I/O bandwidth, processing such big data using these compute pipelines will not scale as the amount of data grows. For this work, we motivate and demonstrate a dynamic interception-based I/O analysis tool to assess the file read and write characteristics of a protein mutation generation pipeline. We discuss how our analysis tool can be further extended to apply compression and in-site analysis and has the potential to scale I/O-intensive bioinformatics applications on high performance computing (HPC) systems.

Read the paper · More papers on PaperTik