Combining multiple methods to improve differential expression testing
Luke R. Zappia, Fanny Grillet, Christina Mølck, Kym Pham, Julie Pannequin, Graham R. Taylor, Frédéric Hollande, Arthur L. Hsu · 2015
RNA-seq has rapidly become the experiment of choice for many biological investigations. Perhaps the most common analysis conducted on RNA-seq data is differential expression testing in order to identify changes in gene product regulation between conditions. This type of analysis typically consists of several stages: alignment to a reference genome, summarisation by the feature of interest, normalisation to minimise batch effects, testing for differential expression and functional analysis to examine the biological implications. At each stage the analyst must grapple with the choice of which many existing bioinformatics tools to use. A range of factors can affect this decision. Is the tool easy to use and readily available? Has it been used in previously published work? Is it designed for this data? Has it been based on good theory? Ideally choices should be made based on proven real-world performance, however this is hard to establish given the broad range of experiments and the lack of a ground truth to test against. Often analysts, particularly those new to the field, are left to pick a tool and trust that it is being used correctly and produces the required results. Focusing on the testing stage I will demonstrate some of the ways in which results can be affected by tool choice through analysis of a complex colorectal cancer dataset. I will present my experience in selecting which tool to use and suggest that where time permits it may be preferable to run multiple methods in parallel and combine the results.