CNprep: Copy number event detection
Pascal Belleau, Astrid Deschênes, Guoli Sun, David Arthur Tuveson, Alex Krasnitz · 2020
Genome-wide DNA copy number profiles are a form of molecular data used in multiple areas of genomic analysis and derived from a variety of platforms, including microarrays and next-generation sequencing. Here we present an improved, Bioconductor-compliant version of CNprep (Sun and Krasnitz, 2014), an R package for inference of significant gains and losses in genome-wide DNA copy number profiles. Such inference is relatively simple for profiles derived from normal genomes of an organism, where germline copy number variants (CNV) occupy narrow genomic regions and copy number values are integer. In cancer, these profiles often represent mixtures of cells with different CNV patterns, and CNV extend over large portions of the genome. As a result, it may be difficult to distinguish genuine CNV from platform-dependent noise. In CNprep , CNV are identified based on the assumption that the largest portion of the genome is in a normal, CNV-free state. To identify this portion, CNprep uses data from two sources: the CNV profile of interest and, optionally, as a reference, a set of CNV profiles from cancer-free tissues, with nearly all of the genome in a normal state. Two data types are used as input: copy number values as a noisy function of the genomic location and a piecewise-constant approximation of that function called segmentation. The package employs a clustering approach to identify the portion of the genome in the normal copy number state, and computes, for each segment, its marginal probability to belong to that state. Segments with low marginal probabilities are flagged as amplified or deleted.