An R package for discovering subtypes within complex heterogeneous datasets

Song Gao, Stefan Mutter, Ville‐Petteri Mäkinen · 2016

Many human diseases share risk factors and involve overlapping biological processes. Moreover, multiple diseases often co-occur in vulnerable individuals, and we need a knowledge discovery framework that can simultaneously handle combinations of morbidities, molecular and physiological features and environmental factors across diverse study designs to obtain more accurate disease subtypes. We have developed an R package, Numero, specifically designed for defining subtypes of samples with partially overlapping features or continuum of phenotypic characteristics. In our framework, the self-organizing map (SOM), an unsupervised pattern recognition technique, is adopted to organize high-dimensional data on a 2D canvas according to rank-based similarity criteria. The obtained map is coloured according to locally averaged values for a particular variable, thus revealing the differences in the phenotypic profiles between specific subpopulations in an easily observable visual format. Here, we demonstrate its use in two case studies (Phosphorylation dataset and the Framingham Cohort).

Read the paper · More papers on PaperTik