partition: A fast and flexible framework for data reduction in R

Malcolm Barrett, Joshua Millstein · The Journal of Open Source Software · 2020

Data are increasingly feature-rich, including many variables for each observation; in modern genomics, for example, high-resolution genetic data captures much more information than it did just a decade ago.While improved measurements contribute immensely to science, they also increase computational burden and complicate interpretability (Karczewski & Snyder, 2018).Data reduction techniques, such as principal component analysis (PCA) and K-Means clustering, are vital tools used to address these issues, particularly for noise and redundancy.However, these techniques may lead to problems in scalability, information loss, and interpretability (Malod-Dognin, Petschnigg, & Pržulj, 2018).The Partition framework is an approach to data reduction that is flexible, scalable, and interpretable, developed to address information loss while maintaining speed (Millstein et al., 2020).As opposed to other data reduction strategies, Partition only reduces data if the data reduction retains a specified amount of information.This framework is also agnostic to how partitions form; users can easily use other tools such as PCA and t-Distributed Stochastic Neighbor Embedding (t-SNE) to create a partition or summarize data within a partition subset and thus reduce data while constraining information loss.

Read the paper · More papers on PaperTik