SimCorrMix: Simulation of Correlated Data with Multiple Variable Types Including Continuous and Count Mixture Distributions

Allison Fialkowski, Hemant K. Tiwari · The R Journal · 2019

The SimCorrMix package generates correlated continuous (normal, non-normal, and mixture), binary, ordinal, and count (regular and zero-inflated, Poisson and Negative Binomial) variables that mimic real-world data sets.Continuous variables are simulated using either Fleishman's thirdorder or Headrick's fifth-order power method transformation.Simulation occurs at the component level for continuous mixture distributions, and the target correlation matrix is specified in terms of correlations with components.However, the package contains functions to approximate expected correlations with continuous mixture variables.There are two simulation pathways which calculate intermediate correlations involving count variables differently, increasing accuracy under a wide range of parameters.The package also provides functions to calculate cumulants of continuous mixture distributions, check parameter inputs, calculate feasible correlation boundaries, and summarize and plot simulated variables.SimCorrMix is an important addition to existing R simulation packages because it is the first to include continuous mixture and zero-inflated count variables in correlated data sets.(Mohammadi, 2017); CAMAN, which provides tools for the analysis of finite semiparametric mixtures in univariate and bivariate data (Schlattmann et al., 2016); flexmix, which implements mixtures of standard linear models, generalized linear models and model-based clustering (Gruen and Leisch, 2017); mixdist, which applies to grouped or conditional data (MacDonald and with contributions from Juan Du, 2012); mixtools and nspmix, which analyze a variety of parametric and semiparametric models (Young et al., 2017;Wang, 2017); MixtureInf, which conducts model inference (Li et al., 2016); and Rmixmod, which provides an interface to the MIXMOD software and permits Gaussian or multinomial mixtures (Langrognet et al., 2016).With regards to count mixtures, the BhGLM, hurdlr, and zic packages model zero-inflated distributions with Bayesian methods (Yi, 2017;Balderama and Trippe, 2017;Jochmann, 2017).Given component parameters, there are existing R packages which simulate mixture distributions.The mixpack package generates univariate random Gaussian mixtures (Comas-Cufí et al., 2017).The distr package produces univariate mixtures with components specified by name from stats distributions (Kohl, 2017;R Core Team, 2017).The rebmix package simulates univariate or multivariate random datasets for mixtures of conditionally independent Normal, Lognormal, Weibull, Gamma, Binomial, Poisson, Dirac, Uniform, or von Mises component densities.It also simulates multivariate random datasets for Gaussian mixtures with unrestricted variance-covariance matrices (Nagode, 2017).Existing simulation packages are limited by: 1) the variety of available component distributions and 2) the inability to produce correlated data sets with multiple variable types.Clinical and genetic studies which involve variables with mixture distributions frequently incorporate influential covariates, such as gender, race, drug treatment, and age.These covariates are correlated with the mixture variables and maintaining this correlation structure is necessary when simulating data based on real data sets (plasmodes, as in Vaughan et al., 2009).The simulated data sets can then be used to accurately perform hypothesis testing and power calculations with the desired type-I or type-II error.SimCorrMix is an important addition to existing R simulation packages because it is the first to include continuous mixture and zero-inflated count variables in correlated data sets.Therefore, the package can be used to simulate data sets that mimic real-world clinical or genetic data.SimCorrMix generates continuous (normal, non-normal, or mixture distributions), binary, ordinal, and count (regular or zero-inflated, Poisson or Negative Binomial) variables with a specified correlation matrix via the functions corrvar and corrvar2.The user may also generate one continuous mixture variable with the contmixvar1 function.The methods extend those found in the SimMultiCorrData package (version ≥ 0.2.1, Fialkowski, 2017; Fialkowski and Tiwari, 2017).Standard normal variables with an imposed intermediate correlation matrix are transformed to generate the desired distributions.Continuous variables are simulated using either Fleishman (1978)'s third-order or Headrick (2002)'s fifth-order polynomial transformation method (the power method transformation, PMT).The fifth-order PMT accurately reproduces non-normal data up to the sixth moment, produces more random variables with valid PDF's, and generates data with a wider range of standardized kurtoses.Simulation occurs at the component-level for continuous mixture distributions.These components are transformed into the desired mixture variables using random multinomial variables based on the mixing probabilities.The target correlation matrix is specified in terms of correlations with components of continuous mixture variables.However, SimCorrMix provides functions to approximate expected correlations with continuous mixture variables given target correlations with the components.Binary and ordinal variables are simulated using a modification of GenOrd's ordsample function (Barbiero and Ferrari, 2015b).Count variables are simulated using the inverse cumulative density function (CDF) method with distribution functions imported from VGAM (Yee, 2017).Two simulation pathways (correlation method 1 and correlation method 2 ) within SimCorrMix provide two different techniques for calculating intermediate correlations involving count variables.Each pathway is associated with functions to calculate feasible correlation boundaries and/or validate a target correlation matrix rho, calculate intermediate correlations (during simulation), and generate correlated variables.Correlation method 1 uses validcorr, intercorr, and corrvar.Correlation method 2 uses validcorr2, intercorr2, and corrvar2.The order of the variables in rho must be 1 st ordinal (r ≥ 2 categories), 2 nd continuous non-mixture, 3 rd components of continuous mixture, 4 th regular Poisson, 5 th zero-inflated Poisson, 6 th regular Negative Binomial (NB), and 7 th zero-inflated NB.This ordering is integral for the simulation process.Each simulation pathway shows greater accuracy under different parameter ranges and Calculation of intermediate correlations for count variables details the differences in the methods.The optional error loop can improve the accuracy of the final correlation matrix in most situations.The simulation functions do not contain parameter checks or variable summaries in order to decrease simulation time.All parameters should be checked first with validpar in order to prevent errors.The function summary_var generates summaries by variable type and calculates the final correlation matrix and maximum correlation error.The package also provides the functions calc_mixmoments to calculate the standardized cumulants of continuous mixture distributions, plot_simpdf_theory to plot simulated PDF's, and plot_simtheory to plot simulated data values.The plotting functions work

Read the paper · More papers on PaperTik