Challenges in Cluster Analyses for Longitudinal Data
Liesbeth M. Bruckers · Document Server@UHasselt (UHasselt) · 2014
In this dissertation, we have addressed clustering for high dimensional data, possibly subject to missingness. The research was inspired by a number of data sets, ranging from data collected in a mental care setting, studies in patients with abdominal aortic aneurysm or heart failure, to an EEG study in rats. The communality in these studies is the believe that the population under investigation is not homogenous, but instead consists of subpopulations. A direct labelling of these subpopulations is not available. But given that these sub-populations are characterized by different structures in the collected data, it is possible to uncover the latent subpopulations. Model-based clustering is a statistical tool that can be entertained for this purpose. However, for the given data sets, clustering is impeded by the high dimensionality and longitudinal character of the data, and by the fact data is not always fully observed. Often a set of outcomes is measured over time, resulting in multivariate longitudinal data. Due to the dimension of the joint distribution of the random effects, computational problems are likely to occur when mixture models are applied to a multivariate longitudinal setting. In this dissertation, we have proposed an algorithm to reveal latent subgroups for multivariate repeated outcomes. The approach is inspired by work of Fieuws and Verbeke (2008), the authors perform a discriminant analysis for repeatedly measured data. Instead of maximizing the full joint model a pseudolikelihood approach, based on bivariate joint models for the repeated outcomes, was utilized. The iterative algorithm mimics a partition cluster method. The performance of the proposed algorithm was looked into by means of a simulation study. Complexity is enhanced when observations are densely sampled over a continuum, e.g., time. In such a situation, the data are generated by an underlying smooth function or by a set of smooth functions that are not easily described by a mathematical expression. Functional data analysis methods are used to reduce the dimensionality of the data and latent subgroups are then discovered for the reduced data. Such an approach was, e.g., used in Jacques and Preda (2013). The fact that their approach uses a data reduction technique, requiring a complete data structure, limits the practical usefulness of the cluster algorithm. In this dissertation, we combined methods from functional data analysis, missing data and ensemble clustering to discover latent subgroups in high-dimensional data, in terms of the number of responses and the number of repeated measurements, contaminated by missing observations. Data were completed by means of multiple imputation, whereupon the model-based clustering of Jacques and Preda (2013) was used to find latent subgroups in the principal components, and finally ensemble clustering was employed to summarize the set of partitions into a final data partition. The amalgamation of statistical techniques allows to cluster complex data and at the same time to quantify the influence of the missing data on the composed groups. Ensemble clustering has to our knowledge not yet been used in combination with multiple imputation. A small simulation study was designed to explore its utility. When the missing-data mechanism is believed to be non-random, the joint distribution of the data and the missing-data indicators should be considered. In this work, we have investigated various mixture models for non-random missingness as proposed by Muth´en et al. (2011). We assessed the vulnerability of the results not only in terms of the number of clusters, the cluster-specific profiles, but also in terms of the group-membership probabilities. It is however impossible to decide on the best model, since all models rely on non-verifiable assumptions. We have illustrated how an ultimate outcome, related to the growth curves, can be supportive in choosing between the models. Cluster results are of course also sensitive to outlying and influential observations. We used ideas presented by Lesaffre and Verbeke (1998) for a mixed model, and applied local-influence diagnostics to a mixture model. This allowed quantification of the influence an observation has on the cluster-specific profiles and on the groupmembership probabilities of the other observations. A number of issues were not or partially addressed in this dissertation and could be topic for further deepening. The local-influence diagnostics, described in Chapter 8, are obtained by introducing weights for the log-likelihood contributions of single subjects, where focuss was on the influence of a single subject. Other perturbation schemes could be worthwhile to consider. The method of local influence could for example be used to study the impact of MNAR mechanisms on the cluster result. The approach presented in Chapter 6, to cluster sets of smooth but incomplete functions, has a number of flaws. The method is sensitive to the class-specific orders to approximate the pseudo-likelihood for the functional data. A heuristic test is used to determine these orders. More formal procedures could be implemented. Determination of the number of clusters is also difficult. An information criterion similar to the one proposed by Breaban and Luchian (2011) could be developed for functional data. This would address the selection of the class-specific orders and the optimal number of clusters at the same time. The set of partitions is reduced into a final partition by means of consensus clustering. This step could be replaced by other techniques, for example a latent class analysis with the cluster-indicators as variables. We have focussed on repeated measurements for continuous responses. But the methods presented in this dissertation, can be applied to non-continuous responses/data with other structures. It would be interesting to see how the methods perform in spatial or temporal-spatial settings and for combinations of responses not belonging to the same parametric family.