A Bayesian approach to Nested Clade Analysis

Ioanna Manolopoulou · 2009

The purpose of this study is to identify genetically distinct clusters of individuals based on related characteristic traits (namely phenotypic data) or geographical locations (namely phylogeographic data). There are 2 main steps to this process: inferring the genetic history of the sequences under study, and subsequently identifying significant clusters according to the phenotypic/phylogeographic measurements. Based on an evolutionary model and an appropriate model for the distribution of the phenotype, such inference is possible in a number of different ways. However, due to the multiple level uncertainty and the complexity of the models, it is essential that the methods avoid stepwise optimization in order to give statistically reliable conclusions. The main methods currently used for analysis of this type are called Nested Clade Analysis (NCA) and Nested Clade Phylogeographic Analysis (NCPA) for phenotypic and phylogeographic data respectively. In short, they rely on finding the optimal genetic history based on a simplified evolutionary model, and identifying significantly different clusters for the phenotype/geography (assuming the inferred genetic history as fixed) by using Nested Analysis of Variance and permutation tests. Such methods do not allow for the uncertainty of each step

Read the paper · More papers on PaperTik