Identifying character non-independence in phylogenetic data using parallelized rule induction from coverings
Jennifer L. Leopold, Anne M. Maglia, M. Thakur, Bhadresh Patel, Fikret Erçal · WIT transactions on information and communication technologies · 2007
Undiscovered relationships in a data set may confound analyses, particularly those that assume data independence.Such problems occur when characters used for phylogenetic analyses are not independent of one another.Although a data mining technique known as rule induction from coverings has earlier been shown to be a promising approach for identifying such non-independence, its inherent computational complexity has limited its application for large phylogenetic data sets.Herein we present a parallelized implementation of the rule induction from coverings strategy which overcomes some of these limitations.We also discuss two heuristics that have been applied to the algorithm to further improve its efficiency.