CTA: Stata module for conducting Classification Tree Analysis

Ariel Linden · RePEc: Research Papers in Economics · 2020

Classification tree analysis (CTA) models use one or more attributes to classify a sample of observations into two or more subgroups that are represented as model endpoints (these are called “terminal nodes” in alternative decision-tree methods). Subgroups are known as “sample strata” because the CTA model stratifies the sample into subgroups of observations that -- with respect to model attributes -- are homogeneous within and heterogeneous between strata (Yarnold & Soltysik 2016). The pruned CTA algorithm involves chained optimal discriminant analysis (ODA) models in which the initial (“root”) node represents the attribute achieving the highest effect size for sensitivity (ESS) value for the entire sample, and additional nodes yielding greatest ESS are iteratively added at every step on all model branches while retaining statistical significance (defined by the prune() option). In contrast, the enumerated-optimal CTA algorithm explicitly evaluates all possible combinations of the first three nodes, which dominate the solution. cta is a wrapper program for the Classification Tree Analysis (CTA) software (Yarnold & Soltysik 2016). Therefore, CTA must be installed in order for the cta Stata package to work. CTA software is available at https://odajournal.com/resources/

Read the paper · More papers on PaperTik