Data mining using genetic programming: the implications of parsimony on generalization error

Michael J. Cavaretta, Kumar Chellapilla · 2003

A common data mining heuristic is, "when choosing between models with the same training error, less complex models should be preferred as they perform better on unseen data". This heuristic may not always hold. In genetic programming a preference for less complex models is implemented as: (i) placing a limit on the size of the evolved program; (ii) penalizing more complex individuals, or both. The paper presents a GP-variant with no limit on the complexity of the evolved program that generates highly accurate models on a common dataset.

Read the paper · More papers on PaperTik