Data Transformation (Pre-processing)
Steven Finlay · Palgrave Macmillan UK eBooks · 2012
Model construction techniques display varying degrees of sensitivity to the way data is presented to them. Data transformation is undertaken to provide an alternative representation of the data, that it is hoped will lead to a better (more predictive) model than would result from using the data in its original form. Data transformation typically achieves the following outcomes: Linearization. Transformations are applied so that the relationships between the predictor variables and the dependent variable are (approximately) linear. Having linear relationships is important for methods such as linear regression and logistic regression. If the relationships in the data are highly non-linear then poor models will result using these methods. Linearization is less important for non-linear techniques such as CART and neural networks. Standardization. If one predictor variable takes values in the range 10,000 to 1,000,000 and another takes values in the range 0.01 to 1, then the parameter coefficients (the model weights) will be very different, even if the two variables contribute equally to the model. This is not an issue for all model construction techniques, but as a rule, it is good practice to transform interval variables so that they all take values that lie on the same scale. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.