Data Transformation
Adam P. Tashman · 2024
Raw data will generally need to be transformed before it can be useful. We start by reviewing the practice of smoothing noisy data with a moving average. Methods for treating outliers are discussed, including winsorizing and trimming. It is difficult to visualize data that changes in magnitude, and it is difficult to compare variables of different magnitudes. We study transformations to handle scale, such as the use of the logarithm, standardization, and normalization. A discussion of when to standardize and when to normalize data is included. We review methods for converting text data to a quantitative form. This includes the count vectorizer, which can be used to represent a set of documents as a term-document matrix. Additional techniques for representing data are discussed, including one hot encoding, binarization, and discretization. One hot encoding can convert categorical data to a vector indicating the variable&s;s level. Binarization works by accentuating certain parts of the data. Discretization converts a continuous variable into discrete buckets, which sometimes yields a powerful predictor. Finally, transforming data can be an intensive process. We discuss how best to store and share the results.