Covariate Shift in Machine Learning

Santosh Chapaneri, Deepak Jayaswal · Apple Academic Press eBooks · 2022

Covariate shift occurs in machine learning when the input train and test probability distributions are different even though the conditional distribution of output given the train and test inputs remain unchanged. Most existing supervised machine learning techniques make an assumption that the train and test data samples follow the same probability distribution, but this is violated in practice for many real-world applications. In this work, the covariate shift is corrected by learning the importance weights for re-weighting the train data such that the training samples closer to the test data samples get more importance during modeling. To estimate the importance weights, a computationally efficient Frank-Wolfe optimization algorithm is used. Structured prediction is required to model the dependency that can exist in the multi-dimensional target variables, for example, in the application of human pose estimation. Twin Gaussian process (TGP) structured regression is used to model this dependency for improving the prediction performance relative to learning the multiple target dimensions separately. A computationally efficient method of TGP is presented for covariate shift correction and the performance is evaluated on the benchmark HumanEva dataset resulting in significantly reduced regression error.

Read the paper · More papers on PaperTik