When Training and Test Sets are Different: Characterising Learning Transfer

Amos Storkey · 2013

In this chapter, a number of common forms of dataset shift are introduced, and each is related to a particular form of causal probabilistic model. Examples are given for the different types of shift, and some corresponding modelling approaches. By characterising dataset shift in this way, there is potential for the development of models which capture the specific types of variations, combine different modes of variation, or do model selection to assess whether dataset shift is an issue in particular circumstances. As an example of how such models can be developed, an illustration is provided for one approach to adapting Gaussian process methods for a particular type of dataset shift called Mixture Component Shift. After the issue of dataset shift is introduced, the distinction between conditional and unconditional models is elaborated in Section 3. This difference is important in the context of dataset shift, as it will be argued in Section 5 that dataset shift makes no difference for causally conditional models. This form of dataset has been called covariate shift. In Section 6, another simple form of dataset shift is introduced: prior probability shift. This is followed by Section 7 on sample selection bias, Section 8 on imbalanced data and Section 9 on domain shift. Finally three different types of source component shift are given in Section 10. One example of modifying Gaussian process models to apply to one form of source component shift is given in Section 11. A brief discussion on the issue of determining whether shift occurs (Section 12) and on the relationship to Transfer Learning (Section 13) concludes the chapter. 2

Read the paper · More papers on PaperTik