Diversity-Based Sampling for Imbalanced Domain Adaptation

Andrea Napoli, Paul Robert White · 2024

Many unsupervised domain adaptation (UDA) algorithms have been shown to fail when the target data is class-imbalanced – a scenario known as imbalanced domain adaptation (IDA). However, the lack of labels means this data cannot be balanced in the usual way. The most common workaround – to balance the data using pseudo labels – makes the strong assumption that the predicted labels are accurate. In this paper, we propose instead to balance the target data implicitly, by performing a diversity-based resampling of that data. This reduces the strength of the assumptions required to obtain an accurate balancing. We describe three diversity-based sampling algorithms, two based on k-means clustering and the other on the determinantal point process (DPP), which can be used to implicitly balance unlabelled data. These algorithms readily scale to large datasets, and are compatible with both shallow and deep UDA methods. Extensive analysis on a cross-dataset bioacoustic event detection task shows that this approach significantly outperforms prior work on IDA, and can achieve up to oracle-level performance for a wide range of imbalances.

Read the paper · More papers on PaperTik