BoTTA: Benchmarking On-device Test Time Adaptation
Michal Danilowski, Soumyajit Chatterjee, Abhirup Ghosh · 2025
The performance of deep learning models at inference time is susceptible to distribution shifts from the training data, often leading to substantial accuracy drops. Test-time adaptation (TTA) mitigates this by adapting models during inference without requiring labeled test data or access to the training set. While prior work has examined TTA from various perspectives, including algorithmic design, distributional shifts, and learning paradigms, the unique constraints of mobile and edge deployments remain underexplored. We introduce BoTTA, a benchmark tailored to evaluate TTA methods under real-world constraints of mobile and edge devices. BoTTA focuses on four core challenges shaped by limited resources and usage conditions: (i) few test samples, (ii) exposure to a limited set of categories, (iii) diverse distribution shifts, and (iv) compound shifts within a single sample. We evaluate state-of-the-art TTA methods on these axes using standard datasets and report accuracy and system-level performance on real hardware testbeds. Notably, BoTTA departs from continuous inference-time adaptation, promoting periodic adaptation that is more suitable for on-device scenarios.