Rethinking data-efficient artificial intelligence for low-resource settings
Ronald Katende · Machine Learning with Applications · 2025
Recent progress in artificial intelligence has been shaped by data abundance and computational scale, yet these assumptions rarely hold in low-resource environments. This paper examines how constraints in data availability, compute, connectivity, and institutional capacity redefine what effective AI should look like. Through a structured, mixed-methods review, we identify and compare data-efficient approaches such as physics-informed models, few-shot and self-supervised learning, parameter-efficient fine-tuning, TinyML, and federated learning. Across sectors including health, agriculture, climate, and education, we show that lean, operator-informed, and locally validated methods outperform conventional large-scale models under real-world constraints. We argue that data-efficient AI is not a stopgap but a foundational paradigm for equitable and sustainable innovation, offering a roadmap for research, policy, and deployment in the Global South. • Provides a structured synthesis of over 300 studies on data-efficient machine learning for low-resource environments, using a PRISMA-based methodology. • Maps the alignment between structural constraints; data scarcity, energy, connectivity, and governance; and methods such as physics-informed learning, self-supervised and few-shot training, TinyML, and federated systems. • Quantifies key regional asymmetries (R&D intensity, electricity access, internet connectivity, and research capacity) to show why classical big-data paradigms underperform in the Global South. • Demonstrates through health, agriculture, climate, and education case studies how data-efficient AI enables practical, equitable deployments under infrastructural constraints. • Discusses performance trade-offs and ethical considerations, including energy efficiency, fairness, and institutional capacity for responsible AI scaling in low-resource regions.