Data-driven multi-objective optimization of ML inference hardware configurations for energy, performance and cost
Joel Castaño, Jaime Bustillo, Xavier Franch, Silverio Martínez‐Fernández · Sustainable Computing Informatics and Systems · 2026
Selecting ML inference hardware requires balancing energy, performance, and cost, a core sustainable computing challenge. We present a data-driven framework that provides actionable decision support by identifying Pareto-optimal hardware configurations. We leverage a combined MLPerf™ Inference (v4.1/v5.1) corpus, enriching it with standardized hardware descriptors, proxy energy (TDP-based with a conservative CPU factor), and indicative component cost. We then train task-specific regressors to predict throughput and energy per unit, using grouped cross-validation to avoid system-level leakage. The framework combines these predictions with cost to generate workload-specific, multi-objective recommendations. Across most tasks, results show divergent feature importance: accelerator scale and generation dominate throughput, while energy efficiency depends on architecture, memory subsystem, host CPU, and the Offline vs. Server scenario. Predictive models demonstrate strong accuracy: predicted Pareto sets match true fronts on held-out systems (median precision 0.92, recall 0.89). For computer vision, a balanced recommendation reduced energy per unit by 75.0% and cost by 86.5% compared to a max-throughput baseline, retaining 46.2% performance. Furthermore, our Top-3 predicted balanced configurations included a true Pareto-optimal system in 100% of test cases, versus 34.5% for random selection. For generative language tasks, prediction accuracy is robust for modern LLMs, though specialized tasks show variability due to unmodeled software features. Sensitivity analyses confirm recommendations remain robust to energy proxy and pricing uncertainty. This framework makes complex sustainable hardware selection trade-offs explicit and actionable. • Recommender balances throughput, energy, and cost for ML inference. • Enriched MLPerf data with standardized hardware and energy descriptors. • Predicts throughput and energy per unit; finds Pareto-optimal configs. • Balanced picks cut energy 75% and cost 86.5% vs max-throughput baseline. • Enables sustainable computing decisions by turning public benchmarks into energy and cost-aware hardware recommendations.