Optimizing Agentic AI Applications Under Budget Constraints
Hamta Sedghani, Federica Filippini, Zahra Seyedi, Blerina Spahiu, Michele Ciavotta, Danilo Ardagna · 2026
AI systems increasingly rely on multi-step agentic workflows that orchestrate heterogeneous computational agents, including Large and Small Language Models (LLMs and SLMs), each configurable at runtime with different inference methods and parameters. While this flexibility enables specialization and resource efficiency, it introduces complex trade-offs between accuracy, latency, and energy consumption, particularly under strict operational budgets. The sequential nature of these workflows further amplifies such challenges: errors propagate across steps, and resource decisions made early in the pipeline constrain downstream options. To address these challenges, this paper presents a formal optimization framework for configuring multi-step agentic pipelines under global time and energy constraints. Given a task decomposition and a heterogeneous pool of agent configurations, the framework selects exactly one agent-inference method-parameter triple per task to maximize end-to-end workflow quality. We formalize this as a Mixed-Integer Linear Programming problem and introduce two complementary objective formulations: a max-min accuracy objective that prioritizes robustness by improving the least accurate step, and a multiplicative accuracy objective that captures cumulative performance across the workflow. We further propose an iterative solution strategy that re-optimizes at each step, adapting to deviations between predicted and realized resource consumption. Experimental results demonstrate distinct accuracy-resource trade-offs between the two formulations and show that the approach scales effectively to workflows with up to a thousand steps.