Mallitason viiveen optimointi agenttimaisissa tekoälyjärjestelmissä sääennustedatakontekstissa

Otso Aksela · Aaltodoc (Aalto University) · 2026

Agentic AI systems consist of multiple LLMs with access to tools defined by an ability to independently pursue goals. These systems suffer from significant cost and latency challenges that limit adoption. Especially latency is a direct disruptor of utility for real-world applications. The issue is greatest when using a simple generalist agent, as it is required to be tuned to the most difficult challenges, making it overkill for most. This thesis explores model-level latency optimization strategies for AI agents in a weather forecast context for Vaisala Xweather. The thesis introduces two frameworks for improving model-level latency. The first framework centers around LLM-to-RNN knowledge transfer, where Claude Haiku 3.5 is used as a teacher and an RNN as a student for forecast reliability estimation. The second framework constructs an embedding-based model cascade with Mistral Ministral 3 8B for predictive tool calling preceding the generalist agent. The results indicate that an LLM can successfully estimate forecast reliability, and that an RNN can complete the same task with comparable results while being two orders of magnitude faster and three orders of magnitude cheaper. Additionally, it is found that the predictive tool calling system offers latency reduction by a factor of 8.6 and cost reduction of two orders of magnitude. The RNN is limited by regional specificity, and the predictive tool calling is only 82.5\% accurate. Overall, the thesis contributes two reproducible frameworks for model-level latency optimization that can be combined to form a hybrid model with the benefits of both, offering a partial solution to the latency and cost problem of AI agents.

Read the paper · More papers on PaperTik