An evaluation framework for measuring prompt wise metrics for large language models in resource-constrained edge

Partha Pratim Ray, Mohan Pratap Pradhan · BenchCouncil Transactions on Benchmarks Standards and Evaluations · 2025

Existing challenges in deploying large language models (LLMs) on resource-constrained devices stem from limited CPU throughput, memory capacity, and power budgets. Motivated by the lack of edge-specific evaluation tools, we introduce LLMEvaluator , a framework that profiles quantized LLMs — Qwen2.5, Llama3.2, Smollm2, and Granite3 — on a Raspberry Pi 4B using a suite of core and derived metrics. Our contributions include (i) a unified taxonomy that integrates latency, throughput, power variation, memory stability, and thermal behavior; (ii) prompt-wise analyses across ten NLP tasks; and (iii) correlation studies guiding optimizations. Key results show that Qwen2.5 leads in energy efficiency and throughput with a 68.44 MB memory standard deviation; Granite3 excels in memory stability , minimal load overhead, and per-token latency; Smollm2 suffers the highest total duration, longest prompt overhead, and lowest power efficiency; and Llama3.2 balances latency, throughput (8.12 tokens/s), and energy per token with moderate power variability (1.05 W std dev). Correlation analysis reveals that reducing model load time yields the largest improvement in end-to-end latency ( r > 0 . 9 ), and that throughput gains directly translate into energy savings ( r ≈ − 0 . 81 ). LLMEvaluator empowers selection and tuning of LLMs for low-power environments.

Read the paper · More papers on PaperTik