Energy-efficient large language models

Javier Jareño, José Miguel Aragón-Jurado, Juan Carlos de la Torre, Patricia Ruiz, Bernabè Dorronsoro · Future Generation Computer Systems · 2026

Large Language Models (LLMs) have become widely used tools across domains such as education and scientific research, following recent progress in computing infrastructure, data availability, and training techniques. However, their increasing deployment raises significant concerns about their high computational and energy demands at inference time. This work tackles the problem from a software-stack perspective, presenting an automatic optimization pipeline for LLM inference engines. This task is formulated as a combinatorial optimization problem and addressed using a cellular genetic algorithm that evolves variable-length sequences of code optimizations. The algorithm proposes a novel density-guided mutation and a longest common subsequence crossover operator. Experiments conducted on a Mac Mini with Apple Silicon M4 using llama.cpp demonstrate that, for two Llama 3.2 variants (with 1 and 3 billion parameters) quantized to 4-bit precision, the optimized engine reduces total energy consumption by 13.5% and runtime by 19.5% for the 1-billion-parameter model, compared to the non-optimized version, while the corresponding reductions for the 3-billion-parameter model were 5.3% in energy and time. If we compare against the aggressive -O3 compilation optimization flag, 4.4% and 2.7% energy savings are achieved for the two optimized engines, respectively, with comparable runtimes. It was also studied how the optimized engines generalize to their 8-bit versions, and their improvements over the non-optimized baseline were around 12% in energy consumption and 11.5% in runtime. These results demonstrate cross-quantization generalization and confirm that hardware-aware optimizations can deliver energy savings without modifying the neural model, offering a portable optimization methodology toward greener LLM deployment.

Read the paper · More papers on PaperTik