• 14 October, 2025
  • GreenCode

New paper: EnergyBench: Holistic Benchmark for Correct and Energy-Efficient LLM-Generated Code

The increasing use of Large Language Models (LLMs) for automated software development creates a paradox: LLM-generated code can boost energy efficiency across industries through digital transformation but their often unsupervised usage unintentionally generates energy-inefficient code when functional correctness is prioritized. Current benchmarks show that LLMs generate energy-efficient code when prompted with optimized instructions. However, it is unclear for which problems, programming languages, and prompting strategies lead to the best trade-off between correctness and energy efficiency. To address the current gaps in untested prompts, this work uses a systematic approach to analyze which prompt elements lead to LLMs generating correct and energy-efficient code. Results reveal that prompt engineering increases code energy efficiency by up to 91.9% in some cases, but often reduces accuracy. Our novel approach, inspired by Holistic Evaluation of Language Models (HELM), reveals that LLMs are sensitive to prompt contents. One result shows that the energy efficiency increases by more than 4 × when half of the programming problem description is removed, while in other cases accuracy drops to zero.

Authors: Dragoș Ionescu, Søren Kejser Jensen, Bent Thomsen

Contributing partner: Aalborg University

Publication page and citation View at the publisher