Green AI and Energy-Aware Computing

Green AI and energy-aware computing: poster for the GreenCode film
Video created by explainapaper.com

A pipeline that uses AI to save energy has to account for its own. GreenCode measures the carbon of every run, keeps it to a small fraction of the saving, and schedules the heavy work for when the power is clean.

The tool has to pay for itself

The energy cost of running the optimisation against the saving it produces over time, with a run or do-not-run gateView full size
The cure has to cost less than the disease, and the run that cannot show it should not happen.

AI is the fastest-growing consumer of energy in the ICT sector. The International Energy Agency put data centre consumption at about 460 TWh in 2022 and projected that it could exceed 1,000 TWh by 2026, roughly the electricity consumption of Japan, with AI workloads a principal reason for the climb. Any project proposing to use AI to reduce software energy has an obvious obligation: prove the cure costs less than the disease.

GreenCode takes that obligation literally. The carbon of the pipeline's own training and inference is modelled, measured and reported alongside the saving it produces, and the pipeline is only worth running where the ratio is decisively favourable. Research on ICT's climate impact has repeatedly found that claimed digital climate solutions rarely publish evidence of a net carbon benefit. GreenCode is built so that it can.

How the carbon is counted

The accounting is deliberately simple enough to audit, and it uses the providers' own figures. Google has measured a median of 0.24 Wh per Gemini text prompt and OpenAI has given about 0.34 Wh per average ChatGPT query. Their average over a conservative 500-token response is about 0.58 mWh per generated token. At roughly 15 tokens per refactored line and the European grid average of 242 g CO2e per kWh, that is about 2 mg of CO2e per line, before the extra tokens a model spends reading context.

Set against the outcome, the ratio is stark. Refactoring studies indicate that roughly a quarter of a codebase is actually changed, and that efficiency refactoring delivers energy improvements of 10 to 30 per cent. For a large open-source database of about 7 million lines running across tens of thousands of deployments, a 23 per cent improvement avoids in the region of 3,700 tonnes of CO2e a year, for around 4 kg of CO2e of AI inference on the modelled tokens. Even at a thousand times that, the inference is about a thousandth of one year's saving; the saving repeats every year the optimised version is deployed, and the carbon is spent once.

Methods that keep inference small

The ratio only holds because the pipeline is engineered to spend as little inference as possible. Four choices do most of that work.

  • Small and fine-tuned models are used in preference to large general-purpose ones, because a model tuned to a narrow code task can match a far larger model on that task at a fraction of the energy per token.
  • Retrieval-augmented generation supplies the relevant fragment of a codebase from an index rather than pushing whole files through a context window, so the token count reflects what the task needs.
  • Mixture-of-experts routing sends each request to the smallest component capable of answering it, instead of activating a full model for every prompt.
  • Successive runs process only the deltas between them, so a codebase is paid for in full once and thereafter only where it has changed.

Scheduling against the energy mix

The second half of energy-aware computing is when the work happens, not just how much of it there is. A tonne of computation run on a grid saturated with wind costs a fraction of the carbon of the same computation run on peaking gas. Because GreenCode analyses a codebase asynchronously rather than interactively, it can hold intensive jobs until clean power is available and release them when carbon intensity falls. Nothing is waiting on the result in the next second, so the latency is free.

That flexibility is also the point of difference from copilot-style assistants. Those tools run inference on essentially every keystroke, in the moment, wherever the developer happens to be, and the energy accumulates invisibly across thousands of silos with no view of the whole application. They optimise the typing. GreenCode optimises the system, in one scheduled pass, with the cost of that pass counted.

Deciding whether a run is worth it

Scale decides the answer, and it varies enormously. A compact content management system of about a million lines, installed on hundreds of millions of sites, can carry a potential saving in the millions of tonnes of CO2e for a run costing under two tonnes. A 34 million line browser codebase costs around 50 tonnes to process and still returns a saving thousands of times larger, because of its install base.

So GreenCode runs a pre-assessment before it commits. It estimates the lines to be processed, the likely energy improvement and the size of the deployed base, and compares that against the modelled carbon of the run. Where the return does not justify the inference, the honest answer is not to run it. That check applies to GreenCode's own operation as much as to anything it processes. The resulting saving is measured and reported as part of the reduced carbon footprint of the software treated.

Why this generalises

The same methods matter well beyond a code optimisation pipeline. Small tuned models, retrieval instead of brute-force context, routing to the cheapest capable component and scheduling against the grid are the techniques that make AI viable at the edge, on battery power and on constrained hardware, where there is no option to spend more energy. Work on efficient training, inference and scheduling is one of the deliverables GreenCode will bring out of the project, and it transfers directly to anyone trying to run useful AI inside a fixed energy budget.

To discuss the methods or contribute a use case, get in touch through the contact page.

  • Benefit area Climate
  • Who it is for AI and platform engineering leads
  • Inference energy per token about 0.58 mWh (Google and OpenAI figures)
  • Inference per refactored line about 2 mg CO2e, before context
  • Delivered through Small models, retrieval, routing and scheduling
GreenCode

Want this benefit for your own systems? Talk to the GreenCode team about a trial.

Get in touch