Research software and the energy of a failed job: poster for the GreenCode film
Video created by explainapaper.com

Research software and the energy of a failed job

On one national university cluster, about half the energy went on jobs that ended unsuccessfully, though not all of it was wasted. This film follows research code written by scientists under deadline and shows what failed and slow jobs cost on shared supercomputers. With the machines largely optimised already, the code is what remains to improve.

Chapters

In this film

  • Why most code on university clusters is written by scientists under deadline, with efficiency low on the list.
  • What a failed or stalled job costs in energy, and in machine time other researchers are queuing for.
  • How national facilities already record each job's energy, and why expert code review cannot reach every group.
  • Why a saving in the code is worth more at the meter once cooling is counted.
  • How GreenCode is exploring a check that improves a job's code before it runs, with the scientist deciding.

Transcript

Written to answer a question

Most of the code on a university cluster is written by the people doing the science. Nearly all researchers rely on research software, more than half write their own, and one in five of those has had no training in programming. A chemist or a physicist is paid for results, not tidy code. Scripts are written under deadline, extended by the next student, and run until the paper is out. Efficiency, error handling and testing come a long way down the list, and code treated as disposable is often reused for years.

What a failed job costs

On a laptop, slow code costs an afternoon. On a shared cluster it wastes energy, and machine time that other researchers are queuing for. Some jobs fail part-way through, on an unhandled edge case, a memory limit or a bad input, and have to be fixed and resubmitted. Others loop or stall until the clock runs out or an operator kills them. On one national cluster, about half the energy went on jobs that ended unsuccessfully, though not all of it was wasted. And a computation resubmitted three times has, in effect, run four.

IT shows up in the accounting

National facilities already measure this. One such service draws megawatts under load, and its scheduler records how many kilowatt hours each job consumed. Operators have pulled the levers they control: capping the default processor frequency made most benchmarks slightly slower while using noticeably less energy. Even on renewable electricity, every node-hour still carries the embodied carbon of the hardware, failed runs included. With the machines and their configuration largely optimised, the code is what remains.

Expert review, slowly

Universities have responded by employing research software engineers, specialists who sit between researchers and the machine and improve code before it goes into production. Expert review works. On one astrophysics code, an audit found thousands of tiny messages between processes and too little parallelism, and once those were fixed it ran almost five times faster. But there are far fewer of these engineers than research groups, and each review is slow work. Their attention goes to the largest, most visible projects, while the long tail of student and postdoc scripts rarely gets looked at.

The cooling arithmetic

A saving in the code does not stop at the server. Energy spent on computation ends up as heat, and removing it is most of a building's overhead. In an efficient, liquid-cooled facility, every hundred kilowatt hours taken out of the jobs saves around a hundred and eight at the meter. At the industry average, it can be as much as a hundred and fifty-eight. The less efficient the building, the more that saving is worth. Ordinary code improvement cuts ten to thirty per cent in general applications, and GreenCode has not yet published results for research software. GreenCode is exploring a check that runs before anything reaches the scheduler.

Before a job is admitted

At that gate, a job's code is analysed for likely failures and obvious inefficiencies, its energy is estimated, and a reviewed, improved version is offered back to the researcher. Correctness comes first, because an optimisation that changes the science is worse than none, so the scientist decides what is accepted. Calculators already let researchers gauge a job's footprint. The step GreenCode is working on goes from estimate to change. Research fields are siloed, but their code is not. Chemists, physicists and biologists all write nested loops over large arrays, handle files inefficiently, and allocate memory they never release. A pattern fixed in one discipline is often present in another, and an automated pipeline that learns from each fix can carry it across. Reach out to us with the jobs that keep failing or overrunning on your cluster, and we will help you understand how GreenCode can check and improve them before they run.

More films

All GreenCode videos »

  • Length 3:53
  • Type Article film
  • Chapters 7
  • Captions English
  • Also on YouTube
GreenCode

Want to see what GreenCode finds in your own code? Talk to the GreenCode team.

Get in touch