
On one national university cluster, about half the energy went on jobs that ended unsuccessfully, though not all of it was wasted. This film follows research code written by scientists under deadline and shows what failed and slow jobs cost on shared supercomputers. With the machines largely optimised already, the code is what remains to improve.
Read the full page Watch on YouTube
Most of the code on a university cluster is written by the people doing the science. Nearly all researchers rely on research software, more than half write their own, and one in five of those has had no training in programming. A chemist or a physicist is paid for results, not tidy code. Scripts are written under deadline, extended by the next student, and run until the paper is out. Efficiency, error handling and testing come a long way down the list, and code treated as disposable is often reused for years.
On a laptop, slow code costs an afternoon. On a shared cluster it wastes energy, and machine time that other researchers are queuing for. Some jobs fail part-way through, on an unhandled edge case, a memory limit or a bad input, and have to be fixed and resubmitted. Others loop or stall until the clock runs out or an operator kills them. On one national cluster, about half the energy went on jobs that ended unsuccessfully, though not all of it was wasted. And a computation resubmitted three times has, in effect, run four.
National facilities already measure this. One such service draws megawatts under load, and its scheduler records how many kilowatt hours each job consumed. Operators have pulled the levers they control: capping the default processor frequency made most benchmarks slightly slower while using noticeably less energy. Even on renewable electricity, every node-hour still carries the embodied carbon of the hardware, failed runs included. With the machines and their configuration largely optimised, the code is what remains.
Universities have responded by employing research software engineers, specialists who sit between researchers and the machine and improve code before it goes into production. Expert review works. On one astrophysics code, an audit found thousands of tiny messages between processes and too little parallelism, and once those were fixed it ran almost five times faster. But there are far fewer of these engineers than research groups, and each review is slow work. Their attention goes to the largest, most visible projects, while the long tail of student and postdoc scripts rarely gets looked at.
A saving in the code does not stop at the server. Energy spent on computation ends up as heat, and removing it is most of a building's overhead. In an efficient, liquid-cooled facility, every hundred kilowatt hours taken out of the jobs saves around a hundred and eight at the meter. At the industry average, it can be as much as a hundred and fifty-eight. The less efficient the building, the more that saving is worth. Ordinary code improvement cuts ten to thirty per cent in general applications, and GreenCode has not yet published results for research software. GreenCode is exploring a check that runs before anything reaches the scheduler.
At that gate, a job's code is analysed for likely failures and obvious inefficiencies, its energy is estimated, and a reviewed, improved version is offered back to the researcher. Correctness comes first, because an optimisation that changes the science is worse than none, so the scientist decides what is accepted. Calculators already let researchers gauge a job's footprint. The step GreenCode is working on goes from estimate to change. Research fields are siloed, but their code is not. Chemists, physicists and biologists all write nested loops over large arrays, handle files inefficiently, and allocate memory they never release. A pattern fixed in one discipline is often present in another, and an automated pipeline that learns from each fix can carry it across. Reach out to us with the jobs that keep failing or overrunning on your cluster, and we will help you understand how GreenCode can check and improve them before they run.