As a former software engineer, turned quality assurance practitioner, turned functional safety analyst, and now basking in my retirement, I find myself with both time and reason to reflect. And when those reflections overlap with concerns for climate change, there is one inevitable question that keeps cropping up: Do climate modellers apply anywhere near decent standards in their software development? There is both a short and a long answer to this question. The short answer is ‘no’. Here is the long answer.
Some context
Before I get into detail, I should point out that the problem of inadequate standards is by no means restricted to climate science. It is, in fact, applicable to most scientific research software, and it stems from the attitudes and priorities that exist within the scientific community as a whole. The problem is so well-known that its impact has been allocated its own jargon; it is known as technical debt.
In due course I will fully explain what is meant by technical debt, but first I would like to illustrate by offering a notorious example. And in order to emphasise that technical debt is not a problem exclusive to climate scientists, the example I offer comes from the world of epidemiology.
During the Covid pandemic, a great deal of government policy relied upon predictions made by epidemiological models. Specifically, the UK placed great store in a model known as CovidSim, developed by Neil Ferguson of the Imperial College London. However, when the source code for this model was finally made public on GitHub in mid-2020, the wider software engineering community found it to be wanting.
Actually, the technical term used by many of that community was ‘horrendous bag of crap’. Firstly, it was a monolithic file of over 15,000 lines of code, enjoying the structure of a plate of spaghetti bolognese. Upon detailed examination, a total of over 1,200 ‘issues’ were documented, not the least of which was the lack of reproducibility. Basically, different outputs were produced depending upon the platform upon which it was run. The indeterminism isn’t a bug, it’s a feature, opined Ferguson. We don’t care, replied the software engineers; a lack of determinism is the last feature you would want in software that would be used to draft sweeping national laws destined to cause grievous damage to the economy.
In light of these problems, a rescue mission was launched in April and May of 2020. Software engineers from Microsoft and GitHub, and the odd video game programming legend, volunteered to restructure and refactor the code. An independent peer-review group called Codecheck later verified that the heavily cleaned-up, open-source version was now stable and deterministic. Problem solved, but at a not inconsiderable cost when measured in terms of lost reputation and remediation effort. It was a classic example of technical debt.
The point is this: Ferguson could have avoided all of this had he placed greater importance upon developing his code to widely-recognised, professional standards. He could have made it modular, well-documented and thoroughly tested to eradicate bugs. But he didn’t. And the reason was because he was initially developing a tool for his own scientific research. It wasn’t necessary for the software to meet the standards expected of commercial software — only that it should provide him with the wherewithal to pursue his scientific objectives. Such corner-cutting seems perfectly reasonable, until the goalposts shift and public scrutiny, wider applicability, maintainability and high levels of required integrity become important issues. It is at that point that the corner-cutting has to be paid for. That is what is known as technical debt.
As I have said, this potential problem exists throughout the world of scientific software. Scientists are always focussed upon what is required to fulfil the goals of scientific research; all other goals will be an afterthought at best. One can express this in terms of the quality assurance concepts of validation and verification. Scientists are very concerned that their software models should be valid, i.e. serve the required purpose. This can be determined simply by scientifically evaluating the outputs. What they are not so hot on is producing software in a manner that facilitates the verification of its correct development. This is a problem because validation without verification is a recipe for disaster. Without sound verification, the validity can be an illusion.
So, what about climate models?
Climate modelling is not miraculously free from technical debt. On the contrary, the scale of the problem is such that no one involved could dream of ignoring it. This is not the stuff of conspiracy theory; climate scientists and computational researchers have produced numerous peer-reviewed papers, hosted endless workshops, and participated in formal studies that highlight the extent of their technical debt. The problems are not difficult to diagnose.
Firstly, there is the problem of the ‘Fortran legacy’. Climate models were first developed back in the time when Fortran was all the rage (a language invented probably before you were born). To be fair, it was ideally suited to the problem. Consequently, even state-of-the-art global climate models still consist of millions of lines of far from state-of-the-art programming that has been passed down and modified across decades – often involving shortcut hacks. Evidence for this can be found in the so-called Self-Admitted Technical Debt (SATD) that takes the form of thousands of IOU-style code comments along the lines of ‘Sorry, I have put this ugly patch in here because I didn’t have the time or confidence to properly refactor this horrible legacy code’. I sympathise with this because I have done it myself in my own software engineering career. It’s expedient, but it just adds to the debt mountain.
Meanwhile, for the new generation of climate scientists, a program written in Fortran might as well be written in Latin. If they have any software training at all, it would not be on how to produce and maintain millions of lines of opaquely complex and entangled code written in an ancient tongue. They might know better, but they are stuck with what they have got and the technical debt involved in sorting this mess out is not to be underestimated. The community openly discusses the immense difficulty of remediation; as one observer has put it:
Re-writing a codebase with over 9000 commits of Fortran is like tearing down and rebuilding a house. Re-writing a climate model is like rebuilding the entire neighbourhood.
However, let’s face it, millennial climate scientists will not have been plucked from the ranks of fully-trained software engineers, let alone the embattled community of software quality assurance professionals. They will likely be accidental programmers that know just enough to get by whilst they pursue their first love – the science. What use have they for ISO9001 with TickIT when they have a model that is churning out the sort of predictions that their scientific livelihood thrives upon? But as I said earlier, validation without sound verification can trip you up. Many climate models are opaque ‘black boxes’ where it must be hard to isolate whether an unexpected output is a breakthrough in physical theory or a silent floating-point rounding error.
However, before you get too excited, I should point out that none of this is reason enough to dismiss climate models as bug-ridden crocks to be avoided at all costs. In fact, there are good reasons to expect that climate models may have a lower defect density than most other programs. In effect, they are open-source software in which code has been passed about so often, and scrutinised by so many peers, that it becomes difficult for bugs to survive. And you can say what you want about climate modellers, they are rather clever people who are highly motivated to construct their tools properly. Furthermore, it should be acknowledged that there is a lot going on to fix the debt.
Where is the community going from here?
Much as the climate modelling community might wish to continue living off its technical debt, it can’t do so forever. The legacy code was written to execute on your average CPU, but that is not how things have remained. Modern-day climate science modelling demands a lot of compute utilising the parallel processing power of the GPU. The time to call in the debt has arrived as code is re-written to execute on fundamentally different architectures. Basically, climate science is moving house, so it is now or never for getting its house in order. Consequently, a lot of effort is now being put into re-writing software just to deal with the hardware advances, without there being any actual advances in the science. I’m afraid that is what technical debt payback looks like.
A good example of cutting-edge modelling, designed to solve the architectural limits of older models, is the Icosahedral Nonhydrostatic Weather and Climate Model (ICON), developed jointly by the Max Planck Institute for Meteorology (MPI-M) and the German Weather Service (DWD). The model was originally written in Fortran but, instead of just continuing to patch old code, the ICON open-source community joined with Swiss supercomputing centre CSCS to rewrite and port the model’s dynamical core to GPUs using OpenACC directives and Domain-Specific Languages (DSLs). This vastly improved maintainability because the same code can now run on either CPUs or GPUs without requiring scientists to write two entirely separate versions of their physics equations.
Another initiative well worth mentioning is the US National Center for Atmospheric Research (NCAR) Common Infrastructure for Modeling the Earth (CIME) framework. NCAR had a substantial technical debt in the form of the Community Earth System Model (CESM) that glues together massive, independent sub-models (atmosphere, ocean, land, sea ice) written by different scientific teams over 30 years. It consists of millions of lines of mostly Fortran code produced in a variety of styles in accordance with various legacy standards. CIME helps to improve maintainability of this code by separating the ‘scientific code’ from the automation/build code. CIME handles the compilation, testing, and system-level setups, thereby removing the burden of software engineering from the climate scientists and leaving them to deal only with the bits that tickle their scientific fancies.
These and other developments are no doubt helping to pay back the debt, but the scale of the problem is daunting. The sad fact is that, in keeping with most other scientific research programming, climate modelling is where good software development practice went to die. In a scientist’s hands, software is just a tool, not a product to be built to a standard. It is often highly experimental; potentially disposable and hence unworthy of future-proofing. It is said that the proof of the pudding is often in the eating; if it is good enough to generate scientifically plausible results, then it is good enough. Little more is asked of it. Climate model programming followed that groove, but it really shouldn’t have. Climate models are as much a product as they are a tool, and they sit at the centre of political strategies that will cost trillions of dollars, should they prove to be flawed. By running up a technical debt, the community developing these models created a ticking time bomb, destined to go off in their communal face, much like CovidSim did for Professor Ferguson.
When the CovidSim debacle hit the streets, the British Computer Society (BCS) was amongst those that called for any scientific research code that had potential safety and national security ramifications to be written to recognised software development standards. I took the BCS’s concerns with me to an online forum comprised mainly of practicing scientists, congregating at the blog And Then There’s Physics (ATTP). The response was hostile and dismissive: The BCS clearly didn’t know what it was talking about. It was a joke and had no idea what scientific software development entailed. Its call for costly standards was entirely unnecessary and could only do harm. Scientists know what they are doing and software engineers and quality assurance time-wasters should keep their advice to themselves. It’s all there for you to read.
All I can say is that I am pleased to see that the climate scientists themselves seem a lot more cognisant of the technical debt, and far more receptive to the idea of repaying it than are those fellow scientists who would rush to their defence.
Further Reading
The Nature of Technical Debt in Research Software
Multi-Artifact Analysis of Self-Admitted Technical Debt in Scientific Software
Tackling Tech Debt in Scientific Research
Climate Models: Challenges for Fortran Development Tools
Why are Climate models written in programming languages from 1950?
How Climate Model Developers Deal With Bugs
Climate Models are Good Quality Software (?!)
Cracking the code: Linking good modeling and coding practices for new ecological modelers
What are the biggest challenges and innovations for new climate models?
CESM Tutorial – Introduction to CESM2
Operational numerical weather prediction with ICON on GPUs (version 2024.10)
Coding that led to lockdown was ‘totally unreliable’ and a ‘buggy mess’, say experts
The Software that Led to the Lockdown
Codecheck confirms reproducibility of COVID-19 model results
BCS calls for computer coding in scientific research to be more professional