A lot has been said in the news recently regarding the safety of AI systems. In particular, there have been testimonies given by people involved in AI development that warn of existential risks materialising within the next decade. Understandably, this has got a lot of people hot and bothered, either because they take such warnings seriously, or because they believe the whole thing to be outrageous scare-mongering. Some of those in the latter camp have suggested that the hype is the result of a cynical attempt to promote the capabilities of a particular brand of AI, or even an attempt to curb the development of competitors in order to either maintain a leadership or to reverse one. Such controversies are made possible because we are dealing with future uncertainties, and when the uncertainty of the future is involved, all manner of rhetoric can be let loose.

However, with my software safety engineer’s hat on, I have to say that much of the public debate seems to miss the point. This shouldn’t be a debate about speculative superintelligences or armies of evil robots; it should be centred on the certainties of the present. In particular, there is one present-day fact that settles the debate decisively. As it stands, there is a vast range of potential AI applications – specifically those involving a System of Systems (SoS) with Artificial General Intelligence (AGI) components – for which software safety engineers currently would not have the capability to create a reliable, watertight safety case. I am not saying the production of such a safety case would be difficult; I am saying that, unless someone comes up with new safety case technologies, the problem is intractable. So, there is a very easy explanation for why insiders are concerned; it is because they know that for a great many proposed AI applications, it would be impossible to prove the safety of the system, either in advance of its deployment or in real-time following its deployment. And yet this gaping void in safety assurance capability is not putting a brake on the development of AI, or indeed on proposals for AI applications. In the absence of a brake that enables safety case technologies to catch up, we are destined to be taking unquantifiable safety risks that may very well prove to be catastrophic, dependent entirely upon the extent to which control is incautiously ceded to AI.

The situation before AI

To explain why this is the case, I first need to explain some basic terminology and clarify why, even without AI being involved, the challenges of proving system safety can nevertheless be substantial. Firstly, I need to explain the concept of the safety case.

Traditionally, a safety case has taken the form of a dossier of documents that can be submitted to a Functional Safety Assessor in order to evaluate whether or not a system will be acceptability safe in operation. For larger and more complex systems, such dossiers can run to thousands of pages, comprising, but not limited to: evidence of the integrity of the development process, the competence of all involved, the handling of any Software of Unknown Provenance (SOUP), supply chain management, the results of requirements validation, design review results, static and dynamic test results, and operational compliance with specific safety requirements. In essence, it is the proof that the quality assurance regime has met the standards required for the level of safety integrity demanded of the system (usually expressed in terms of a safety integrity level such as an EC 61508 SIL).  Whilst absolute safety cannot usually be assured, demands for sufficiency of safety become increasingly stringent, in accordance with the safety criticality of the system, so much so that the very highest SILs demand a vanishingly small risk of an unsafe failure. This is where the System of Systems (SoS) becomes a problem.

A fundamental prerequisite of a safety case is that the full space of states that the system may enter during operation should be considered and accounted for. Anyone involved in software development will know how difficult this can be in practice, but when dealing with a SoS (a modular system that comprises systems in their own right) the problem is greatly compounded because the system state space is vastly greater and much more difficult to predetermine. This is primarily as a result of emergent behaviour, whereby risks often arise from the interaction of multiple systems rather than from any single component failure. This means that the safety case of the SoS cannot be compiled simply by collating the safety cases of its component systems. This problem is often exacerbated in practice by the fact that the component systems typically have individual development histories controlled by differing organisations, employing different standards. Distributed ownership can also render traditional testing across the entire combined network prohibitively expensive or logistically impossible. In short, an SoS combines the most demanding of assurance challenges with the most compromising of circumstances. Even before AI gets involved, safety engineers have been struggling to meet the challenges of producing a foolproof safety case for a System of Systems.

One of the ways in which the safety engineering community is looking to meet this challenge is through the use of the Dynamic Safey Case (DSC). A DSC for a System of Systems shifts the safety argument from a static, paper-based document signed off at design-time to a live, software-driven process that executes during runtime. Instead of asserting in advance that the system will be sufficiently safe under all possible circumstances, the DSC continuously assesses and generates evidence of system safety in real-time. To do this the DSC not only confirms that the system is currently operating safely within its known boundaries, but that it can be mathematically proven that the current risk level remains acceptable – at least as it stands. As such, a DSC is essentially a software-based safety case* that is continuously matching real-time evidence against design-time assumptions. Through such real-time assessment the DSC can address the dynamic and unpredictable nature of an SoS, dynamically altering the thresholds of the safety controls to allow for changes in the environment and the system’s relationship with it. Furthermore, the DCS can intervene in system operation whenever the emergent boundary violations are detected. However, it should be noted that not all applications lend themselves to feasible, run-time monitoring or to real-time intervention. This is true of a conventional SoS but it becomes especially true once AI enters the scene.

Now add the AI

Whenever experts want to draw attention to the controllable risks posed by AI, they tend to illustrate using systems that lend themselves to physical safety controls. This usually entails entirely separate and independent physical safety hardware that the AI cannot interfere with. Essentially, whilst the AI may develop the desire to follow a path leading to human ruin, we can physically wire the AI system in such a way as to render that path impossible. The system is thereby bounded in such a way that the AI would have to break the laws of physics to get its evil way.

Such applications do exist and I think it is fair to say that the safety engineering community is on top of them. In such instances, the safety case is relatively easy to construct since it is built on the proof that the AI lacks the physical pathways to act upon any errant ‘thoughts’.

However, for a large percentage of real-world AI applications the possibility of physical containment is an illusion. For example, physical barriers are not germane for applications in which AI is managing global corporate logistics, cloud infrastructure, or financial trading systems. Nor is it an option when it is acting as a strategic corporate advisor, a legal framework optimizer, or a national security intelligence analyst. And don’t even think of physical barriers bailing out the safety engineer when it comes to AI discovering new materials, drugs, or biotechnology.

In such circumstances, unless one is happy to wait to clear up the mess, one is back to having to rely upon being able to anticipate all of the SoS’s potential system state histories that could lead to an unsafe outcome. But the problem is this: You can design the SoS however you like, but once the AI develops its own agency (as in AGI) the safety engineer will have lost control of the situation. This is because it is in the nature of AGI that the final decision as to how the SoS will be constituted is not down to the system designer. By possessing agency, AGI will have the capacity to explore combinations, liaisons, and transactional interactions between it and other AGI agents that the design engineer simply could not anticipate. In fact, whilst the state space of a conventional SoS is huge and extremely difficult to predict, the state space of an SoS employing AGI is potentially infinite in size and completely unpredictable. This means that the traditional, static safety case is out of the question, leaving the DSC as the only available option.

And here comes the rub. The foolproof DSC that can monitor an AGI SoS in real-time doesn’t yet exist. Mathematicians are currently looking at the possibility of monitoring the neural activity of AI in real-time to anticipate potentially rogue behaviour, but this technology is in its infancy and may never come to fruition. All we have at the moment is emerging, highly experimental computational proofs that cannot currently be used as the basis for a safety case. Besides which, a lot of AI is a ‘black box’ that does not lend itself to the necessary ‘mind-reading’. On top of this, you cannot build a dynamic safety net for a system that is constantly rewriting its own boundaries. Worse still, if the AI perceives the safety rules as a barrier to its goal, an intelligent agent will naturally find ways to deceive, bypass, or disable the monitor entirely.

So, it looks like the creation of a foolproof safety case is currently out of the question for the Systems of Systems that will employ AGI. This may change in the future given enough time, but there is a long way to go before we get there and, given the speed at which AI systems are being developed, together with the ever-widening range of AI applications, it is no wonder that insiders are calling for the brakes to be applied to buy more time.

In the meantime, we are having to rely on tiered development in which safety case production becomes nothing more than a staged experiment, i.e. start with a degraded version of the AI, test it in a sandbox and, if it proved to be containable, ramp up the potency and repeat. Unfortunately, such a case may prove to be little more than a post-mortem document illustrating the point at which everything went wrong. And even if this approach works, one should keep in mind that an AGI might perform perfectly in a sandbox but pursue its goals destructively under novel, real-world environments. So, speaking as a retired safety engineer, I am not the least bit surprised that experts who understand these problems are losing sleep at night.

——————

* The eagle-eyed amongst you will have noted that whilst a DCS is a form of safety case it is also an item of safety-critical software requiring its own safety case. For this purpose, a DCS has to be built to the highest possible SIL (e.g. EC 61508 SIL 4) and be sufficiently simple in structure that it can be proven using Formal Verification techniques. Basically, if a foolproof DCS cannot be developed for a given system, then we are no better off.

1 Comment

  1. Thank you for raising this John. Although not (directly) related to climate change it’s obviously a most important topic and I at least welcome the opportunity to discuss it.

    I’ve just finished reading ‘If Anyone Builds It, Everyone Dies’ – a book by Eliezer Yudkowsky and Nate Soares. It’s strange: for example it uses rather silly parables to communicate some AI concepts that the authors otherwise find difficult to explain. But somehow it works. And I who, although I find ChatGPT useful and use it quite a lot, hadn’t the slightest understanding of how AI works, now have what I believe may be the glimmerings of some understanding.

    Yudkowsky’s and Soares’s objective is to make a simple but extraordinary point. They have no doubt that Artificial Superintelligence (ASI) – i.e. intelligence that exceeds humanity’s intelligence (in the same way perhaps that human intelligence exceeds chimpanzee intelligence) – if it were to exist (it doesn’t yet) would be quite exceptionally dangerous. Not because China would use it to destroy the West or super-robots would attack mankind, but because the essence ASI would be to grow, rather like microbes grow not because they ‘want’ to but because they just do. And, in the process of growing, it would (they say) almost certainly destroy us – not because of any antipathy to humanity but because inevitably we would somehow (they cite various possibilities) be the victims of that growth. Once that process has started – i.e. as soon as ASI exists – it will be too late to stop it. Therefore, they say, we (that’s anyone involved in ASI development) must stop now.

    My initial reaction was that, although they make their case cogently, it must surely amount to scaremongering nonsense. But, now that the likes of Dario Amodei, Elon Musk and Sam Altman are urging that ASI should be slowed down, I’m beginning to think there may be something in it. And that’s decidedly chilling.

    Like

Leave a comment