The moral hazard at the centre of AI risk — and the hard instruments that close it
Those exposed to AI harm can't prevent it; those who can face competitive pressure against caution. Why voluntary self-governance fails, the hard instruments that fix incentives, and how to be ready.
This is the third of four deep-dives behind our summary of the [272-expert AI risk Delphi study]. The first two looked at the risks themselves — the capabilities a bad actor can point at a target and the structural harms no model-level fix can touch. This one is about the finding that ties them together and, in our reading, matters most: a textbook moral hazard sitting at the centre of the whole system.
The study rated who is vulnerable against who is responsible, and the two did not line up. AI users and the general public were judged the most vulnerable, near-unanimously. Responsibility was assigned upstream — to foundation-model developers and to governance actors. The people who bear the harm cannot prevent it, and the people who can prevent it face intense competitive pressure not to. That gap does not close on its own. This post is about the instruments that can close it, and what a financial institution should do while they are still being built.
10-minute read · Updated August 4, 2026
Key takeaways
- The study's central finding is an incentive failure: those most exposed to AI harm cannot prevent it, while those who can — upstream developers — face a competitive gradient that punishes caution and externalises the cost of failure onto the vulnerable.
- Voluntary self-governance is therefore structurally insufficient. Not for lack of goodwill, but because race dynamics and tragedy-of-the-commons pressures make individual restraint individually irrational.
- The study's prescription is hard instruments — strict liability, mandatory insurance, transparency mandates, and monitoring — that force upstream actors to internalise the costs they currently externalise. The legal groundwork is already moving.
- You cannot legislate; you can prepare. The transparency and evidence these instruments demand are exactly what continuous, runtime-generated governance produces as a by-product — so the firms that already document from runtime state are the ones that will absorb the shift cheaply.
The asymmetry, restated as an incentive problem
Strip the study's finding to its economics and it is a familiar shape. A moral hazard exists when the party that can reduce a risk does not bear its cost. Here the party best placed to reduce catastrophic AI risk — through weight security, alignment, and release discipline — is the upstream developer, but the cost of a failure lands on a downstream public that had no lever on any of those decisions. Precaution is expensive to the developer and (in a competitive field) strategically costly, while the failure it prevents is paid for by someone else. Under those incentives, the rational move is to precaution less than is socially optimal, and the actor who precautions least is rewarded with speed.
This is not a claim about anyone's character. It is a claim about the gradient people are standing on. The same gradient shows up in the study as its own top-five risk — competitive dynamics — and it is why the structural risks in our second article are so sticky.
Why voluntary self-governance is not enough
The instinctive fix — ask frontier developers to govern themselves — runs straight into the incentive problem. Voluntary commitments are valuable signals, but a commitment that is individually costly and collectively beneficial is exactly the kind of thing a competitive field erodes: whoever relaxes it first gains an advantage, so the equilibrium drifts toward the floor. This is the tragedy of the commons with model releases in place of grazing land. Good intentions, as we like to put it, do not survive a competitive gradient.
That is why the serious policy conversation has moved from exhortation to mechanism. If the problem is that costs are externalised, the fix is to internalise them — to change the payoff so that precaution becomes the rational choice rather than the noble one.

The hard instruments — and where they already stand
The study points to four instruments, and the legal machinery behind them is further along than most risk committees assume.

Strict liability is the sharpest. Under the EU's revised Product Liability framework, software and AI systems are treated as products, with strict liability applying from late 2026 — meaning a claimant does not have to prove negligence, only defect and harm. That single move reallocates the cost of failure toward the party that put the system on the market. (Notably, the separate, fault-based AI Liability Directive proposed in 2022 was withdrawn in 2025, which concentrates attention on the product-liability route rather than a bespoke AI regime.)
Mandatory insurance is the natural complement: if you are strictly liable for catastrophic harm, you must be able to pay for it, and an insurer pricing that risk becomes a second, market-based supervisor of your safety practices. Academic work on catastrophic liability for frontier AI has begun to formalise how liability and insurance can be combined to price systemic risk rather than merely allocate blame after the fact.
Transparency mandates force the information that safety depends on into the open — disclosure of capabilities, incidents, and controls. And monitoring is the standing counterpart to transparency: continuous supervision rather than point-in-time attestation, for the reason our first article gave — the tests themselves are becoming less trustworthy.
Operationalised: notice what these four have in common. Every one of them runs on evidence. Strict liability turns on what you can show about defect and control. Insurance is priced on what you can demonstrate about your practices. Transparency and monitoring are evidence by definition. An instrument regime is, underneath, a demand for documentation that stands up under scrutiny.

What "internalising the cost" looks like on your side of the fence
Most of our readers are not frontier developers; they are the deployers and institutions caught in the middle, who will feel these instruments as new obligations rather than new powers. The good news is that the work the instruments demand is work worth doing anyway, and it is the work our platform is built around.
We have argued for governance by default — a model registry with owners, challenger models, automatic drift monitoring, policy-bounded agents, and documentation generated from runtime state rather than authored at validation cycles. Read that list again against the instruments above and the fit is exact. Strict liability wants a defensible record of controls; continuous MRM produces one as a by-product of running. Insurance wants evidence of practice; the evidence chain is already attached to every signal. Transparency and monitoring mandates want current, inspectable documentation; runtime-generated documentation is current by construction, because it is regenerated from the system's actual state rather than reconstructed at audit time.

The firms that will absorb the coming instrument regime cheaply are the ones for whom producing this evidence is not a project but a property of how the platform runs. The firms that will struggle are the ones still treating documentation as a periodic deliverable — because a periodic deliverable is, almost by definition, out of date the day a strict-liability claim or a supervisory request arrives.
Frequently asked questions
What is the AI "moral hazard" the study identifies?
The actors most vulnerable to AI harm — users and the general public — are not the ones who can prevent it. Prevention sits upstream with foundation-model developers and governance actors, who face competitive pressure against caution. So the cost of precaution falls on them while the cost of failure lands on the vulnerable, and the incentive is to precaution too little.
Why isn't voluntary self-regulation sufficient?
Because a commitment that is individually costly and collectively beneficial erodes in a competitive field — whoever relaxes it first gains an edge. That is a tragedy of the commons, and it is why the study argues the gap will not close without hard instruments that change the underlying payoff.
What are the "hard instruments"?
Strict liability, mandatory insurance, transparency mandates, and monitoring. Together they force upstream actors to internalise the costs they currently externalise — making precaution the rational choice. The EU's move to treat software and AI as products under strict liability from late 2026 is a concrete example already in motion.
How should a deployer prepare if it can't set the rules?
By building governance that produces evidence continuously rather than periodically. Strict liability, insurance pricing, transparency and monitoring all run on documentation that must be current and inspectable — which is exactly what continuous, runtime-generated model-risk-management produces as a by-product of operating.
Related reading
- Experts put a number on AI catastrophe: the 272-expert Delphi study
- Structural AI risks: the harms no model-level fix can touch
- Governance by default: 9 principles for AI in finance
External references
- EU Product Liability Directive (Directive (EU) 2024/2853)
- Catastrophic Liability: Managing Systemic Risks in Frontier AI Development (arXiv)
- International AI Safety Report 2026 (arXiv)
About the author — Model Risk Management — Deep Finance Analytics. MRM operates the continuous model-risk-management layer — registry, challengers, drift monitoring, and runtime-generated documentation. See the Insights hub for the full archive, or book a discovery call to discuss this post with the team.