Thursday, September 17, 2026

What is the AI Equivalent of "Mutually Assured Destruction?"

One of the problems with calls to “regulate AI” is that we cannot seem to agree yet on what needs to be done. Much of the discussion seems to center on reining in frontier model capabilities, mostly along the line of “human in the loop” protections. 


But there are all sorts of other possible approaches that do not deal directly with model capabilities but social impact, economic impact, labor and job impact, consumer protection, transparency, forensic chains, import-export controls or prohibited uses. 


As much as we might agree that some of those areas might make sense, precisely how to regulate frontier models to minimize existential threats remains contentious. 


The central concern isn't simply that an AI gives a bad answer. It is that a sufficiently capable system could plan, acquire resources, replicate, manipulate people, conduct cyber operations, develop dangerous technologies, or resist attempts to shut it down faster than humans can respond.


Perhaps all the main suggestions fall into three buckets. First, make the models safer or constrain them. The former attempts to build an AI that does not want to do dangerous things. 


The second approach assumes misalignment can happen and creates safeguards so harm is prevented. 


The third approach aims to prevent sufficiently dangerous capability from being developed or released. 


Perhaps all three approaches assume all developers agree with the goals. Bad actors will still attempt to create weaponized capabilities, we might well assume. 


This is closest to traditional nuclear/biological-weapons governance. Of course, the issue is that no matter what U.S. firms might do, compliance is voluntary. 


Google DeepMind's Frontier Safety Framework explicitly tracks "Critical Capability Levels" and uses early-warning evaluations followed by mitigation when those capabilities emerge. Its current framework explicitly includes the possibility that a misaligned AI could interfere with operators' ability to direct, modify or shut down it.


OpenAI's Preparedness Framework similarly evaluates severe risks and requires safeguards as capabilities increase; its current governance framework covers cyber, CBRN, manipulation and loss of control, together with security, incident response and external expert input.


Anthropic's policies explicitly separate security, safeguards and alignment. Its roadmap includes access controls, red-teaming, threat intelligence, automated attack investigation and research into ensuring models don't autonomously cause harm.


So there is considerable convergence around a defense-in-depth architecture. The most-consequential unresolved question is whether increasing AI capability eventually makes the defensive layers themselves harder to maintain faster than they become more effective. 


If AI becomes substantially better at cybersecurity, persuasion, strategic planning and AI research, then sandboxing, monitoring and alignment progressively harder.


The main point is that it remains unclear what, exactly, we ought to do about model development and regulating it. Simply banning or slowing development might not work unless all developers globally agree to do so, and abide by those agreements. 


Proposed safeguard

How it would work

Threat it is intended to address

Main advantage

Main weakness

Capability evaluations / early-warning tests

Regularly test models for autonomy, cyber, biological, persuasion, self-replication and AI-R&D capabilities before deployment

Detecting when a model crosses a dangerous capability threshold

Relatively concrete and measurable; can trigger additional safeguards before deployment

Tests can be gamed or incomplete; capability can emerge unexpectedly

"Safety case" before deployment

Developer must construct an evidence-based argument that a model's catastrophic risks have been sufficiently mitigated before release

Preventing deployment of inadequately controlled frontier systems

Forces an explicit connection between capabilities, risks and safeguards

Ultimately depends on whether the safety argument is actually convincing

Scalable oversight

Use weaker AI systems, humans, debate, critique, monitoring, etc. to supervise a much more capable model

Humans becoming unable to understand or evaluate what a highly capable AI is doing

Potentially allows humans to supervise systems beyond their own cognitive capabilities

The supervisor itself may fail; "AI supervising AI" can create correlated errors

Constitutional/value alignment

Train the model around explicit principles governing acceptable behavior rather than relying only on individual human feedback

Model pursuing goals inconsistent with human interests

Gives the model a persistent behavioral framework

There is no universally agreed "human constitution"; a sufficiently capable model might learn to circumvent the intended rules

Interpretability

Analyze the model's internal representations and computations to identify deceptive plans, dangerous objectives or other problematic internal states

Hidden goals, deception and strategic behavior

Could eventually allow researchers to see why a model is behaving dangerously rather than merely observing outputs

Current interpretability is far from providing a complete readout of a frontier model's "intentions"

Deception / scheming evaluations

Specifically test whether models behave differently when they know they're being evaluated, conceal objectives, manipulate overseers, or attempt to preserve their ability to act

Loss of control and deceptive alignment

Targets arguably one of the most important existential-risk mechanisms directly

A model could behave safely during testing and differently in deployment

Sandboxing / containment

Run an autonomous model in an isolated computing environment with restricted filesystem, network, credentials and external communication

Preventing an AI from escaping its environment or acquiring resources

Very powerful because it limits what the model can actually do

A sufficiently capable system may find vulnerabilities or persuade humans/operators to grant access

Least-privilege agents

Give an AI only the minimum tools, permissions, accounts, money, API access and computer privileges necessary for a task

Preventing uncontrolled accumulation of power

Simple cybersecurity principle applied to AI

Complex agents may need increasingly broad permissions to be useful

Human approval gates

Require explicit human authorization before consequential actions—sending money, executing code, changing infrastructure, accessing sensitive systems, etc.

Autonomous harmful action

Maintains a human decision point

Humans can become rubber stamps; sufficiently fast attacks may bypass human review

Continuous monitoring

Monitor agent actions, network traffic, tool calls, unusual planning, attempts to circumvent restrictions, etc.

Detecting dangerous autonomous behavior

Can catch behavior that ordinary output filtering misses

Monitoring itself becomes harder as agents become more sophisticated

Automated tripwires / shutdown

If specified dangerous behavior occurs, terminate the process, revoke credentials or isolate the system

Runaway agents

Fast and potentially automatic

A model capable of anticipating the tripwire might avoid triggering it

Weight / model-security controls

Protect model weights with extremely strong cybersecurity and physical security

Preventing theft of a dangerous model and removal of its safeguards

Particularly important because stolen weights could eliminate deployment controls

Extraordinary security is expensive; insiders and sophisticated attackers remain risks

Controlled deployment

Restrict access to particularly capable models to vetted users, trusted environments or particular applications

Misuse and uncontrolled proliferation

Can substantially reduce the number of pathways to catastrophic use

Reduces openness and economic value; eventually difficult to maintain if models proliferate

Rate/capability limits

Limit how much autonomous work an agent can perform, how many instances can run, how quickly it can act, or how much compute it can consume

Rapid escalation and runaway autonomous activity

Creates time for humans to intervene

Limits can become less meaningful as models become more efficient

Secure model provenance / attestation

Cryptographically establish that a deployed model actually corresponds to the tested and approved model weights

Preventing tampering or substitution after safety testing

Addresses the gap between "safe model tested" and "model actually deployed"

Doesn't establish that the original model itself is safe

Multiple independent AI systems

Have independent systems monitor, critique or approve each other's actions

Single-system failure or deception

Avoids relying on one monitor

Several systems can share the same blind spots or collude/coordinate

AI-to-AI containment / watchdogs

A specialized AI continuously searches another AI's behavior for deception, hacking, dangerous planning, etc.

Monitoring increasingly complex agents

Potentially scales better than human monitoring

Creates an adversarial arms race between the agent and watchdog

Limit autonomous self-improvement

Restrict an AI's ability to modify itself, train successor models, access training infrastructure or conduct unrestricted AI R&D

Recursive self-improvement / rapid capability escalation

Directly attacks one of the more extreme loss-of-control scenarios

Difficult to define "self-improvement"; ordinary AI research increasingly involves AI assistance

Compute controls

Monitor or license very large training runs and/or concentrations of advanced compute

Preventing uncontrolled creation of extremely capable systems

Compute is a physical bottleneck and therefore potentially governable

Algorithmic efficiency can reduce the compute threshold; activity can migrate geographically

International controls

Agreements establishing red lines, inspections, incident communication and human control over especially dangerous AI uses

Global arms-race dynamics and military escalation

Addresses the fact that one country cannot solve the problem alone

Verification and geopolitical competition are extraordinarily difficult


The problem with the analogy to biological or nuclear weapons frameworks is trust. 


If we treat loss of control from highly autonomous AI as analogous to proliferation of nuclear or biological weapons, the policy choices might include: 

  • identify dangerous capabilities

  • control access to the enabling resources

  • verify compliance

  • create international mechanisms for transparency and response.


But all that requires international trust and willingness to comply. And that always seems a tough problem to overcome. “Mutually assured destruction,” oddly enough, was what kept the world safe from nuclear warfare. 


The weapons became too dangerous to use, as their party survives the nuclear exchange. What mechanisms can we actually create for AI software, when global trust cannot be relied upon?


No comments:

Lots of AI Regulations are Conceivable; Few Will Address Existential Threats

One problem with calls for “regulating artificial intelligence” is that it is not entirely clear what should be done, especially on the core...