One of the problems with calls to “regulate AI” is that we cannot seem to agree yet on what needs to be done. Much of the discussion seems to center on reining in frontier model capabilities, mostly along the line of “human in the loop” protections.
But there are all sorts of other possible approaches that do not deal directly with model capabilities but social impact, economic impact, labor and job impact, consumer protection, transparency, forensic chains, import-export controls or prohibited uses.
As much as we might agree that some of those areas might make sense, precisely how to regulate frontier models to minimize existential threats remains contentious.
The central concern isn't simply that an AI gives a bad answer. It is that a sufficiently capable system could plan, acquire resources, replicate, manipulate people, conduct cyber operations, develop dangerous technologies, or resist attempts to shut it down faster than humans can respond.
Perhaps all the main suggestions fall into three buckets. First, make the models safer or constrain them. The former attempts to build an AI that does not want to do dangerous things.
The second approach assumes misalignment can happen and creates safeguards so harm is prevented.
The third approach aims to prevent sufficiently dangerous capability from being developed or released.
Perhaps all three approaches assume all developers agree with the goals. Bad actors will still attempt to create weaponized capabilities, we might well assume.
This is closest to traditional nuclear/biological-weapons governance. Of course, the issue is that no matter what U.S. firms might do, compliance is voluntary.
Google DeepMind's Frontier Safety Framework explicitly tracks "Critical Capability Levels" and uses early-warning evaluations followed by mitigation when those capabilities emerge. Its current framework explicitly includes the possibility that a misaligned AI could interfere with operators' ability to direct, modify or shut down it.
OpenAI's Preparedness Framework similarly evaluates severe risks and requires safeguards as capabilities increase; its current governance framework covers cyber, CBRN, manipulation and loss of control, together with security, incident response and external expert input.
Anthropic's policies explicitly separate security, safeguards and alignment. Its roadmap includes access controls, red-teaming, threat intelligence, automated attack investigation and research into ensuring models don't autonomously cause harm.
So there is considerable convergence around a defense-in-depth architecture. The most-consequential unresolved question is whether increasing AI capability eventually makes the defensive layers themselves harder to maintain faster than they become more effective.
If AI becomes substantially better at cybersecurity, persuasion, strategic planning and AI research, then sandboxing, monitoring and alignment progressively harder.
The main point is that it remains unclear what, exactly, we ought to do about model development and regulating it. Simply banning or slowing development might not work unless all developers globally agree to do so, and abide by those agreements.
The problem with the analogy to biological or nuclear weapons frameworks is trust.
If we treat loss of control from highly autonomous AI as analogous to proliferation of nuclear or biological weapons, the policy choices might include:
identify dangerous capabilities
control access to the enabling resources
verify compliance
create international mechanisms for transparency and response.
But all that requires international trust and willingness to comply. And that always seems a tough problem to overcome. “Mutually assured destruction,” oddly enough, was what kept the world safe from nuclear warfare.
The weapons became too dangerous to use, as their party survives the nuclear exchange. What mechanisms can we actually create for AI software, when global trust cannot be relied upon?
No comments:
Post a Comment