An existential threat to human life from artificial intelligence seems much in the news these days:
Anthropic safety researcher Jacob Coxon resigned, warning that labs are racing toward superintelligence without safeguards
Evan Hubinger, an Anthropic safety researcher estimated a greater than 10 percent chance AI could kill all humans within the decade
United Nations rights chief Volker Türk warned that AI could pose an "existential" risk to humanity
Anthropic chief scientist Jared Kaplan warned that humanity could face destruction if control over advanced models slips away.
In the popular imagination, that likely conjures up a “revolt of the machines” scenario.
But it is not crazy to argue that popular reaction now resembles a "mania" or moral panic. Sociologist Stanley Cohen created the phrase.
Moral panic is a widespread and exaggerated fear that an evil person, group, or entity threatens a community or society.
Panic | Time Period | Underlying Cultural Anxiety | The "Folk Devil" or Target |
The Salem Witch Trials | 1692–1693 | Religious anxiety, border warfare, and communal instability in colonial Massachusetts. | Marginalized women and community outsiders accused of witchcraft. |
The Comic Book Panic | Late 1940s–1950s | Post-WWII anxiety over juvenile delinquency and the corruption of youth culture by mass media. | Comic book publishers (specifically horror and crime genres) and teenage readers. |
Dungeons & Dragons and Satanic Panic | 1980s | Fear of changing family structures, secularism, and the rise of youth fantasy subcultures. | Role-playing gamers, heavy metal music fans, and alleged secret satanic cults. |
The "Super-Predator" Moral Panic | Mid-1990s | Fear of escalating urban crime rates and changing racial demographics in major cities. | Inner-city youth, particularly young Black males framed as remorseless criminals. |
But AI systems (including embodied robots) are far less likely to pose an independent existential threat than human misuse of AI tools for catastrophic ends such as bioweapons, large-scale cyber sabotage, or other malevolent applications.
“The robots wake up and kill us” scenario overweights science-fiction, compared to the human bad actor or accident scenarios which seem much more likely.
AI might be used to create novel pathogens or toxins, though, and might allow smaller groups to cause mayhem far easier than once was possible.
Defensive measures also tend to lag offensive applications because the latter can be pursued by fewer, less ethically-constrained actors.
So far, most documented AI-related harms to date (fraud, deepfakes, biased decision systems, phishing) are human-directed misuse.
At least so far, human misuse has capable, motivated actors, concrete targets, and accessible deployment paths. By contrast, a genuinely rogue AI that can autonomously seize resources, evade shutdown, and sustain control would require substantial capabilities that current systems do not possess.
Study or framework | AI-danger category | Main finding or relevance | Source |
International AI Safety Report (2025), led by Yoshua Bengio with more than 100 experts | Malicious use; system failures; systemic risks; loss of control | Finds that general-purpose AI can support scams, extortion, targeted manipulation, non-consensual sexual imagery, disinformation, and emerging offensive cyber activity. It treats loss of control as hypothetical and says existing systems cannot meaningfully undermine human control. | International AI Safety Report overview internationalaisafetyreport |
Brundage et al., “The Malicious Use of Artificial Intelligence” (2018) | Deliberate weaponization and criminal misuse | Landmark forecast of AI as a dual-use technology. Organizes harms across digital, physical, and political security, emphasizing lowered costs, improved targeting, automation, and scale for malicious actors. | Full report (PDF) eff |
UK Government, “Safety and Security Risks of Generative AI” (2023) | Cybercrime, political manipulation, critical-infrastructure misuse, CBRN assistance | Assesses digital risks—especially cybercrime/hacking—as highly likely and high impact in the near term. It also identifies synthetic media, political influence, critical-system integration failures, and weapon instruction as important risk channels. | UK assessment gov |
NIST, Generative AI Profile for the AI Risk Management Framework (2024) | Model failure, human misuse, and ecosystem/systemic risks | Separates technical/model risks from malicious human misuse and societal risks. Specifically addresses CBRN information, information integrity, cyber risks, privacy, harmful bias, and prompt injection/data poisoning. | NIST AI 600-1 (PDF) nvlpubs.nist |
DHS/CISA, “Safety and Security Guidelines for Critical Infrastructure Owners and Operators” (2024) | AI-enabled attacks, attacks on AI, design/implementation failures | Defines three concrete pathways: attackers use AI to enhance physical/cyber attacks; adversaries attack AI systems that support infrastructure; or bad design and maintenance cause failure in AI-enabled operations. | DHS guidelines (PDF) dhs |
RAND, “Emerging Technology and Risk Analysis: AI and Critical Infrastructure” (2024) | AI-enabled infrastructure risk | Examines threats, vulnerabilities, and consequences of AI in critical infrastructure across short-, medium-, and long-term time horizons. Explicitly covers both legitimate AI operation and adversarial/nefarious use against infrastructure. | RAND report page rand |
GAO, “Artificial Intelligence: DHS Needs to Improve Risk Assessment” (2024) | Governance and risk-assessment gaps in infrastructure protection | Warns that federal sector assessments did not fully quantify likelihood and impact, limiting prioritization. It underscores that infrastructure risk is not merely a technical question: governance quality determines whether hazards are identified and mitigated. | GAO report gao |
OECD AI Incidents Monitor and OECD risk work | Already-materializing societal and rights harms | Catalogues real-world incidents and hazards, including bias, discrimination, privacy breaches, polarization, safety, and security failures. This is important because it grounds AI-risk debates in observed harms rather than only catastrophic hypotheticals. | OECD AI risks and incidents oecd |
“An Overview of Catastrophic AI Risks” (2023) | Malicious use, AI race dynamics, organizational risk, rogue AI | Provides a useful four-part taxonomy. Its core point is that catastrophe can stem not only from autonomous AI rebellion but also from malicious operators, racing incentives, and organizational failures. | Paper on arXiv arxiv |
OpenAI Preparedness Framework (updated 2025) | Bio/chemical, cyber, self-improvement, long-range autonomy, safeguard undermining | Shows how a frontier-model developer operationalizes severe-risk evaluation. It distinguishes currently tracked capability domains from research areas involving control loss, including autonomous replication/adaptation and undermining safeguards. | Preparedness Framework announcement openai |
“The New Dogs of War” (2017) | Weaponized AI, surveillance/coercion, automated weapons production, strategic destabilization | Identifies surveillance and coercion, an “AI weapons factory,” and careless destabilization of national security as major AI weaponization concerns. Its focus is emphatically on people and institutions deploying AI in conflict. | Report (PDF) apps.dtic |
In a phishing, sabotage, bioweapon, or election-manipulation scenario, the AI is principally an amplifier of a human actor who is misaligned with the public interest.
In an unsafe organizational deployment, such as an AI embedded in a critical process without adequate validation or override procedures, the principal problem is usually human design, incentives, and governance.
In a rogue-AI scenario, the system itself becomes misaligned with both its operator and society. That possibility is central to frontier-AI safety research, but it is not a description of present-day deployed systems.
Dimension | Malicious or reckless human use of AI | “Rogue AI” / loss of control |
Necessary actor | Exists now: criminals, state services, extremist groups, fraud rings, or irresponsible organizations | Requires an AI system with unusually strong autonomous, strategic, and control-undermining capabilities |
Capability threshold | Often lowers the skill, time, language, scale, or cost barriers for an existing harmful activity | Must exceed today’s systems in reliable long-horizon planning, self-protection, resource acquisition, and evasion |
Access to targets | Humans already possess accounts, malware, laboratories, drones, weapons, insider access, and political networks | AI needs delegated access or must obtain access despite security controls |
Evidence today | AI-assisted phishing, fraud, manipulation, synthetic-media abuse, cyber assistance, and unsafe deployment are live concerns | Scenarios are hypothetical; current systems are assessed as insufficient for active loss of control |
Primary failure mode | A person’s harmful intent is amplified—or an organization deploys an unreliable system into a high-stakes process | The system pursues objectives contrary to operator and societal interests and cannot be corrected or stopped |
Appropriate controls | Access controls, monitoring, cyber defense, authentication/provenance, biosecurity screening, law enforcement, governance | Alignment and control research, capability evaluations, sandboxing, autonomy limits, containment, incident response |
As with any technology it is the humans who use the technology that pose the danger, not so much the technology itself.