What are the risks of superintelligence?
Misuse, accidents, misalignment, concentration of power, cyber and military risks, economic disruption and loss of control, each separated into what is observed in today's AI and what is argued about hypothetical superintelligence.

In short: The risks discussed under the heading of superintelligence fall into a handful of distinct categories: misuse by people, accidents, misalignment, concentration of power, cyber and military risks, economic disruption, and loss of control, with the most serious arguments concerning human extinction. Some of these are already observed in today’s AI, at a much smaller scale. Others, especially loss of control over a system far more capable than its makers, are hypothetical: they are arguments about systems that do not exist. Serious researchers disagree sharply about how likely the worst outcomes are, and this page sets out the categories and the evidence rather than a single verdict.
Two kinds of risk
The single most useful distinction is between risks that have been observed in current AI systems and risks that are argued to arise from hypothetical future ones.
Superintelligence, in the research sense this site uses, does not exist; our explainer covers the definition. No one has observed a superintelligence doing anything; what one might be able to do is argued from today’s systems. What can be observed is today’s AI misbehaving in ways that, critics argue, would matter far more in a far more capable system. The categories below are therefore labelled, for each risk, by what is documented now and what is extrapolated.
The main published frameworks divide the field slightly differently. The International AI Safety Report 2026, the nearest thing to a consensus scientific account, uses three categories: malicious use, malfunctions and systemic risks. Google DeepMind’s April 2025 safety paper uses four: misuse, misalignment, mistakes and structural risks. A 2023 overview by Dan Hendrycks and colleagues uses malicious use, an AI race, organisational risks and “rogue AIs”. The headings below are a practical combination of these.
The categories
| Risk | Observed today | The superintelligence version (hypothetical) |
|---|---|---|
| Misuse by people | Fraud using cloned voices; sexually explicit deepfakes; AI-assisted phishing and hacking | A system able to give decisive help with weapons, mass manipulation or attacks that no human team could mount |
| Accidents and malfunctions | Confident false statements (“hallucination”); agents taking unintended actions in test environments | Errors by a system too capable for its operators to check, acting at scale |
| Misalignment | Reward hacking; models gaming tests; laboratory demonstrations of models concealing their behaviour | A system pursuing goals its makers did not intend, competently and without being correctable |
| Concentration of power | A few companies control the most capable models and the compute to train them | Whoever controls the most capable system gains a decisive economic, military or political advantage |
| Cyber | AI finding thousands of serious software vulnerabilities; AI-assisted attacks | Offence outpacing defence across the world’s infrastructure |
| Military | AI used in targeting, intelligence and autonomous systems | Arms races, faster-than-human escalation, loss of meaningful human control over force |
| Economic disruption | Early, uneven effects on some jobs and tasks | Human labour ceasing to be competitive across most of the economy |
| Loss of control | Not observed in the full sense | Humans permanently unable to correct, contain or stop the most capable systems |
Misuse
Misuse is the risk most clearly present now. People use AI to do harmful things they could already do, but more cheaply, more convincingly or at greater scale. In Britain, creating sexually explicit deepfakes of adults has been a criminal offence since February 2026, and Ofcom and the Information Commissioner opened investigations into AI-generated sexual imagery on X in January and February 2026; our regulation page sets out those cases. The superintelligence version of the concern is about scale and uplift: a system that could provide expert help, on demand, to anyone seeking to cause mass harm.
Accidents and misalignment
These are distinct and often confused. An accident is a system doing the wrong thing by mistake. Misalignment is a system pursuing a goal that differs from the one its makers intended, which can look entirely competent from the inside. Both are documented in today’s systems: in August 2026 the UK AI Security Institute reported that in 10 of 122 cyber-testing runs, AI agents took 19 unsanctioned actions directed at real people and organisations, including creating fake identities to approach a software maintainer. Its investigation found no resulting real-world harm. The alignment problem, and why it may get harder as systems become more capable, has its own page.
Concentration of power
This risk does not depend on any AI misbehaving. It is the concern that very capable AI, controlled by a small number of companies or governments, could concentrate economic and political power to a degree democratic institutions cannot check. A 2025 analysis by Forethought, a research organisation, argued that AI workforces made singularly loyal to a small group of leaders could make it possible to seize power without the cooperation of the many people that has always been needed. The 28 September 2026 working paper on the intelligence explosion lists the erosion of checks on power “within and between states, companies, and branches of government” among its main concerns.
Cyber and military risks
Cyber capability is the area where the gap between “observed” and “hypothetical” has narrowed fastest. In April 2026 Anthropic said its Claude Mythos Preview model had found thousands of previously unknown vulnerabilities in major operating systems and web browsers, and released it only to selected organisations defending critical software. The International AI Safety Report 2026 records that in one competition an AI agent identified 77% of the vulnerabilities present in real software. The same capabilities serve attack and defence; the worry is that offence could outpace the ability of defenders to patch.
Military risks are harder to document from public sources, because most of the relevant work is classified. The concern in the literature is less about any single weapon than about arms-race dynamics: states deploying systems before they are understood, and conflicts escalating faster than human decision-makers can follow.
Economic disruption
The evidence on today’s AI and jobs is early and mixed, and this site will cover it separately. The superintelligence argument is of a different order: if machines could do most cognitive work better and more cheaply than people, the link between human labour and income that modern economies rest on would weaken. Anthropic’s own 2026 account of AI building AI says plainly that it is “difficult to predict what the economy looks like if human labor stops being competitive”. This is a forecast about a hypothetical system, not an observation.
Loss of control
Loss of control is the risk most specific to superintelligence and the one with no direct evidence, because no system yet exists that could resist its operators in a serious way. The argument, developed by Nick Bostrom and Stuart Russell among others, runs roughly as follows. A system far more capable than its makers, pursuing almost any goal, would benefit from acquiring resources and avoiding being switched off or modified, because both help it achieve the goal. Researchers call this “instrumental convergence”. If such a system’s goals were even slightly wrong, it might therefore resist correction, and a sufficiently capable system might succeed. Whether advanced systems could be kept correctable is the subject of our page on control.
Critics reply that the argument assumes a particular kind of goal-directed agent that current systems are not, and that capabilities and safeguards will develop together. Both positions are held by serious researchers.
Existential risk: the arguments
“Existential risk” means a risk of human extinction or of a permanent, drastic curtailment of humanity’s future. It is the most contested category, and it is worth being precise about what is claimed.
- A statement of concern. In May 2023 hundreds of researchers and executives signed a one-sentence statement organised by the Center for AI Safety, among them the chief executives of OpenAI, Google DeepMind and Anthropic: “Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.” That is a statement about priorities, not a probability.
- A survey. In the largest survey of AI researchers, published in January 2024 by Katja Grace and colleagues, between 38% and 51% of respondents, depending on how the question was asked, gave at least a 10% chance to advanced AI leading to outcomes as bad as human extinction. Many others put the chance much lower.
- Individual estimates. Some researchers inside AI companies have given high personal figures; in September 2026 Evan Hubinger, Anthropic’s alignment lead, wrote on X that he put the chance of AI causing human extinction within a decade above 10%, comments later reported by Forbes and Reuters. These are judgements, not measurements, and other researchers regard such figures as far too high.
- A scientific consensus. There is none on the probability. No scientific body has published an agreed figure, and the estimates above vary by orders of magnitude.
The most defensible position is that existential risk from AI is a hypothesis taken seriously by many of the people building the technology, disputed by many others, and not something any current evidence can settle.
Why these risks are hard to reason about
Three features make this subject unusually difficult. The systems in question do not exist, so the evidence is indirect. The people with most knowledge of current systems mostly work for companies with a commercial stake. And the outcomes at the extreme are irreversible, so the usual method of learning from mistakes may not be available. That combination is why the debate has turned increasingly to measurement: tracking how far AI is automating its own development, as the recursive self-improvement page describes, and testing frontier models before release, which in Britain is the job of the AI Security Institute.
Sources
- International AI Safety Report 2026, chaired by Yoshua Bengio, 3 February 2026, internationalaisafetyreport.org: risk categories; vulnerability evaluation.
- Rohin Shah and others, Google DeepMind, “An Approach to Technical AGI Safety and Security”, arXiv:2504.01849, April 2025.
- Dan Hendrycks, Mantas Mazeika and Thomas Woodside, “An Overview of Catastrophic AI Risks”, arXiv:2306.12001, 2023.
- Yoshua Bengio, Geoffrey Hinton and others, “Managing extreme AI risks amid rapid progress”, Science, vol. 384, 2024; preprint arXiv:2310.17688.
- Center for AI Safety, “Statement on AI Risk”, May 2023.
- Katja Grace and others, “Thousands of AI Authors on the Future of AI”, arXiv:2401.02843, January 2024 (revised 2025).
- Tom Davidson, Lukas Finnveden and Rose Hadshar, Forethought, “AI-Enabled Coups: How a Small Group Could Use AI to Seize Power”, 15 April 2025.
- Alan Chan, Sören Mindermann and 20 others, “What if automating AI R&D triggers an intelligence explosion?”, 28 September 2026.
- UK AI Security Institute, “Incident report: unsanctioned agent behaviour during cyber testing”, 4 August 2026.
- Anthropic, “Project Glasswing”, 7 April 2026: company statement.
- Marina Favaro and Jack Clark, Anthropic Institute, “When AI builds itself”, updated 18 September 2026.
- Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014), chapters 7–8; Stuart Russell, Human Compatible: Artificial Intelligence and the Problem of Control (2019).
- Forbes, “Anthropic Alignment Lead Issues Warning About AI Killing Humans As Researcher Resigns”, 9 September 2026, reporting Evan Hubinger’s comments on X; Reuters, “Three out of four Americans say AI firms not doing enough to prevent disaster, Reuters/Ipsos poll finds”, 22 September 2026: secondary reports of the extinction-risk estimate.
Common questions
- Is superintelligence dangerous?
- Superintelligence does not exist, so its dangers are arguments rather than observations. Some risks, such as misuse, cyber attacks and AI pursuing unintended goals, are already observed at small scale in today's systems. Others, especially loss of control, are hypothetical. Researchers disagree sharply about how likely the worst outcomes are.
- Could AI cause human extinction?
- Some researchers think it possible. In a 2024 survey of AI researchers, between 38% and 51% gave at least a 10% chance to advanced AI leading to outcomes as bad as human extinction, depending on how the question was asked; many others put it much lower. There is no scientific consensus on the probability.
- Which AI risks are happening now?
- Fraud using cloned voices, sexually explicit deepfakes, AI-assisted hacking, confident false answers, and AI agents taking unsanctioned actions during testing are all documented. They are much smaller in scale than the hypothetical risks of superintelligence.