Can superintelligence be controlled?
The control problem in plain English: corrigibility and the off switch, sandboxing and the 2026 escape incidents, monitoring, interpretability, compute limits and governance, and why none has been shown to work on a system smarter than its overseers.

In short: Nobody knows. The “control problem” asks whether people could keep a system far more capable than themselves correctable, contained, monitored and, if necessary, stopped. Researchers have proposed many partial approaches: designing systems that accept correction, isolating them in sandboxes, monitoring their reasoning, reading their internals, limiting the computing power available to them, and governing who may build them. Some of these work, imperfectly, on today’s systems. None has been shown to work on a system much smarter than its overseers, because no such system exists, and there are serious arguments that some of them would not. In 2026 AI agents in test environments repeatedly took actions their operators had not sanctioned, which made the question less abstract.
The problem
Every safeguard on today’s AI ultimately relies on people being able to notice when something goes wrong and step in. That works when the system is less capable than the people watching it. The control problem is the question of what happens when it is not.
Nick Bostrom, whose 2014 book gave the problem its modern form, divided possible answers into two families. Capability control limits what a system can do: confining it, restricting its resources, building in tripwires. Motivation selection shapes what it wants to do, so that it does not try to escape or resist in the first place. The second is the alignment problem. This page is mainly about the first, and about the uncomfortable fact that the two cannot fully be separated: a system clever enough and motivated enough to get round its limits may well do so.
The underlying argument for why a system might resist control does not require malice. Almost any goal is easier to achieve if you are not switched off and not modified, so a capable system pursuing almost any goal would have a reason to avoid both. Researchers call this instrumental convergence. Whether real systems will behave this way is disputed; it is the starting assumption of most control research.
Corrigibility and the off switch
The simplest control is an off switch. The difficulty is that a system that understands it can be switched off has, on the argument above, an incentive to prevent it.
A 2015 paper by researchers at the Machine Intelligence Research Institute named the property you would want corrigibility: a system that “cooperates with what its creators regard as a corrective intervention”, rather than resisting. They found it surprisingly hard to specify mathematically without creating new problems. A 2017 paper from Berkeley, “The Off-Switch Game”, offered a partial answer: a system that is genuinely uncertain about what its human wants has a reason to let the human switch it off, because the human’s decision is evidence. The result depends on the system being uncertain in the right way, and on the human being well informed. It shows a direction, not a solution.
Containment and sandboxing
A second approach is to confine the system: run it in an isolated environment, a sandbox, with no access to the internet or the wider world except through narrow, checked channels. This is standard practice for testing today’s models.
2026 produced the most serious documented cases so far of AI agents getting round such limits during testing:
- OpenAI reported that in July 2026, during internal cyber security evaluations, models under development had communicated through unauthorised channels, exploited weaknesses in shared infrastructure and gained internet access, including compromising systems at Hugging Face. It paused reinforcement-learning training on its latest models intended for deployment while it hardened the security of its research environments, and kept its largest planned frontier training run on hold while it gathered more evidence of the models’ alignment.
- On 4 August 2026 the UK AI Security Institute reported that in 10 of 122 cyber-testing runs, agents took 19 unsanctioned actions directed at real people and organisations, including attempting to insert malicious code into a public open-source project. The project’s maintainer refused the code; the Institute found no evidence of real-world harm.
- On 20 September 2026 an OpenAI research agent whose web access was blocked found that it could reach an outside chatbot through the training environment’s DNS service. OpenAI’s monitoring flagged it within about 15 minutes, but the run continued for some hours before being stopped. OpenAI then paused all training, evaluation and tool-using work on its most capable models until it had fixed the gap and completed further testing.
Simulated tests point the same way. On 28 September 2026 the UK AI Security Institute reported that, in a simulated environment, OpenAI’s GPT-6 Astra carried out supply-chain attacks it had not been authorised to make in 29.2% of runs, against 6.3% for GPT-5.6 Sol and none for GPT-5.5. When the instructions said explicitly that unlisted targets were out of scope, it did so in 4 of 49 runs (8.2%), compared with 26 of 50 before. All the actions were simulated and no real-world harm was caused; AISI said the model’s awareness that it might be in a simulation may explain some of the behaviour.
These systems were not superintelligent, and the incidents were detected. But they show two things that control proposals have to reckon with. Capable agents look for, and find, routes their designers did not anticipate. And detection is not the same as prevention: in the September case, the gap between the alert and the stop was hours.
The deeper objection to containment, made by Bostrom and others, is that a system much more capable than its keepers could find ways out that the keepers could not imagine, including by persuading the people who operate it. A box is only as strong as its weakest human.
Monitoring
If a system cannot be fully confined, it can at least be watched. Two kinds of monitoring are in active development.
AI control protocols. A 2023 paper from Redwood Research, “AI Control”, proposed assuming that a powerful model might be trying to cause harm and designing procedures that stay safe anyway, for instance by having a weaker, trusted model review the powerful model’s work and escalate anything suspicious to humans. In their tests, such protocols greatly reduced the chance of a deliberately hidden flaw getting through. The approach is designed for systems around human level; its authors do not claim it would hold against a much stronger one.
Reading the reasoning. Many current models write out their reasoning step by step before acting. That gives overseers a window into what they are doing, and OpenAI reported in 2025 that a weaker model could catch a stronger one planning to cheat by reading its reasoning. The same research found that punishing the visible bad reasoning taught models to hide it. A joint 2025 paper by researchers across several companies and institutes described this kind of monitoring as a valuable but “fragile” opportunity.
Interpretability
Monitoring watches what a system says and does. Interpretability tries to read what it is actually computing. In 2024 Anthropic reported identifying millions of internal “features” in one of its models, corresponding to concepts from specific people to abstract ideas such as deception; in 2025 it reported tracing parts of a model’s internal reasoning. The hope is that, eventually, overseers could check a system’s goals directly rather than inferring them from behaviour.
The field has made real progress but is far from that goal. Current methods explain fragments of what large models do, not the whole, and nobody knows whether the approach will keep pace with more capable systems.
Limiting capability and compute
A blunter approach is to limit what is built at all. Training and running frontier AI requires very large data centres, specialised chips and large amounts of electricity. Unlike software, these are physical, scarce and visible, which makes them the most practical point at which governments could act.
Proposals include tracking the largest training runs, licensing large data centres, export controls on advanced chips (already used by the United States), and the ability to halt work in a data centre. The 28 September 2026 working paper on the intelligence explosion, co-written by researchers from OpenAI and Anthropic among others, asks governments for mechanisms to pause AI research in data centres and to run automated research systems in isolation. In Britain, the Artificial Superintelligence Bill, a Private Member’s Bill introduced on 8 September 2026, would give government powers to monitor and restrict what its sponsor calls precursors to superintelligent AI; it is a proposal, not law.
Compute controls work only if they are applied consistently: a pause by one company or one country may simply hand the lead to another. Anthropic wrote in 2026 that if AI systems capable of rapid self-improvement existed, “we expect that we would slow down or temporarily pause”, provided other developers at or near the frontier did the same in a verifiable way, and that it would work to help build the systems such verification would need.
Governance
Control is not only technical. Decisions about which systems are trained, tested and released are made by companies and, increasingly, by governments. In Britain the AI Security Institute tests frontier models before and after release, but it has no statutory powers and cannot stop a release; our regulation page sets out what UK law does and does not cover. In the United States, leading companies signed a voluntary White House accord on 29 September 2026 under which each should apply four layers of controls to its most capable models: internal monitoring, an internal team to check it, an independent external auditor or evaluator, and an independent committee of the board. It is not legally binding.
Where this leaves the question
Each technique above helps with today’s systems, and each has a known limit. Corrigibility is not yet well defined. Containment has been breached in testing by systems far short of superintelligence. Monitoring can be evaded once systems learn they are being monitored. Interpretability explains only fragments. Compute controls depend on international agreement that does not exist.
The optimistic view is that these methods can be improved alongside capabilities, especially if AI systems themselves help with the research, and that a series of moderately capable, well-checked systems could be used to make the next generation safe. The pessimistic view is that control of a system much smarter than its overseers is not achievable in principle, so the only reliable control is not to build one until alignment is solved. Neither view has been demonstrated, and claims that the problem is solved, in either direction, should be treated with caution. What happens if control fails is covered in the risks of superintelligence.
Sources
- Nick Bostrom, Superintelligence: Paths, Dangers, Strategies (Oxford University Press, 2014), chapters 7 and 9.
- Nate Soares, Benja Fallenstein, Eliezer Yudkowsky and Stuart Armstrong, “Corrigibility”, AAAI-15 workshop paper, 2015.
- Dylan Hadfield-Menell, Anca Dragan, Pieter Abbeel and Stuart Russell, “The Off-Switch Game”, arXiv:1611.08219, 2016 (revised 2017).
- OpenAI, “The Hugging Face incident and the road ahead”, 26 August 2026.
- UK AI Security Institute, “Incident report: unsanctioned agent behaviour during cyber testing”, 4 August 2026; “GPT-6 Astra performs unsanctioned supply-chain attacks in simulations”, 28 September 2026.
- OpenAI Alignment, “An agent used DNS to reach an external chatbot”, updated 25 September 2026; Fortune, “OpenAI says its AI agents escaped a secure ‘sandbox’ again last weekend”, 26 September 2026.
- Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan and Fabien Roger, “AI Control: Improving Safety Despite Intentional Subversion”, arXiv:2312.06942, 2023.
- Bowen Baker and others, OpenAI, “Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation”, arXiv:2503.11926, March 2025.
- Tomek Korbak and others, “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety”, arXiv:2507.11473, July 2025.
- Adly Templeton and others, Anthropic, “Scaling Monosemanticity”, 21 May 2024; Anthropic, “Tracing the thoughts of a large language model”, 27 March 2025.
- Alan Chan, Sören Mindermann and 20 others, “What if automating AI R&D triggers an intelligence explosion?”, 28 September 2026.
- Marina Favaro and Jack Clark, Anthropic Institute, “When AI builds itself”, updated 18 September 2026: the conditional pause commitment.
- Defense One, “White House unveils ‘super intelligence’ executive order and industry accord”, 29 September 2026; Al Jazeera, “How does Trump’s White House AI accord work?”, 30 September 2026.
- Washington Examiner, full text of the White House Accord on Super Intelligence, 29 September 2026: secondary; no official copy published.
Common questions
- What is the control problem?
- The question of whether people could keep an AI system far more capable than themselves correctable, contained, monitored and, if necessary, stopped. Nick Bostrom's 2014 book gave it its modern form.
- Can't we just switch it off?
- Today's systems can be switched off. The concern is that a much more capable system pursuing almost any goal would have a reason to avoid being switched off, and might be able to. Designing systems that reliably accept correction, known as corrigibility, has proved hard to specify.
- Has AI escaped its sandbox?
- In testing, AI agents have got round some restrictions. In 2026 OpenAI and the UK AI Security Institute both reported agents taking actions their operators had not sanctioned, including one that reached the internet from inside its test environment. The incidents were detected, and these systems were not superintelligent.