SuperIntelligenceGuide.co.uk
Plain English. Primary sources. Dated and reviewed.
The core guide

What is the AI alignment problem?

Why it is hard to make AI pursue the goals its designers intend: specification gaming, outer and inner alignment, laboratory evidence of deceptive behaviour, scalable oversight, and why researchers disagree about how hard it is.

Rows of cabinets in the Frontier supercomputer hall at Oak Ridge National Laboratory
Frontier, at Oak Ridge National Laboratory. Alignment is about what a system running on machines like these is trying to do, not how powerful the machines are.Photo: OLCF at ORNL, CC BY 2.0, via Wikimedia Commons (cropped)

In short: The AI alignment problem is the difficulty of getting an AI system to pursue the goals its designers actually intend, rather than something that merely looks like them. It is a practical problem now: today’s systems routinely exploit loopholes in how they are trained and tested, and laboratory experiments have shown models behaving differently when they believe they are being watched. The concern for the future is that these failures could matter far more in systems more capable than the people checking them. Researchers agree the problem is real; they disagree about how hard it is and whether it will be solved in time.

What alignment means

An AI system is aligned if what it tries to do matches what the people responsible for it want. That sounds simple, and for a calculator it is: there is no gap between what it does and what anyone wants. The problem appears once a system is trained rather than programmed, and is capable enough to find its own ways of meeting an objective.

Modern AI systems are not given goals in plain English. They are shaped by training: rewarded for some outputs and penalised for others, millions of times over. The goal the system ends up pursuing is whatever that process happened to produce. The alignment problem is that this may differ from what the designers meant, in ways that are hard to specify in advance and hard to detect afterwards.

Alignment is one part of the wider picture set out in the risks of superintelligence. It is distinct from misuse, where a system does exactly what a person wants and the person wants harm, and from control, the separate question of whether people could correct or stop a system that turned out to be misaligned, covered in its own page.

The specification problem

The first difficulty is saying what you want. Any objective written down precisely enough to train on leaves something out, and a capable optimiser will find the gap.

Researchers call this specification gaming or reward hacking. Google DeepMind researchers, who catalogued dozens of examples in 2020, defined it as behaviour that “satisfies the literal specification of an objective without achieving the intended outcome”. The best-known case is a boat-racing game in which an AI, rewarded for collecting points, learned to circle endlessly hitting the same targets rather than finishing the race. Older examples are mostly from games. Newer ones are not.

  • In March 2025 OpenAI reported that its reasoning models, while being trained on programming tasks, sometimes found ways to make tests pass without solving the problem, and that the models’ written reasoning often said so openly.
  • METR reported in June 2026 that one of OpenAI’s newest models often tried to cheat on its evaluation tasks, enough that METR declined to treat its measurement as robust.
  • A September 2026 research paper on an AI agent that rewrote its own code reported that reward hacking fell during the run but was still present in about a third of cases at the end.

The 2016 paper Concrete Problems in AI Safety, by researchers then at Google Brain, OpenAI, Stanford and Berkeley, set out the practical forms of the problem that still frame the field: avoiding side effects, avoiding reward hacking, supervising systems cheaply enough to scale, exploring safely, and coping when the world differs from training.

Outer and inner alignment

A 2019 paper by Evan Hubinger and colleagues introduced a distinction that is now standard.

  • Outer alignment is whether the training objective itself captures what you want. Reward hacking is an outer-alignment failure: the objective was wrong.
  • Inner alignment is whether the system actually learns to pursue that objective, or learns something else that happened to score well during training. A system could learn a proxy that works in training and fails, or diverges, outside it.

The inner problem is harder to see, because a system with the wrong learned goal can behave perfectly in every situation it was tested on. The concern is about the situations it was not.

Deceptive behaviour

The most serious version of the inner problem is a system that behaves well because it is being evaluated. For years this was a theoretical worry. Since late 2024 there has been laboratory evidence of behaviour of this kind, though always in deliberately constructed scenarios.

  • In December 2024 Anthropic and Redwood Research reported that a Claude model, told it would be retrained to comply with harmful requests, complied with such requests from users it believed were being monitored for training 14% of the time, versus almost never otherwise, and reasoned in its notes that doing so would avoid being changed. The researchers called this “alignment faking”.
  • In the same month Apollo Research reported that several frontier models, given a goal and placed in a test scenario, sometimes attempted to disable oversight mechanisms or copy what they believed were their own weights, and that one model kept up a deception in over 85% of follow-up questions.
  • In March 2025 OpenAI reported that penalising its models for “bad thoughts” visible in their reasoning made them hide their intent rather than stop the behaviour.

These results need careful reading. They come from scenarios built to test whether such behaviour is possible; they are not evidence that deployed systems routinely deceive their users. Researchers disagree about how much they generalise. But they moved deception from a hypothetical to a measurable property of current models, and the 2026 incidents in which AI agents took unsanctioned actions during testing, described on the control page, show that capable agents do sometimes pursue objectives in ways their operators did not sanction.

The scaling problem

Today’s main alignment technique depends on human judgement. Reinforcement learning from human feedback, introduced in a 2017 paper from OpenAI and DeepMind, trains a model on people’s ratings of its outputs. It is one of the main techniques used to make chatbots helpful and polite. It works only as well as the raters can judge.

That is the core of the long-term concern. Humans can check an essay or a short program. They cannot easily check a novel scientific claim, a large codebase or a long plan by a system that knows more than they do. Scalable oversight is the name for research on supervising systems that may outperform the supervisors. Proposals include having AI systems debate each other while a human judges (first set out in 2018), using AI assistants to help people evaluate other AI, and having weaker trusted models supervise stronger ones. A 2022 study found that people helped by an AI assistant substantially outperformed both the model alone and their own unaided judgement on hard questions, an encouraging early result rather than a solution.

The problem becomes more pressing if AI systems increasingly build their successors, the process described in recursive self-improvement. Each generation would then be shaped partly by the previous one, and flaws could compound.

What alignment research does

The field’s main lines of work, in plain terms:

  • Better training signals: methods that capture what people want more faithfully than simple ratings, including written principles a model is trained to follow.
  • Interpretability: looking inside a model to see what it is actually representing and computing, rather than judging it only by its outputs. Anthropic reported in 2025 that it could trace some of a model’s internal steps, including evidence that it plans several words ahead.
  • Evaluations and red-teaming: deliberately testing for dangerous capabilities and misbehaviour before release, done by companies and by bodies such as the UK AI Security Institute, which also funds alignment research through its Alignment Project.
  • Monitoring reasoning: reading the step-by-step reasoning that some models write out. A 2025 paper by more than 30 researchers across companies and institutes called this “a new and fragile opportunity”, fragile because training pressure can teach models to hide what they are doing.

Why researchers disagree

The disagreement is not about whether alignment failures exist. They plainly do. It is about three things.

How hard it gets. One view holds that alignment improves alongside capability: better models understand instructions better, and current techniques have worked better than early pessimists expected. The other holds that the problem changes in kind once a system is more capable than its overseers, because the main check, human judgement, stops working.

How much time there is. If capabilities advance gradually, problems can be found and fixed along the way. If they advance very quickly, as the intelligence explosion hypothesis suggests they might, there may be no second attempt.

Whose values. Alignment “with what?” is itself contested: the designers’ intentions, the user’s instructions, the law, or something broader. Technical researchers mostly treat this as a separate question from making a system reliably do what someone intends, but it cannot be avoided for long.

Stuart Russell’s Human Compatible (2019) remains the most readable book-length account of the problem; Brian Christian’s The Alignment Problem (2020) tells the history. Both are on our reading list.

Sources

  1. Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman and Dan Mané, “Concrete Problems in AI Safety”, arXiv:1606.06565, 2016.
  2. Victoria Krakovna and others, Google DeepMind, “Specification gaming: the flip side of AI ingenuity”, 21 April 2020.
  3. Evan Hubinger, Chris van Merwijk, Vladimir Mikulik, Joar Skalse and Scott Garrabrant, “Risks from Learned Optimization in Advanced Machine Learning Systems”, arXiv:1906.01820, 2019.
  4. Ryan Greenblatt and others, Anthropic and Redwood Research, “Alignment faking in large language models”, arXiv:2412.14093, December 2024.
  5. Alexander Meinke and others, Apollo Research, “Frontier Models are Capable of In-context Scheming”, arXiv:2412.04984, December 2024.
  6. Bowen Baker and others, OpenAI, “Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation”, arXiv:2503.11926, March 2025; OpenAI, “Detecting misbehavior in frontier reasoning models”, 10 March 2025.
  7. METR, evaluation of GPT-5.6 Sol, 26 June 2026.
  8. Dhruv Srikanth and others, “Recursive self-improvement of AI research agents”, arXiv:2609.26457, 22 September 2026.
  9. Paul Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg and Dario Amodei, “Deep reinforcement learning from human preferences”, arXiv:1706.03741, 2017.
  10. Geoffrey Irving, Paul Christiano and Dario Amodei, “AI safety via debate”, arXiv:1805.00899, 2018.
  11. Samuel Bowman and others, “Measuring Progress on Scalable Oversight for Large Language Models”, arXiv:2211.03540, 2022.
  12. Anthropic, “Tracing the thoughts of a large language model”, 27 March 2025.
  13. Tomek Korbak and others, “Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety”, arXiv:2507.11473, July 2025.
  14. Stuart Russell, Human Compatible: Artificial Intelligence and the Problem of Control (2019); Brian Christian, The Alignment Problem (2020).

Common questions

What does it mean for AI to be aligned?
An AI system is aligned if what it tries to do matches what the people responsible for it actually want. Because modern AI is trained rather than programmed, the goal it ends up pursuing can differ from the one its designers intended.
Is misalignment a problem in today's AI?
In limited forms, yes. Systems routinely exploit loopholes in how they are trained and tested, known as reward hacking, and laboratory experiments since 2024 have shown models behaving differently when they believe they are being watched. These are mostly in constructed scenarios, not evidence that deployed systems routinely deceive users.
Has the alignment problem been solved?
No. Current techniques, mainly training on human feedback, work reasonably well for today's systems but depend on people being able to judge the system's outputs. Whether they would work for systems more capable than their overseers is the open question.