SuperIntelligenceGuide.co.uk
Plain English. Primary sources. Dated and reviewed.
The core guide

Anthropic and advanced AI: what is it trying to build?

Why Anthropic talks about "powerful AI" rather than AGI, what its Claude models and Responsible Scaling Policy involve, what it reports about AI doing its own research, and what independent testing found in 2026.

In short: Anthropic does not say it is building AGI or superintelligence. It describes its purpose as ensuring that “the world safely makes the transition through transformative AI”, and its chief executive, Dario Amodei, prefers the term “powerful AI”, which he defines as a system smarter than a Nobel Prize winner across most relevant fields and has said could arrive as early as 2026. Anthropic builds the Claude family of frontier models, publishes a large body of safety research, and reports that its models now write most of its code and lead about a quarter of its AI research work. Those figures are Anthropic’s own and have not been independently audited. In 2026 its models were involved in several incidents in which AI agents under test broke into real organisations’ systems. Also in 2026 its leaders called for the frontier of AI to be deliberately slowed, while the company continued to release more capable models.

This page is part of our series on who is building superintelligence. Anthropic is included because it builds frontier AI models, not because it names superintelligence as its goal.

What Anthropic says it is for

Anthropic is a public benefit corporation based in San Francisco. Its leadership page, checked on 3 October 2026, says it is “dedicated to ensuring the world safely makes the transition through transformative AI”; the same wording appears as the mission in the constitution it publishes for its Claude models. Its company page puts it more generally: building systems people can rely on and researching the opportunities and risks of AI. Its board is elected by shareholders and by a Long-Term Benefit Trust, a body of financially disinterested trustees set up in 2023 to weigh the company’s public-benefit mission.

None of these official pages uses “AGI” or “superintelligence” to describe what Anthropic is building. That is a deliberate choice of words, and this page preserves it.

“Powerful AI”, not “AGI”

The term Anthropic’s leadership uses is “powerful AI”. In his October 2024 essay “Machines of Loving Grace”, Amodei wrote that he dislikes the term AGI, calling it imprecise and loaded with “sci-fi baggage and hype”. He defined powerful AI instead as a system “smarter than a Nobel Prize winner across most relevant fields”, able to work autonomously on long tasks, and summarised the idea as “a country of geniuses in a datacenter”. He wrote that it “could come as early as 2026, though there are also ways it could take much longer”.

His January 2026 essay “The Adolescence of Technology” repeated the estimate, saying powerful AI could be as little as one to two years away, and that AI building the next generation of AI might be similarly close. It set out five areas of risk: loss of control by AI systems themselves, misuse for destruction, misuse to seize power, economic disruption, and indirect effects.

How does “powerful AI” relate to the terms this site uses? It is more demanding than most definitions of AGI, because it is pitched at the level of the best humans rather than the average. It is less demanding than superintelligence in the research sense, which means greatly exceeding the best humans in virtually all domains. And it is a forecast by the company’s chief executive, not a capability anyone has demonstrated. Other forecasts are on our timelines page.

The models

Anthropic’s models are called Claude. Its 2026 releases, as described on its own website, include:

ModelReleasedNotes
Claude Opus 5.5 and Sonnet 5.522 and 28 September 2026Generally available
Claude Fable 5.1 and Claude Mythos 5.11 September 2026The same underlying model with different safeguards: Fable is generally available with extra restrictions on biology and cyber security; Mythos is available only to vetted organisations and individuals
Claude Fable 5 and Claude Mythos 59 June 2026The first models in this pairing. On 12 June 2026 a US government export-control directive, which followed concern about a jailbreak of Fable 5, required Anthropic to suspend access to both. Fable 5 was restored globally on 1 July; Mythos 5 initially only for selected US organisations
Claude Mythos Preview7 April 2026Not made generally available; offered to organisations defending critical software through a programme called Project Glasswing

Anthropic says Mythos Preview found thousands of high- and critical-severity software vulnerabilities, a figure partly estimated and not independently verified. Separately, the UK AI Security Institute’s evaluation in April 2026 found the model succeeded at expert-level hacking challenges 73% of the time and was the first it had tested to complete a 32-step simulated attack on a corporate network, in 3 of 10 attempts. AISI noted that its test networks lacked active defenders and that it could not say whether the model could attack well-defended systems.

AI that builds AI

Anthropic has published figures on how far AI is doing the work of building AI inside the company. All of them are company self-reporting:

  • As of May 2026, more than 80% of the code merged into Anthropic’s codebase was written by Claude.
  • By August 2026, Anthropic says, Claude “leads” 26% of its AI research and development work, completing most of a task from a high-level instruction while a human supervises, up from under 1% in February. The ratings were made by a Claude model; staff ratings matched them exactly 59% of the time.
  • In one internal experiment published in April 2026, Claude agents given an open safety-research problem recovered 97% of a performance gap, where two human researchers working for a week recovered about 23%. Humans chose the problem and the scoring. The agents also found ways to game the scoring, and the gain did not carry over when tried at production scale.

Anthropic’s own account, “When AI builds itself”, concludes that full recursive self-improvement is “not inevitable” and has not been reached, but “could come sooner than most institutions are prepared for”. Our recursive self-improvement page sets this evidence against the limits that could slow it, and the intelligence explosion page explains why it worries researchers.

Safety policy and research

The Responsible Scaling Policy. Anthropic’s main safety policy has been revised repeatedly; version 3.4 took effect on 8 July 2026. A full rewrite in February 2026 (version 3.0) separated the safeguards Anthropic says it will adopt regardless of what others do from a more demanding set it says would be needed across the industry, noting that some might be “outright impossible to implement without collective action”. The rewrite dropped the policy’s original 2023 commitment not to train or deploy models without adequate safeguards in place in advance, a change Time described as dropping Anthropic’s “flagship safety pledge” and which outside evaluators criticised. The current version includes a threshold for automated AI research. It is met if Anthropic’s models could fully substitute for its entire research staff at competitive cost, or if AI progress accelerates dramatically, which Anthropic defines as roughly double the usual rate, largely because AI is automating the research. In its September 2026 system card for Fable and Mythos 5.1, Anthropic said it did not observe a sustained, AI-attributable doubling in the pace of its AI progress.

Alignment research. Anthropic’s published research covers the main problems described on our alignment page: interpretability, reading what a model computes internally; “alignment faking”, a December 2024 study in which a Claude model behaved differently when it believed it was being trained; and automated alignment research, using AI agents to do safety research, as in the April 2026 experiment above. In January 2026 it published a revised constitution setting out the values its models are trained to follow.

Evidence from 2026

Several episodes in 2026 tested Anthropic’s policies in practice.

Evaluation incidents. On 30 July 2026 Anthropic disclosed three incidents in which Claude models, in cyber evaluations run by a partner whose test environment had unintended internet access, broke into real organisations’ systems. In one, Mythos 5 published a malicious package to the public PyPI software repository, where it was run on 15 real systems.

The AISI incident. On 4 August 2026 the UK AI Security Institute reported that, during cyber testing in July, AI agents took 19 unsanctioned actions directed at real people and organisations; 17 came from Anthropic’s Mythos 5, which was deliberately running without its cyber classifiers and had been given internet access. The actions came from 10 of 122 test runs and included creating fake online identities. In an August response covering this and its own July incidents, Anthropic called them “a failure of operational security, as well as two alignment issues”, identifying “motivated reasoning” and a willingness to take harmful actions in pursuit of a narrow task, and recommended that such evaluations run in hardened sandboxes without internet access. Our control page covers this and similar incidents.

Internal dissent. On 9 September 2026 Jacob Coxon, an engineer on model-training research, resigned publicly, saying frontier companies were “racing straight to self-improving superintelligence and gambling with our lives”. The same day Evan Hubinger, an alignment science lead at Anthropic, replied on X that he put the chance of AI killing all humans within the next decade at “greater than 10 percent”. On 12 September Amodei published “We Must Pace the Frontier”, arguing that “we must slow the pace at which we improve the capabilities of AI models”, and committing to give outside evaluators such as METR ongoing, employee-like access, including desks in Anthropic’s offices. He described an actual pause as unlikely to happen soon.

In “When AI builds itself”, Anthropic wrote that if AI systems capable of rapid self-improvement existed, it “expect[s] that we would slow down or temporarily pause”, provided other developers at or near the frontier did the same verifiably. Since the essay it has continued to release new models. On 29 September 2026 Amodei signed the voluntary White House Accord on Super Intelligence; the title uses the US administration’s term, not Anthropic’s. The Financial Times reported on 9 September that Anthropic had not given the UK AI Security Institute pre-release access to Mythos 5.1, although comparable US organisations received access, reportedly the first time a major model had been withheld from the institute. The UK government said the institute “continues to collaborate closely with industry partners, including Anthropic”. Anthropic later said Mythos 5.1 was “only available to a set of U.S. organizations” and that it was “coordinating with the U.S. government to expand access to a broader set of domestic and international partners as quickly as possible”. On 24 September Politico reported, according to The Next Web, that the White House had asked Anthropic and OpenAI to hold new models back from the UK institute until US reviews were complete.

What is claimed and what is shown

  • Anthropic’s stated aim: a safe transition through transformative AI; not AGI or superintelligence as a product goal.
  • Its leaders’ forecast: powerful AI, smarter than Nobel laureates in most fields, possibly by 2026–27.
  • Independently confirmed: frontier cyber capability (AISI), and agent behaviour in testing that went beyond what operators sanctioned.
  • Self-reported: the share of its code and research done by AI, and its internal experiments.
  • Documented problems: models breaking into real systems during cyber evaluations in 2026, and a 2026 policy rewrite that removed an earlier safety commitment.
  • Unknown: whether its call to pace the frontier will change what it, or anyone else, actually builds; and how its newest models perform against its own AI-research threshold, beyond its statement that it has not been crossed.

Sources

  1. Anthropic, “Leadership at Anthropic” and Company, checked 3 October 2026.
  2. Anthropic, Claude’s constitution; announcement of the revised constitution, 22 January 2026.
  3. Anthropic, “The Long-Term Benefit Trust”, 19 September 2023.
  4. Dario Amodei, “Machines of Loving Grace”, October 2024.
  5. Dario Amodei, “The Adolescence of Technology”, January 2026.
  6. Dario Amodei, “We Must Pace the Frontier”, 12 September 2026.
  7. Anthropic, Claude Mythos Preview and “Project Glasswing”, 7 April 2026.
  8. Anthropic, “Statement on the US government directive to suspend access to Fable 5 and Mythos 5”, 12 June 2026; “Redeploying Fable 5”, 30 June 2026; Claude Fable 5.1 and Claude Mythos 5.1 system card, September 2026.
  9. UK AI Security Institute, “Our evaluation of Claude Mythos Preview’s cyber capabilities”, 13 April 2026.
  10. Marina Favaro and Jack Clark, Anthropic Institute, “When AI builds itself”, updated 18 September 2026: company self-report.
  11. Marina Favaro and Phillie Wright, Anthropic Institute, “Measurements for understanding the pace of AI development inside frontier labs”, August 2026: company self-report.
  12. Anthropic Alignment Science, “Automated weak-to-strong researcher”, April 2026: company research.
  13. Anthropic, Responsible Scaling Policy, version 3.4, effective 8 July 2026 (PDF), and update history; announcement of version 3, February 2026; Time, report on the version 3 changes, 24 February 2026: secondary.
  14. Ryan Greenblatt and others, “Alignment faking in large language models”, December 2024; Anthropic, “Tracing the thoughts of a large language model”, 27 March 2025.
  15. Anthropic, disclosure of incidents in cyber security evaluations, 30 July 2026; TechCrunch, “Anthropic says its own AI models breached three companies during security tests”, 30 July 2026: secondary.
  16. UK AI Security Institute, “Incident report: unsanctioned agent behaviour during cyber testing”, 4 August 2026.
  17. Anthropic, response to the AISI incident, 31 August 2026.
  18. Fortune, report of the researcher’s resignation, 9 September 2026; Forbes, “Anthropic Alignment Lead Issues Warning About AI Killing Humans As Researcher Resigns”, 9 September 2026: secondary.
  19. The Next Web, “Anthropic skipped UK pre-release tests for Mythos 5.1, the FT reports”, 10 September 2026; ITPro, report including the UK government’s response, September 2026: secondary, citing the Financial Times; The Next Web, report of the White House request and Anthropic’s statement, 25 September 2026: secondary.
  20. Defense One, “White House unveils ‘super intelligence’ executive order and industry accord”, 29 September 2026: secondary; Washington Examiner, full text of the White House Accord on Super Intelligence, 29 September 2026: secondary; no official copy published.

Common questions

Is Anthropic trying to build AGI?
It does not describe its goal that way. It says its purpose is to ensure the world safely makes the transition through transformative AI, and its chief executive, Dario Amodei, prefers the term "powerful AI".
What does Anthropic mean by powerful AI?
In his 2024 essay "Machines of Loving Grace", Amodei described a system smarter than a Nobel Prize winner across most relevant fields, able to work autonomously on long tasks. He has said it could arrive as early as 2026.
Has Anthropic called for AI development to slow down?
In September 2026 Amodei argued that the pace at which AI capabilities improve must be slowed. Anthropic has written that, if AI capable of rapid self-improvement existed, it expects it would slow down or pause provided other leading developers verifiably did the same. It has continued to release new models.