SuperIntelligenceGuide.co.uk
Plain English. Primary sources. Dated and reviewed.
The core guide

Frontier AI safety frameworks: how AI companies decide a model is too dangerous

What a frontier AI safety framework is, how capability thresholds, dangerous-capability tests and safeguards work, when a company says it would slow down, what is voluntary and what is law, and what these documents cannot tell you.

In short: A frontier AI safety framework is a policy, written and published by an AI company, that sets out:

  • which dangerous capabilities it will test its most advanced models for;
  • what level of capability or risk would be too high without extra protection;
  • what it will do when a model gets there.

The best known are Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework and Google DeepMind’s Frontier Safety Framework. Meta, xAI, Microsoft, Amazon and others publish their own.

They share a basic shape: risk thresholds, tests and safeguards. They differ on:

  • what triggers action;
  • whether the company commits to pausing;
  • who decides.

Most began as voluntary pledges. Since 2026, California law has required the largest developers to publish such a framework and follow it, and the EU requires similar risk management. Neither law sets the thresholds for them.

The frameworks have been revised repeatedly as models have improved, including changes to some companies’ earlier commitments to pause or stop. A published framework shows what a company says it will do. It does not show that a model is safe.

What a frontier AI safety framework is

The documents go by different names:

  • Responsible Scaling Policy (Anthropic)
  • Preparedness Framework (OpenAI)
  • Frontier Safety Framework (Google DeepMind)
  • Advanced AI Scaling Framework (Meta)
  • Frontier Artificial Intelligence Framework (xAI)

Most follow the same if-then logic. If a model reaches a defined level of dangerous capability or risk, then the company applies defined protections before it develops the model further or releases it.

They apply only to “frontier” models, the most capable systems a company builds. Their concern is severe, large-scale harm, such as:

  • helping someone make a chemical or biological weapon;
  • serious cyber attacks;
  • AI systems that undermine human control.

They are not general product-safety policies. They do not cover everyday problems such as bias or misinformation, which companies handle separately.

They are also separate from a company’s mission. A framework says how a company will manage risk on the way to whatever it is building:

Why companies wrote them

The first ones. Anthropic published the first, its Responsible Scaling Policy, in September 2023. OpenAI’s Preparedness Framework followed in December 2023, and Google DeepMind’s Frontier Safety Framework in May 2024.

The idea. As capabilities grow, decisions about training and release should be tied in advance to measured risk. They should not be made case by case under commercial pressure.

The AI Seoul Summit, May 2024. The practice spread when 16 companies signed the voluntary Frontier AI Safety Commitments. The list has since grown to 20, including the Chinese developers Zhipu AI and MiniMax. Signatories promised to:

  • publish a safety framework;
  • set out “thresholds at which severe risks … would be deemed intolerable”;
  • in the extreme, “not to develop or deploy a model or system at all” if risks could not be kept below those thresholds.

Who has published one. METR, an independent AI evaluation organisation, tracks published frontier safety policies. Its list covers 12 companies.

Who has not. This site found no comparable published framework from:

How they work: thresholds, tests, safeguards

Capability thresholds

Most frameworks define risk domains, and capability or risk thresholds within them. The domains recur:

  • chemical and biological weapons (some frameworks add radiological and nuclear);
  • cyber attack;
  • AI research and development, meaning AI that speeds up or automates the building of more capable AI;
  • loss of human control, misalignment or harmful manipulation.

Most thresholds are written in words, not numbers.

  • OpenAI sets two levels.
    • “High” means capabilities that “significantly increase existing risk vectors for severe harm”.
    • “Critical” means capabilities that present “a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent”.
    • Its Critical cyber threshold includes a model that can find and develop working exploits for previously unknown vulnerabilities “in many hardened real-world critical systems without human intervention”.
  • Google DeepMind sets “Critical Capability Levels” (CCLs). These are levels at which, “absent mitigation measures”, a model “may pose heightened risk of severe harm”. It also defines lower “Tracked Capability Levels” for significant but not severe harm.
  • Anthropic’s threshold for automated AI research is met if its models could fully substitute for all its research scientists and engineers at competitive cost. It is also met if AI progress accelerates dramatically, which Anthropic defines as roughly double the usual rate, largely because AI is automating the research.
  • Meta sorts models into three risk tiers: moderate or lower, high and critical.
  • xAI’s current framework does not set fixed capability lines. It applies what it calls “systemic risk acceptance criteria” and proceeds only if risks are “determined to be acceptable”. Its 2025 framework had used numbers, for example a model answering fewer than 1 in 20 restricted queries.

The AI-research thresholds link these documents directly to the questions on our pages about recursive self-improvement and the intelligence explosion.

Dangerous-capability evaluations

Companies decide whether a model has crossed a threshold by testing it. The tests include:

  • question sets on weapons-relevant biology;
  • hacking challenges and simulated network attacks;
  • tests of how much of an AI researcher’s job a model can do;
  • tests of whether a model behaves differently when it believes it is being watched.

Most frameworks add an early-warning step. Google DeepMind sets “alert thresholds” “marginally earlier than” each CCL. Crossing one is a sign that the CCL itself may be reached soon, and prompts a response plan.

Outside testers sometimes take part:

  • the UK AI Security Institute;
  • the US Center for AI Standards and Innovation;
  • independent organisations such as METR and Apollo Research.

Access is negotiated with each company and is voluntary, and much of the evidence behind a decision is the company’s own.

The tests have known limits. A model can be more capable than a test reveals, especially if the test does not push it hard or the model behaves differently when evaluated. Benchmark results also do not settle questions of general capability, for reasons set out on our page on recognising AGI.

Safeguards and deployment restrictions

When a threshold is reached, the frameworks call for two kinds of protection:

  • Deployment safeguards stop the dangerous capability being misused. They include refusals, filters that block certain outputs, monitoring and account bans.
  • Security safeguards stop the model’s weights from being stolen, since a stolen model carries none of its developer’s safeguards. Google DeepMind ties its security levels to a scale published by RAND.

In 2026, restricted releases became more common:

  • Anthropic offered Claude Mythos Preview only to organisations that maintain critical software, through Project Glasswing. It said it did not plan to make that model generally available.
  • OpenAI released GPT-6 Astra, which it treats as Critical for cyber security, with cyber safeguards for general users. It is giving vetted defenders phased access through a programme called Daybreak.
  • Google released Gemini 4 Argon first to “trusted cyber defenders”. Those defenders get the model without its cyber guardrails, and wider availability is to follow. Google said Argon’s refusal of harmful requests followed its Frontier Safety Framework. It did not say whether the decision to release first to defenders was taken under the framework.

When would a company slow down or stop?

This is where the frameworks differ most, and where they have changed most.

Explicit halt or pause wording.

  • OpenAI’s framework requires safeguards for a Critical-level model “during development, irrespective of deployment plans”. For each Critical threshold, its guidance is to “halt further development” until it has specified safeguards and security controls that meet a Critical standard.
  • Microsoft’s framework says: “If, during the implementation of this framework, we identify a risk we cannot sufficiently mitigate, we will pause development and deployment until the point at which mitigation practices evolve to meet the risk.”

Narrower wording.

  • Google DeepMind allows a model that has reached a CCL to be deployed externally only once the residual risk is judged acceptable, supported by a safety case. The framework says this requirement covers external deployment, not internal use or further development. It contains no general commitment to pause development.
  • Meta says it will “proceed with development” of a critical-tier model “only if catastrophic risks are within acceptable levels”. Its framework does not use the words “pause” or “stop”.
  • xAI says it will proceed with development or release only if the risks are judged acceptable. Its framework does not use the words “pause” or “halt”.
  • Anthropic has made the largest change.
    • Its 2023 policy committed it to “pause the scaling and/or delay the deployment” of new models whenever it could not keep up with its own safety procedures.
    • Its February 2026 rewrite separates what Anthropic says it will do regardless of others from what it recommends for the whole industry. It states: “we cannot unilaterally and unconditionally commit to staying in line with the industry-wide recommendations”.
    • It says it “would strongly consider pausing development and/or deployment” in some circumstances.
    • Time reported the change as Anthropic dropping its “flagship safety pledge”. Anthropic said government action on AI safety “has moved slowly”.

Competitor clauses. Several frameworks let the company take account of what rivals do.

  • OpenAI may lower its required safeguards if it can rigorously confirm that another developer has released a comparable system without comparable safeguards. It may do so only if:
    • doing so does not meaningfully increase overall risk;
    • it says publicly that it is making the change;
    • it keeps its safeguards more protective than the other developer’s.
  • Google DeepMind says some protections are “most effective when adopted by industry as a whole”. Its risk assessments can consider whether another company’s comparable model is already publicly deployed with weaker protections.

How independent assessors see the changes. In its summer 2026 AI Safety Index, the Future of Life Institute said Anthropic, OpenAI, Google DeepMind and Meta had “weakened or voided” earlier pledges to pause unilaterally if red lines were approached. That is the institute’s own assessment. The companies say their frameworks change as understanding of the risks improves.

Who decides, and who checks

In every framework reviewed here, the final decision is made inside the company.

  • OpenAI. A Safety Advisory Group of staff reviews the evidence and makes recommendations. The chief executive, or someone they designate, decides. A safety and security committee of the board can review decisions and the board can reverse them.
  • Anthropic. Each risk report goes to the chief executive and a Responsible Scaling Officer for final approval. Their decisions and the report are shared with the board and the Long-Term Benefit Trust. The Trust can request an external review.
  • Google DeepMind. The framework refers to “appropriate corporate governance bodies” without naming them.
  • Meta. Its Chief AI Officer oversees the process.
  • xAI. It assigns named “risk owners”.

Independent checking is partial.

  • Outside evaluators test some models before release, by agreement.
  • Anthropic says it will work towards public external review of its risk reports, and requires one in some cases, such as heavily redacted reports on highly capable models.
  • The voluntary White House Accord of 29 September 2026 asks signatories to partner with “an independent external auditor or evaluator”.

None of these frameworks gives an outside body the power to stop a release. In the UK, the AI Security Institute tests models only with companies’ cooperation, and in September 2026 it was reported that a major new model had been kept from it before launch (our page on the institute).

What is voluntary and what is law

InstrumentWhereStatusWhat it requires
Frontier AI Safety Commitments (Seoul)InternationalVoluntaryPublish a framework with thresholds. In the extreme, do not develop or deploy a model at all.
Transparency in Frontier AI Act (SB 53)CaliforniaLaw since 1 January 2026Frontier developers with annual revenue over $500 million must publish a frontier AI framework and comply with it. They must review it at least once a year and publish material changes within 30 days. They must report critical safety incidents to the state within 15 days. Civil penalties are up to $1 million per violation.
RAISE ActNew YorkSigned; in force from 1 January 2027Similar framework requirements, with critical safety incidents reported within 72 hours
EU AI Act and General-Purpose AI Code of PracticeEuropean UnionLaw since August 2025, with fines possible from August 2026. The Code is voluntary.Providers of the most capable general-purpose models must assess and reduce systemic risks and report serious incidents. The Code asks for a written framework. Meta did not sign it, and xAI signed only its safety and security chapter.
White House Accord on Super IntelligenceUnited StatesVoluntaryInternal controls and monitoring, an internal team to check them, an independent external auditor or evaluator, and an independent board committee
United Kingdom—No statutory requirementTesting by the AI Security Institute is voluntary (our regulation page).

What the laws do not do. Neither California nor the EU tells a company where to set its thresholds.

  • California requires a company to follow the framework it has written.
  • It makes materially false or misleading statements about the company’s management of catastrophic risk unlawful.
  • The company still decides what the framework says, and can revise it if it publishes the changes.

Compliance documents. Anthropic and OpenAI now publish separate documents explaining how they comply with these laws:

  • Anthropic’s Frontier Compliance Framework, December 2025;
  • OpenAI’s Frontier Governance Framework, May 2026.

Both say their original policy continues alongside the new document. Anthropic calls the Responsible Scaling Policy its “voluntary safety policy”. OpenAI says the Preparedness Framework “remains the foundation” of its approach.

How the main frameworks differ

OrganisationDocument and versionMain domainsLevelsAt the top levelPause wordingFinal decision
OpenAIPreparedness Framework v2, April 2025 (OpenAI said in August 2026 that it would “evolve” it)Biological and chemical; cyber security; AI self-improvementHigh, CriticalSafeguards required during development; “halt further development” until Critical-standard safeguards are specifiedHalt at Critical until safeguards are specifiedChief executive or delegate, on Safety Advisory Group advice; board can reverse
Google DeepMindFrontier Safety Framework v3.1, April 2026CBRN; cyber; harmful manipulation; ML R&D and misalignmentAlert thresholds; Tracked and Critical Capability LevelsExternal deployment only once residual risk is judged acceptable, with a safety caseApplies to deployment; no general development pauseUnnamed governance bodies
AnthropicResponsible Scaling Policy v3.4, July 2026Chemical and biological weapons; misaligned AI in high-stakes settings; automated R&D in key domains, including AINamed capability thresholds; safeguard sets called AI Safety LevelsCompany plan, plus industry-wide recommendations it does not commit to meeting alone“Strongly consider pausing”; no unconditional commitmentChief executive and Responsible Scaling Officer; board and Long-Term Benefit Trust informed
MetaAdvanced AI Scaling Framework v2, April 2026Cyber; chemical and biological; loss of control (radiological, nuclear and physical autonomy listed as emerging)Moderate or lower, high, criticalDevelopment “only if catastrophic risks are within acceptable levels”None explicitChief AI Officer
xAIFrontier Artificial Intelligence Framework, effective June 2026CBRN; offensive cyber; loss of control; harmful manipulationRisk acceptance criteria; the 2025 version used numeric limitsProceeds only if systemic risks are judged acceptableNone explicitDesignated risk owners

Microsoft and Amazon. Their frameworks cover similar domains.

  • Microsoft commits to pause development and deployment if it identifies a risk it cannot sufficiently mitigate.
  • Amazon’s framework was updated in September 2026. It commits not to deploy a model that meets a Critical Capability Threshold until safeguards appropriately reduce the risk.

What has happened in practice

Thresholds triggered, as the companies describe it.

  • May 2025. Anthropic released Claude Opus 4 under its stricter ASL-3 protections as a “precautionary and provisional action”. It said that clearly ruling out the relevant chemical, biological, radiological and nuclear (CBRN) risks was no longer possible.
  • July 2025. OpenAI treated ChatGPT agent as High capability in biology and chemistry.
  • August 2026. Google DeepMind’s report on Gemini 3.7 Flash found the model above the alert thresholds for CBRN and cyber, but below the critical levels. It judged the model acceptable for deployment.
  • August–September 2026. OpenAI said on 7 August that it “cannot rule out critical cyber capabilities” in its upcoming Astra model. It paused internal work on Astra that did not meet strengthened security controls. Its September system card describes GPT-6 Astra as “our first model to reach the Critical level of cybersecurity capability”. OpenAI released the model publicly with the safeguards and phased access described above. For the later GPT-6.1 Sol, which it also treats as Critical in cyber, OpenAI said its Safety Advisory Group had recommended, and its leadership had determined, that the safeguards were sufficient for public launch.

Revisions and criticism.

  • April 2025. OpenAI moved persuasion risks out of its framework, saying it would handle them through other policies.
  • May 2025. xAI missed its own deadline to publish a revised framework, TechCrunch reported.
  • August 2025. Sixty UK parliamentarians accused Google of breaking the Seoul commitments. Google had published detailed safety information on Gemini 2.5 Pro weeks after releasing it. Google said it was fulfilling its commitments.
  • February and April 2026. Anthropic and Meta rewrote their frameworks.

Incidents outside the framework process. Some of the most serious 2026 incidents happened during testing itself: models under development reached systems they should not have. These are described on our page on whether superintelligence can be controlled.

  • OpenAI’s response included:
    • a two-week pause in some training;
    • keeping its largest planned training run on hold;
    • stricter security.
  • OpenAI said it would “evolve” its Preparedness Framework to cover safeguards across training and deployment.

September 2026. After an internal research agent used its test environment’s DNS service to reach an outside chatbot on 20 September, OpenAI paused “all training, evaluation, and inference with tool-use” of its most capable models. Later that month it cancelled the planned release of GPT-6.1 Astra, telling The Register that the model “performed worse than GPT-6 Astra on alignment evaluations”. OpenAI published the first as a misalignment report and explained the second by alignment results; it did not describe either as a Preparedness Framework threshold decision.

What these documents do not tell you

They do not show that a model is safe. A framework is a plan. Whether it was followed, and whether its tests caught what mattered, are separate questions, and the evidence is mostly the company’s own.

They do not fix the bar. Each company sets its own thresholds, in words that leave room for judgement. It can revise them, and several companies have.

They do not guarantee a stop. Only some frameworks commit to pausing development, and several let competitors’ behaviour change what the company will do.

They do not cover everything. They focus on a few catastrophic risks. Other harms, and risks that no one has thought to test for, fall outside them.

They are not regulation. California and the EU now require frameworks to exist and to be followed, but each company still writes its own. Independent ratings of the frameworks reflect each rater’s own method and vary:

  • in its summer 2026 index, the Future of Life Institute graded the frameworks of the five companies compared here between B- (Anthropic) and D (xAI);
  • in July 2026, SaferAI’s assessment of risk-management practices scored no company above 35%.

The underlying problems these frameworks try to manage are set out on our pages on the risks of superintelligence, alignment and control.

Sources

  1. OpenAI, Preparedness Framework, version 2, 15 April 2025 (PDF); “Our updated Preparedness Framework”, 15 April 2025.
  2. OpenAI, “OpenAI Frontier Governance Framework”, 28 May 2026.
  3. OpenAI, ChatGPT agent system card, 17 July 2025; “Responding to the next frontier of critical cyber capabilities”, 7 August 2026; “Pacing model development in an era of cyber-critical capabilities”, 18 August 2026; GPT-6 Astra system card, 3 September 2026 (since updated); Addendum to GPT-6 Astra system card: GPT-6.1 Sol, 29 September 2026; OpenAI Alignment, “An agent used DNS to reach an external chatbot”, updated 25 September 2026.
  4. Google DeepMind, Frontier Safety Framework, version 3.1, 17 April 2026 (PDF); “Strengthening our Frontier Safety Framework”.
  5. Google DeepMind, Gemini 3.7 Flash Frontier Safety Framework report, August 2026 (PDF); Google, “Gemini 4 Argon: our next era of frontier intelligence”, 30 September 2026.
  6. Anthropic, Responsible Scaling Policy, version 3.4, effective 8 July 2026 (PDF), and update history.
  7. Anthropic, “Responsible Scaling Policy version 3”, 24 February 2026; Responsible Scaling Policy, version 1.0, effective 19 September 2023 (PDF).
  8. Anthropic, “Activating AI Safety Level 3 protections”, 22 May 2025; “Project Glasswing”, 7 April 2026.
  9. Anthropic, “Our compliance framework for California’s Transparency in Frontier AI Act”, 19 December 2025.
  10. Meta, Advanced AI Scaling Framework, version 2, 7 April 2026, and “Scaling how we build and test our most advanced AI”.
  11. xAI, Risk Management Framework, 20 August 2025 (PDF); xAI LLC, Frontier Artificial Intelligence Framework, effective 30 June 2026 (PDF).
  12. Microsoft, Frontier Governance Framework, February 2026 (PDF); Amazon, Frontier Model Safety Framework, February 2025, updated September 2026.
  13. UK Government, Frontier AI Safety Commitments, AI Seoul Summit 2024 (updated 7 February 2025).
  14. California, Transparency in Frontier Artificial Intelligence Act (SB 53), Business and Professions Code §22757.10 onwards; Goodwin, “California Moves to Regulate Frontier AI With a Focus on Catastrophic Risk”, November 2025: secondary.
  15. New York, RAISE Act, General Business Law Article 44-B; Davis Polk, “New York joins California in regulating development of frontier AI models”, May 2026: secondary.
  16. EU AI Act, Article 55; European Commission, General-Purpose AI Code of Practice and signatory list; The Register, “Meta declines to abide by voluntary EU AI safety guidelines”, 18 July 2025: secondary.
  17. White House Accord on Super Intelligence, 29 September 2026 (text via the American Presidency Project).
  18. METR, Frontier AI Safety Policies tracker.
  19. Independent ratings, reflecting each organisation’s own methodology: Future of Life Institute, AI Safety Index, Summer 2026; SaferAI, Frontier Risk Management Tracker, July 2026.
  20. Time, “Exclusive: Anthropic Drops Flagship Safety Pledge”, 24 February 2026; Time, “Exclusive: 60 U.K. Lawmakers Accuse Google of Breaking AI Safety Pledge”, 29 August 2025; TechCrunch, “xAI’s promised safety report is MIA”, 13 May 2025; The Register, report on GPT-6.1 Astra, 29 September 2026, with OpenAI’s statements: secondary.

Common questions

What is a frontier AI safety framework?
It is a policy published by an AI company. It sets out which dangerous capabilities the company tests its most advanced models for, what level would be too risky without extra protection, and what the company will do if a model reaches that level. Examples include Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework and Google DeepMind’s Frontier Safety Framework.
Are AI safety frameworks legally required?
In some places. Since January 2026 California has required the largest frontier developers to publish a framework and follow it. New York’s similar law takes effect in 2027. The EU requires risk management for the most capable general-purpose models. The UK has no such requirement, and the international Seoul commitments are voluntary.
Do these frameworks mean AI companies will stop if a model is too dangerous?
Not necessarily. Some, such as OpenAI’s and Microsoft’s, say development would halt or pause in certain conditions. Others commit only to restricting deployment, or to proceeding only if risks are judged acceptable. Several take account of what competitors do, and some companies revised their stopping commitments in 2026.