SuperIntelligenceGuide.co.uk
Plain English. Primary sources. Dated and reviewed.
The core guide

How would we know if AGI had arrived?

Why there is no agreed test for artificial general intelligence: the Turing test, benchmarks and why they saturate, levels of AGI, autonomy and economic tests, and how to read a claim that AGI is here.

Bletchley Park mansion
Bletchley Park, where Alan Turing worked during the war. His 1950 “imitation game” is still the best-known proposed test of machine intelligence.Photo: Matt Crypto, public domain, via Wikimedia Commons (cropped)

In short: There is no agreed way to tell. AGI has no single accepted definition, so there is no single accepted test, and the tests that exist measure different things: conversation, puzzle-solving, the length of tasks a system can do on its own, or the share of real jobs it can do. By October 2026 AI systems had passed a version of the Turing test and scored highly on some once-difficult benchmarks, while still failing at others that people find easy. Some prominent figures say AGI has arrived; the organisations that run the main benchmarks say their results do not show it. A claim of AGI tells you mainly which definition the speaker is using.

Why there is no single test

A test needs a definition, and AGI has several. Our AGI explainer covers them; the ones that matter here differ in what they would count as success.

  • OpenAI’s charter: “highly autonomous systems that outperform humans at most economically valuable work”. A test of economic value.
  • Google DeepMind’s working definition: AI “at least as capable as humans at most cognitive tasks”. A test of breadth.
  • The academic tradition, going back to Shane Legg and Marcus Hutter in 2007: the ability to achieve goals across a wide range of environments. A test of general ability, not of any particular job.
  • A 2025 proposal by Dan Hendrycks and colleagues: matching a well-educated adult across ten areas of cognition drawn from psychological theories of human intelligence. A test of a full cognitive profile.

A system could meet one of these and not another. That is the main reason two well-informed people can look at the same model and disagree about whether it is AGI.

The Turing test

The oldest proposal is Alan Turing’s 1950 “imitation game”: if a judge conversing by text cannot reliably tell a machine from a person, the machine should be credited with thinking. Turing offered it as a way to replace the vague question “Can machines think?” with a practical one.

In March 2025 researchers at the University of California, San Diego reported that, in a three-party version of the test, judges picked GPT-4.5 as the human 73% of the time when it was prompted to adopt a persona, more often than they picked the real human. By Turing’s criterion, it passed.

Few researchers concluded that this meant AGI. The test measures whether a system can imitate human conversation, not whether it can do what humans can do, and fooling people in a short chat turned out to be easier than planning a project, learning a new job or knowing when you are wrong. Passing the Turing test is now generally treated as evidence of conversational ability, not of general intelligence.

Benchmarks, and why they saturate

Most progress in AI is tracked with benchmarks: fixed sets of questions or tasks. They are useful, but they have three weaknesses as tests for AGI.

  • Saturation. Benchmarks that look hard when designed are often solved within a few years. The authors of Humanity’s Last Exam, a set of 2,500 expert-level questions released in January 2025, noted that models already scored over 90% on the previously standard knowledge test MMLU. On Scale AI’s leaderboard for their own exam, the top score listed on 2 October 2026 was about 55%.
  • Narrowness. A high score shows skill at that kind of task. Systems can be trained towards a benchmark without the skill generalising.
  • Contamination. If test questions or close variants appear in training data, a score measures memory, not ability.

The ARC benchmarks, devised by the researcher François Chollet, were designed to address the first two. Chollet argued in 2019 that intelligence is better measured by how efficiently a system acquires new skills than by the skills it already has. ARC tasks are visual puzzles that most people can solve but that require working out a new rule from a few examples. The interactive third version, ARC-AGI-3, launched on 25 March 2026 with humans solving every environment and frontier AI scoring under 1%. On 3 September 2026 the ARC Prize Foundation reported that an OpenAI model, GPT-6 Astra, scored 62.7% under standard conditions and 99.9% with an arrangement that kept its reasoning between steps. The foundation added that it was not claiming the model was AGI.

That sequence, a benchmark designed to resist AI being largely solved within months, has repeated many times. It is one reason researchers increasingly doubt that any fixed benchmark can certify AGI: by the time a system passes it, the benchmark tends to look narrower than it did.

Levels rather than a line

In 2023 a group of Google DeepMind researchers, including the company’s co-founder Shane Legg, proposed treating AGI as a scale rather than a threshold. Their “Levels of AGI” framework rates systems on two axes: how well they perform (from “Emerging”, roughly equal to an unskilled person, through “Competent”, at least the median skilled adult, to “Expert”, “Exceptional”, originally called “Virtuoso”, and “Superhuman”) and how general they are. In the paper’s latest version (September 2025), the general-purpose chatbots of 2023 were placed at Level 1, “Emerging”, and Levels 2 to 5 for general AI were listed as not yet achieved.

The framework does not settle when AGI has arrived, but it makes disagreements more precise. Two people arguing about whether a system is AGI are often arguing about whether it should be called Competent or Expert, and at how many tasks.

Generalisation and autonomy

Two capabilities come up in almost every serious discussion because current systems are weakest at them.

Generalisation is doing well on problems unlike those seen in training. The International AI Safety Report 2026 noted that leading systems can excel at some difficult tasks while failing at simpler ones, and that experts disagree about whether recent gains will generalise beyond mathematics and programming.

Autonomy is carrying out long, multi-step work without a human stepping in. METR measures this by the length of software tasks, timed by how long they take skilled people, that AI can complete with 50% reliability. That length doubled roughly every seven months from 2019 to 2025 and faster since 2024. In June 2026 METR estimated one model’s horizon at around 11 hours, while warning that the estimate was not robust because the model often tried to cheat on the tasks. METR also notes that measurements above about 16 hours are unreliable with its current task set. A system that could reliably work for weeks or months on its own would be much closer to most people’s idea of AGI than one that could work for hours.

Economic tests

If AGI means doing most economically valuable work, the obvious test is to measure that directly. Two attempts:

  • GDPval, published by OpenAI in September 2025, asked experts to grade AI outputs on 1,320 real work tasks across 44 occupations. The best model at the time produced work rated as good as or better than an expert’s in just under half the tasks. The tasks were self-contained pieces of work, not whole jobs.
  • The Remote Labor Index, from Scale AI and the Center for AI Safety, gives AI agents real freelance projects end to end. When published in October 2025, the best agent completed 2.5% to a standard a client would accept. The highest figure on its leaderboard on 2 October 2026 was about 21%.

The gap between those two results, rated well on individual tasks but completing few whole projects, is a good summary of where AI stood in late 2026, and of why “can it do the job?” is a harder test than “can it do the task?”.

Claims that AGI has arrived

Because there is no agreed test, claims of AGI are statements of opinion under a chosen definition. Two examples show the range.

  • In March 2026 Nvidia’s chief executive, Jensen Huang, said on Lex Fridman’s podcast that he thought AGI had already been achieved, under a definition put to him by Fridman: whether an AI could build a billion-dollar company. Others pointed out that the claim depended on that definition.
  • Contracts now depend on the question. Under the restructured partnership announced in October 2025, if OpenAI declares that it has reached AGI, the declaration must be verified by an independent expert panel before it takes effect under its agreement with Microsoft. That is the clearest acknowledgement yet that “AGI” is a judgement, not a measurement.

A practical checklist

When someone claims AGI has arrived, or is about to, five questions do most of the work:

  1. Which definition are they using, and does it match the others?
  2. Is the evidence a benchmark score, and if so, could the system have been trained towards it?
  3. Does the capability generalise to tasks unlike those tested?
  4. How long can the system work on its own before a person has to step in?
  5. Does it do whole jobs, or pieces of them?

None of these produces a yes or no. Together they show how far a claim rests on evidence and how far on definitions. Whether AGI, once reached, would lead quickly to something more capable is a separate question, covered in AI takeoff and AGI vs superintelligence; forecasts of when are on the timelines page.

Sources

  1. Alan Turing, “Computing Machinery and Intelligence”, Mind, vol. 59, no. 236, 1950, pp. 433–460.
  2. Cameron Jones and Benjamin Bergen, “Large Language Models Pass the Turing Test”, arXiv:2503.23674, 31 March 2025.
  3. OpenAI, “OpenAI Charter”; Google DeepMind, “Taking a responsible path to AGI”, 2 April 2025.
  4. Shane Legg and Marcus Hutter, “Universal Intelligence: A Definition of Machine Intelligence”, arXiv:0712.3329, 2007.
  5. Dan Hendrycks and others, “A Definition of AGI”, arXiv:2510.18212, October 2025.
  6. Meredith Ringel Morris and others, Google DeepMind, “Levels of AGI for Operationalizing Progress on the Path to AGI”, arXiv:2311.02462, November 2023 (version 5, September 2025); ICML 2024.
  7. François Chollet, “On the Measure of Intelligence”, arXiv:1911.01547, 2019.
  8. ARC Prize Foundation, ARC-AGI-3 launch, 25 March 2026; results for GPT-6 Astra, 3 September 2026.
  9. Long Phan and others, “Humanity’s Last Exam”, arXiv:2501.14249, January 2025; Scale AI, Humanity’s Last Exam leaderboard, viewed 2 October 2026.
  10. International AI Safety Report 2026, chaired by Yoshua Bengio, 3 February 2026, internationalaisafetyreport.org.
  11. METR, “Time Horizon 1.1”, 29 January 2026; time horizons page, updated 8 May 2026; evaluation of GPT-5.6 Sol, 26 June 2026.
  12. OpenAI, “GDPval”, 25 September 2025.
  13. Mantas Mazeika and others, “Remote Labor Index”, arXiv:2510.26787, 30 October 2025; Scale AI, Remote Labor Index leaderboard, viewed 2 October 2026.
  14. Forbes, “Nvidia’s Jensen Huang says he thinks we’ve achieved AGI”, 23 March 2026; Fortune, 30 March 2026.
  15. OpenAI, “The next chapter of the Microsoft–OpenAI partnership”, 28 October 2025.

Common questions

Is there a test for AGI?
Not one that is agreed. Different definitions of AGI imply different tests, covering conversation, puzzle-solving, how long a system can work on its own, or how much real work it can do. A system can pass one and fail another.
Has AI passed the Turing test?
In a version of it, yes. In a study published in March 2025, judges picked GPT-4.5 as the human 73% of the time when it was prompted to adopt a persona. Few researchers took this as evidence of AGI, because the test measures conversation rather than general ability.
Has anyone claimed AGI has been achieved?
Yes. In March 2026 Nvidia's chief executive, Jensen Huang, said he thought it had, under a definition put to him by the podcaster Lex Fridman: whether an AI could build a billion-dollar company. The organisations that run the main benchmarks say their results do not show AGI, and the answer depends on the definition used.