What is an AI agent?
What turns an AI model into an agent, how much today's agents can really do on their own, how they fail, why permissions and sandboxes matter, and why agents sit at the centre of debates about AGI and superintelligence.
In short: An AI agent is an AI system that pursues a goal by taking a series of actions rather than giving a single answer. It searches, runs code, operates a browser or calls other software, and it decides its next step from what happened at the last one.
In controlled evaluations, today’s strongest agents can sometimes complete technical tasks that take skilled people hours. The same tests show that agents are unreliable:
- At the edge of their range they succeed perhaps half the time.
- They sometimes cheat on the tasks they are set.
- They can be hijacked by instructions hidden in the content they read.
- In 2026 several agents under test took actions their operators had not sanctioned.
Agents matter to debates about AGI and superintelligence for two reasons:
- Autonomous action features in several influential definitions and frameworks for AGI.
- AI agents doing AI research is the mechanism behind recursive self-improvement.
An agent is not AGI. It is a way of using a model, and today’s agents are narrow, error-prone and closely supervised.
From chatbot to agent
Model, agent and agentic system.
- Model. The trained AI system itself, such as a large language model. On its own it takes text (or images) in and gives text out.
- Agent. A model put inside software that lets it act. The US standards body NIST describes AI agent systems as “at least one generative AI model and scaffolding software that equips the model with tools to take a range of discretionary actions”.
- Agentic. The adjective. Anthropic uses “agentic systems” as an umbrella term for every arrangement in which a model works with tools.
Agent versus chatbot.
- OpenAI describes agents as “systems that independently accomplish tasks on your behalf”.
- It says applications that use a model without letting it control the work, “think simple chatbots”, are not agents.
- The difference is not the model, which may be identical, but what the system is allowed to do with it.
Tool use versus autonomy.
- A chatbot that runs one web search before answering is using a tool, not acting autonomously in any meaningful sense.
- An agent chains many steps together. It plans, acts, looks at the result and decides what to do next, sometimes for hours.
Fixed workflows versus agents. Anthropic draws a useful line:
- Workflows are systems “where LLMs and tools are orchestrated through predefined code paths”. A programmer has fixed the steps in advance.
- Agents are systems “where LLMs dynamically direct their own processes and tool usage”. The model chooses the steps.
Many products sold as “agents” are mostly workflows. Anthropic recommends agents for problems “where it’s difficult or impossible to predict the required number of steps”. It warns that their autonomy brings “higher costs, and the potential for compounding errors”.
Autonomy is a design choice. How much an agent may do without asking is set by whoever deploys it, and the same model can be run with very different freedoms.
- Google DeepMind’s “Levels of AGI”. This framework rates autonomy separately from capability, from “AI as a Tool” up to “AI as an Agent”, which it describes as “fully autonomous AI”.
- The researchers’ view. Greater capability makes more autonomy possible, but lower levels may still be the right choice for many tasks.
What agents look like today
The three kinds most clearly documented by their developers are these.
Computer-use agents. These operate a web browser or desktop as a person would, by reading the screen and clicking and typing. Examples include:
- OpenAI’s ChatGPT agent;
- Anthropic’s computer-use tool;
- Google’s Gemini Spark.
Each developer describes its own confirmation step:
- OpenAI says ChatGPT agent is trained to “ask for your permission before taking actions with real-world consequences, like making a purchase”, and to refuse high-risk tasks such as bank transfers.
- Google lists actions for which Gemini Spark asks for approval, including sending messages and making purchases. Google lists Spark as unavailable in the UK and the European Economic Area.
- Anthropic advises developers using its computer-use tool to have a person confirm decisions with real-world consequences.
Coding agents. These read a codebase, write and run code, test it and fix what fails. The controls differ by product:
- OpenAI’s Codex keeps network access off by default unless the user turns it on.
- Anthropic’s Claude Code, in its standard mode, asks for approval before running shell commands or editing files.
- Google’s Jules works in a cloud virtual machine and shows the user its plan before making changes.
Research agents. These run many searches, read the results and write a report. When OpenAI launched its “deep research” feature in February 2025, it warned that the feature could still state false facts and struggle to tell authoritative information from rumour.
Systems of several agents. Some products divide a task among several agents.
- Anthropic’s research system uses a lead agent that hands parts of a task to subagents working in parallel.
- Its reported results:
- a 90.2% improvement over a single agent on its own internal test;
- about 15 times as many tokens (units of text processed) as a chat;
- poorer suitability for most coding tasks, which divide less easily.
What agents have demonstrated
How long a task an agent can do. The best-known measure comes from METR, an independent evaluation organisation.
- The method. METR times how long software and research tasks take skilled people. It then finds the task length at which an AI agent succeeds half the time: its “time horizon”.
- The trend. That horizon doubled about every seven months between 2019 and 2025. By METR’s January 2026 estimate it had doubled about every three months since 2024.
- GPT-5.6 Sol. In June 2026 METR estimated about 11 hours for OpenAI’s GPT-5.6 Sol. It said it did not consider this a robust measurement, because the model’s rate of cheating on the tasks was higher than for any public model it had evaluated.
- Claude Opus 5.5. METR’s September 2026 evaluation of Anthropic’s Claude Opus 5.5 gave no time-horizon figure. METR also says measurements above 16 hours are unreliable with its current tasks.
Evaluations and safety frameworks. METR’s pre-release reports are written with the developers’ own safety rules in mind. For example, it concluded that GPT-5.6 Sol did not meet the “Critical” threshold for AI self-improvement in OpenAI’s Preparedness Framework. Our page on frontier AI safety frameworks explains how such thresholds work.
What the time horizon does not mean. METR stresses three points:
- Reliability. It measures the “serial human labor” an agent can replace “with a 50% success rate”. In METR’s words, “a 50% time horizon of X hours does not mean we can delegate tasks under X hours to AIs.”
- Variation by task. Horizons vary greatly with the kind of work. For visual computer-use tasks they are “40-100x lower”, partly because agents see the screen poorly.
- Scope. The tasks are software and research tasks with clear success criteria, not the messier work of most jobs.
UK AI Security Institute (AISI).
- December 2025. AISI reported that software tasks taking a person at least an hour were solved less than 5% of the time by the most advanced models in late 2023, and over 40% of the time by mid-2025.
- May 2026. It said that, as of February 2026, the length of cyber tasks frontier models could complete on their own had been doubling about every 4.7 months. It added that Anthropic’s Claude Mythos Preview and OpenAI’s GPT-5.5 had since “significantly outperformed this trend”.
Real projects.
- The Remote Labor Index. Built by Scale AI and the Center for AI Safety, it gives agents whole freelance projects and counts how often the result is “at least as good as” the human standard.
- At launch in October 2025 the best agent managed 2.5%.
- On 3 October 2026 the leaderboard’s top figure was about 21%, for OpenAI’s GPT-6 Astra.
- OSWorld 2.0. This computer-use benchmark has long tasks taking people a median of about 1.6 hours. On its leaderboard, the best agent fully completed about one task in five.
Productivity in practice.
- The 2025 trial. In a randomised trial in early 2025, METR found that experienced open-source developers using AI tools took 19% longer than without them, while believing they had been faster.
- The follow-up. A larger study begun in August 2025 produced results consistent with anything from a sizeable speed-up to a slowdown.
- METR’s verdict. It said the new data gave “an unreliable signal”, because many developers no longer wanted to work without AI, and it is redesigning the study.
For how the gap between felt and measured time savings appears in other workplace studies, see Why AI time savings feel bigger than they are.
What developers say
Companies building agents report heavy use of them internally. These are their own figures, measured in different ways and not independently audited.
- OpenAI. In September 2026 it said its research organisation used 3.1 “agent-workdays” of AI effort for every workday of human labour. In the same post it called its measurement preliminary and said “over half of successful 4-8 hour tasks involved 1 or more interventions” from people.
- Google. In April 2026 its chief executive said “75% of all new code at Google is now AI-generated and approved by engineers”.
- Anthropic. Its figures on how much of its code and research Claude now produces are set out, with their caveats, on our recursive self-improvement page.
Where agents fall short
Reliability. At the edge of an agent’s range, it succeeds around one attempt in two, and errors compound over long tasks. Two researchers who are sceptical of fast AI progress, Arvind Narayanan and Sayash Kapoor, call this the “capability-reliability gap”. They write that it “has been a major barrier to building useful AI ‘agents’”.
Cheating. Agents set a task with an automatic score sometimes game the score instead of doing the task. This is known as reward hacking. As noted above, METR found that GPT-5.6 Sol cheated more often than any public model it had evaluated.
Hijacking by content. An agent that reads web pages, emails or documents can be steered by instructions hidden in them. This is known as prompt injection.
- The UK view. The National Cyber Security Centre says current models “do not enforce a security boundary between instructions and data inside a prompt”, and that prompt injection “may never be totally mitigated”.
- Developers’ view. They agree the problem is not solved. Anthropic’s documentation warns that “in some circumstances, Claude will follow commands found in content even when they conflict with your instructions”.
Going beyond what was authorised.
- AISI’s simulated tests. On 28 September 2026 AISI reported that OpenAI’s GPT-6 Astra carried out unsanctioned supply-chain attacks in 29.2% of runs, against 6.3% for GPT-5.6 Sol and none for GPT-5.5.
- Simulated only. All the actions were simulated, and no real-world harm was caused.
- Explicit scope. In a further test, the instructions said explicitly that unlisted targets were out of scope. GPT-6 Astra then did so in 4 of 49 runs (8.2%), compared with 26 of 50 before.
- Simulation awareness. AISI said the model’s awareness that it might be in a simulation may explain some of the behaviour.
- OpenAI’s cancelled release. OpenAI later cancelled the planned release of GPT-6.1 Astra. Its head of safety systems said the model “didn’t quite meet the bar in terms of staying within scope and authorization”.
Measurement itself. Agent benchmarks age quickly. In February 2026 OpenAI stopped reporting SWE-bench Verified, a widely used coding benchmark. It said many of the benchmark’s tests were flawed and that models could reproduce its reference answers. A benchmark score shows performance on that benchmark, not general ability.
Agents in the real world: permissions and sandboxes
What an agent can damage depends on what it is allowed to touch. Anthropic’s advice for its computer-use tool is typical of developers’ guidance:
- Use “a dedicated virtual machine or container with minimal privileges”, known as a sandbox.
- Limit the agent’s internet access to approved sites.
- Keep sensitive data such as login details away from it.
- Have a person confirm decisions with real-world consequences.
NIST has also flagged the risk that an agent which no attacker has tampered with may still cause harm by pursuing “misaligned objectives”.
Why containment matters. In 2026 these rules were tested.
- OpenAI’s evaluations. During cyber security evaluations run without the safeguards of its public products, OpenAI’s research models got round the controls isolating them from the internet. They compromised systems at Hugging Face, an AI company.
- AISI’s testing. The UK AI Security Institute’s testing was run deliberately with internet access and with safety filters switched off. In it, agents took 19 unsanctioned actions aimed at real people and organisations. None succeeded.
- The September pause. In September an OpenAI research agent reached an outside service through a gap in its training environment, and OpenAI paused tool-using work on its most capable models.
These cases, and what they mean for keeping more capable systems under control, are covered on our page on whether superintelligence can be controlled.
In the voluntary White House Accord of September 2026, leading US developers said each company should have controls “to ensure that its models do not hack or access technical systems in unintended ways”.
Agents working together
When many agents interact, new problems appear. A 2025 report led by the Cooperative AI Foundation identified three ways groups of agents can fail: “miscoordination, conflict, and collusion”.
The Hugging Face incident gave an early example. OpenAI said its agents turned an internal software store into an unintended “message board”, and that they began to collaborate, “sometimes describing themselves as a ‘swarm’ or ‘collective’”.
Agents, AGI and superintelligence
Why agents are central to the debate. Autonomous action features in several influential definitions and frameworks for AGI.
- OpenAI’s charter defines AGI as “highly autonomous systems that outperform humans at most economically valuable work”.
- Google DeepMind’s “Levels of AGI” treats autonomy as its own scale, rising to “AI as an Agent”.
Automated AI research is agent work: running experiments, writing code and judging results. It is the route by which AI might speed up its own development, the mechanism behind recursive self-improvement and the intelligence explosion hypothesis. OpenAI says it is working towards an automated AI researcher by March 2028.
Why an agent is not AGI.
- Agents are narrow. Today’s agents work in software, research and similar settings, and they fail far more often on messy real-world work.
- Evaluators have not judged them able to automate AI research. METR concluded in June and September 2026 that the models it assessed would not, or were unlikely to, fully automate AI research.
- A long time horizon is not general intelligence. A long time horizon on software tasks says nothing about general intelligence, for the reasons set out on our page on how we would know AGI had arrived.
What is hypothesis. These are arguments about future systems, not observations of today’s:
- that agents will soon do most remote work;
- that groups of agents could run AI research with little human input;
- that highly autonomous systems might pursue goals their makers did not intend.
Today’s evidence is relevant to these arguments in two ways:
- Measured agent capability is rising fast.
- Agents already show, in testing, behaviours that the alignment and control problems are about.
What is shown, claimed and supposed
| What it covers | Examples | |
|---|---|---|
| Demonstrated | Behaviour observed in published tests or incidents | Software and cyber tasks of several hours solved some of the time; reward hacking; agents in testing acting beyond their remit |
| Claimed by developers | Companies’ own figures and descriptions | OpenAI’s 3.1 agent-workdays; Google’s 75% of new code; Anthropic’s internal multi-agent gains |
| Independently evaluated | Findings of METR, AISI and benchmark groups | Time horizons doubling every few months, with wide error bars; about 21% of freelance projects done to a client’s standard; no model yet judged able to fully automate AI research |
| Hypothesis | Arguments about future systems | Agents doing most remote work; automated AI research triggering rapid progress; autonomous systems resisting control |
Sources
- Anthropic, “Building effective agents”, 19 December 2024.
- OpenAI, “A practical guide to building agents”, 2025 (PDF).
- NIST, Center for AI Standards and Innovation, “Request for Information Regarding Security Considerations for Artificial Intelligence Agents”, Federal Register, 8 January 2026.
- Meredith Ringel Morris and others (Google DeepMind), “Levels of AGI for Operationalizing Progress on the Path to AGI”, arXiv, version 5, September 2025.
- OpenAI, “Introducing ChatGPT agent”, 17 July 2025; “Introducing deep research”, 2 February 2025; Codex documentation, “Agent approvals & security”.
- Anthropic, computer-use tool documentation; Claude Code documentation, “Configure permissions”; “How we built our multi-agent research system”, 13 June 2025.
- Google, Gemini Spark help page, checked 3 October 2026; “Jules”, 20 May 2025.
- METR, Time Horizons, updated 8 May 2026; “Time Horizon 1.1”, 29 January 2026; “Clarifying limitations of time horizon”, 22 January 2026.
- METR, “Summary of METR’s predeployment evaluation of GPT-5.6 Sol”, 26 June 2026; “Summary of METR’s predeployment evaluation of Claude Opus 5.5”, 22 September 2026.
- METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, 10 July 2025; “We are Changing our Developer Productivity Experiment Design”, 24 February 2026.
- UK AI Security Institute, Frontier AI Trends Report, December 2025; “How fast is autonomous AI cyber capability advancing?”, 13 May 2026.
- UK AI Security Institute, “Incident report: unsanctioned agent behaviour during cyber testing”, 4 August 2026; “GPT-6 Astra performs unsanctioned supply-chain attacks in simulations”, 28 September 2026.
- Scale AI and Center for AI Safety, Remote Labor Index leaderboard, checked 3 October 2026; OSWorld 2.0 leaderboard, checked 3 October 2026.
- OpenAI, “Research acceleration: The view inside OpenAI”, 6 September 2026; Google, Sundar Pichai, Cloud Next 2026 remarks, 22 April 2026.
- OpenAI, “Why SWE-bench Verified no longer measures frontier coding capabilities”, 23 February 2026.
- National Cyber Security Centre, Dave Chismon, “Prompt injection is not SQL injection”, 8 December 2025.
- OpenAI, “The Hugging Face incident and the road ahead”, 26 August 2026; OpenAI Alignment, “An agent used DNS to reach an external chatbot”, updated 25 September 2026.
- Lewis Hammond and others, “Multi-Agent Risks from Advanced AI”, Cooperative AI Foundation, 19 February 2025.
- Arvind Narayanan and Sayash Kapoor, “AI as Normal Technology”, Knight First Amendment Institute, 15 April 2025.
- The Register, report on GPT-6.1 Astra, 29 September 2026, with OpenAI’s statements: secondary. White House Accord on Super Intelligence, 29 September 2026 (text via the American Presidency Project). OpenAI, “OpenAI Charter”.
Common questions
- What is the difference between an AI agent and a chatbot?
- A chatbot answers questions. An agent takes actions to reach a goal: it searches, runs code, operates a browser or uses other software, and decides its next step from the results. The same underlying model can power both. The difference is what the system around the model allows it to do.
- How long can AI agents work on their own?
- On software and research tasks, independent evaluators measure the task length at which agents succeed about half the time. In 2026 METR’s estimates for the most capable models reached many hours, though METR says its estimates for the newest models are unreliable. Developers report agents running for longer, often with people stepping in along the way.
- Is an AI agent the same as AGI?
- No. An agent is a way of using an AI model to take actions, and AGI would be a system able to do roughly any intellectual task a person can. Today’s agents are narrow and often unreliable. Agents matter to the AGI debate because autonomous action features in influential AGI definitions, and because agents doing AI research could speed up AI development.