SuperIntelligenceGuide.co.uk
Plain English. Primary sources. Dated and reviewed.

Where does AI actually help a small business? What the evidence shows

AI tools are sold to small businesses for almost everything: writing, customer service, bookkeeping, quotes, tenders, legal questions and business advice. This page sets out what the published evidence actually shows for each of those jobs, where AI has been found to make work worse, and where the claims have run ahead of the evidence.

We graded 25 common uses against 59 sources:

  • controlled studies;
  • UK government evaluations;
  • court decisions;
  • regulators’ guidance;
  • independent tests;
  • and, clearly labelled, vendors’ own claims.

Each use gets an assessment and a short “because” explaining it. Studies of any kind can be wrong or outdated. We say how confident we are and what each assessment rests on.

We don’t rank or recommend AI products. Vendors’ claims appear only as labelled examples.

The short answer

The strongest evidence we found is concentrated in a small number of everyday jobs:

  • producing first drafts of routine writing that a person then checks;
  • helping newer staff answer customers when there is a high volume of enquiries.

For many other jobs the evidence is promising but limited, or points both ways. For some it suggests AI can make work worse unless an experienced person checks every result: spreadsheet analysis, legal research with chatbots, tax and bookkeeping done by AI on its own, and advice on hard business decisions.

Across many well-bounded tasks, speed gains are common. Better work is conditional; correct work remains the critical question.

By a well-bounded task we mean one with a clear output that the person can check: a draft letter, a meeting summary, a reply to a standard enquiry. Even then, speed is not guaranteed. In some careful studies people using AI were slower, and some believed they were faster when they were not (see Why AI time savings feel bigger than they are).

For some jobs the honest answer is that AI probably isn’t worth prioritising yet. That is a normal conclusion on this page, not a failure to find something positive.

What we found

We assessed 25 common small-business uses of AI against 59 sources (66 source records, because some sources inform more than one task): controlled studies, UK government evaluations, court decisions, regulators’ guidance, independent tests and labelled vendor claims.

  • Speed gains are common; correctness is the weak point. Across many well-bounded tasks, AI made work faster. Better work was conditional, and when AI was wrong, people often didn’t notice. More
  • Feeling faster isn’t the same as being faster. Experienced developers were 19% slower with AI while believing they were faster. UK civil servants’ own estimates of time saved on data analysis were not borne out when people were observed doing the work. More
  • Who gains depends on what the AI does. Less experienced people often gain most when AI supplies knowledge they lack; the most skilled can get worse when it replaces judgement they already exercise well. (Our interpretation across studies.) More
  • Some headline gains are against no service, not against a person. The best-known AI customer-service sales gain, 16.3%, compared a chatbot with an automated notice that customer service was unavailable. It shows AI beat no answer, not that it beat a person. More
  • Some heavily marketed uses have no independent evidence of results. For AI quoting, AI-written tenders and AI phone receptionists, we found only vendors’ claims. More
  • The strongest evidence is narrow. In our classification, two of the 25 uses meet our Established rule: first drafts of routine writing, and AI suggestions for newer staff in busy support teams. Eight are graded Poor fit / caution. The counts depend on how the uses are divided. More

Is AI worth it for a small business?

The evidence can’t tell you whether AI is worth paying for in your particular business. Most studies measure individual tasks, not return on investment. Whether it pays depends on:

  • how often you do the task;
  • what the time saved is worth to you;
  • what the AI costs;
  • how much checking it creates;
  • whether faster work turns into more capacity or more revenue.

The UK Department for Business and Trade’s own evaluation of an AI office assistant concluded: “We did not find robust evidence to suggest that time savings are leading to improved productivity.” That is why this page tells you what is worth testing and what to measure, rather than giving general return-on-investment estimates.

Where should a small business start with AI?

Start with a frequent, low-risk task where the output is easy to check.

Of the 25 uses we reviewed, routine first drafts have the strongest evidence and apply to the widest range of small businesses. If you have newer staff handling a large volume of similar customer enquiries, AI suggestions for those staff also have strong evidence. Beyond those, test one task at a time, and measure the result before expanding.

Reasonable to try: strongest evidence first

  • First drafts of letters, proposals, policies, reports and routine emails, from your own notes, checked and edited by you. (Established)
  • AI reply suggestions for newer staff, if you have a busy support team handling many similar enquiries. (Established)

Then, with promising but more limited evidence:

  • Newsletters and product descriptions, with a person checking every claim. The evidence shows similar results to staff-written copy at much lower cost, not better results.
  • Meeting notes and internal summaries, checked by someone who was there or knows the material.
  • A website chatbot for enquiries that would otherwise go unanswered, with an easy way to reach a person.
  • AI suggestions inside your accounting software, such as categorising bank transactions, reviewed by whoever signs off the books.
  • Brainstorming: options, ideas and objections to your own plan, especially if you have nobody to think aloud with.
  • For businesses that employ developers: routine and well-defined coding work, with review before anything goes live.

How to tell whether AI is working in your business

Don’t ask whether AI feels faster. Measure it. A simple test is enough: time the same kind of task with and without AI for a few weeks, and count the corrections. What to measure:

  • Measured time, not how fast it feels. In two of the most careful studies, people’s sense of time saved was wrong (see Why AI time savings feel bigger than they are).
  • Whether the problem was actually solved, not just how fast the reply was. In one large customer-service experiment, AI help made chats faster and customer ratings slightly higher. It made no difference to how many customers had to come back about the same problem.
  • What actually ships, sells or gets paid, not how much is produced. One large study of software developers found that AI coding tools sharply increased the amount of code written, while the number of finished software releases rose about 30%.
  • Your own “before” figures. Write down how long the task takes now and how often it goes wrong, before you start. Most published success stories have no “before” at all.
  • The checking time downstream. Count the corrections your accountant makes, the edits a reviewer makes, and the follow-up questions customers ask.

What shouldn’t a small business trust AI to do?

Anything where a confident mistake would be costly and you might not spot it. On the evidence, don’t leave these to AI unchecked:

  • Legal research, case citations and contract positions. The High Court has said that freely available AI tools cannot do reliable legal research.
  • Tax calculations, tax-rule answers, reconciliations and final accounts.
  • Spreadsheet analysis behind pricing, cash-flow or lending decisions.
  • Customer answers about prices, refunds, terms or eligibility. A business has been held responsible for what its website chatbot told a customer.
  • Diagnosis of what is wrong with your business, especially when it is struggling.
  • Software that handles logins, payments or personal data, built by someone who is not a developer.
  • Quantities and prices in quotes, and factual claims in tenders.

Where the marketing is ahead of the evidence

  • AI quoting and estimating, AI-written tenders, and AI phone receptionists. We found no independent evidence on results: nothing measuring quote accuracy, tender win rates or bookings gained. Only vendors’ own claims.
  • ”AI does your bookkeeping.” In 2025–26 tests, current AI models could not reliably complete accounting tasks on their own. AI suggestions checked by a person are a different and better-supported claim.
  • ”An AI adviser is as good as a consultant.” The only trial of AI business advice measured against real sales and profits found no improvement on average: stronger businesses gained, struggling ones did worse.
  • Time-saved headlines. Most are based on what users reported rather than on measured productivity. The UK Department for Business and Trade’s own evaluation found that reported time savings did not provide robust evidence of improved productivity (see Is AI worth it?).
  • ”Customers prefer instant AI answers.” The strongest positive result compares an AI chatbot with no answer at all. When AI replaces a person, or customers are told they are talking to a bot, the evidence runs the other way.

Which uses of AI have the strongest evidence?

The strongest evidence is for routine first drafts checked by a person, and for AI suggestions to newer staff in busy support teams. The table compares all 25 uses we reviewed.

Each use has one assessment. Five are steps on an evidence ladder; two describe evidence that cannot be placed on it.

The ladder, strongest first:

AssessmentWhat it means
EstablishedGood evidence from settings reasonably like a small business, pointing the same way in more than one independent source.
PromisingCredible evidence of benefit, but from less similar settings, or limited, or partly reported by the companies involved.
EmergingReal use and a credible capability, but results not yet known.
Weakly evidencedSounds plausible, but we found no convincing evidence of results in comparable businesses.
Poor fit / cautionEvidence of harm or worse work, a high risk of confident mistakes, or requirements most small businesses don’t meet.

Not on the ladder:

AssessmentWhat it means
MixedCredible evidence points in both directions and cannot yet be reconciled.
Insufficient to classifyToo little evidence of any kind to judge.
TaskUseAssessmentIn short
WritingRoutine first drafts, edited by a personEstablishedLarge time savings with equal or better quality in several studies, including one in a small firm.
WritingPersonal or relationship messages written mainly by AIPoor fit / caution (low confidence)Recipients judged the senders of heavily AI-assisted messages as less sincere; no study shows a benefit.
MarketingNewsletters and product copyPromisingAbout as effective as staff-written copy for much less effort, in one small firm and one large marketplace.
MarketingAd headlines and promotional push messagesMixedIndependent tests found no clear sales gain; the positive result is a vendor comparing its new AI with its old one.
Customer contactAI suggestions for less experienced staff in busy support teamsEstablishedThree large studies: faster handling and better ratings, mostly for newer staff.
Customer contactAI help for the experienced person who already answers wellPoor fit / cautionIn two large studies, the most skilled staff got slightly worse with AI help.
Customer contactA chatbot for enquiries that would otherwise go unanswered, with a route to a personPromisingMore sales than an “unavailable” message. Not compared with a person.
Customer contactA chatbot replacing people for complaints, complex or sensitive casesPoor fit / cautionThe business is liable for what the bot says; regulators expect a route to a person.
Customer contactAI phone receptionist for trades and small service firmsWeakly evidencedOnly vendors’ own claims.
SummarisingMeeting notes and internal summaries, checked by someone who knows the materialPromisingLarge time savings and good accuracy ratings in UK government trials, but errors do occur and an older trial found AI summaries worse.
SearchUsing AI chat or AI search for answers you will act onMixedFaster, but people miss its mistakes. Treat as Poor fit / caution if the answer is used without checking the source.
BookkeepingAI suggestions inside accounting software, checked by an accountantPromisingAssociated with accountants serving more clients and closing the month faster; not proven to be the cause.
BookkeepingGeneral AI tools doing reconciliations, month-end or tax on their ownPoor fit / cautionCurrent models fail independent accounting and tax tests far too often.
LegalLegal research with general chatbotsPoor fit / cautionInvented cases have reached UK courts; even specialist legal tools made things up 17–33% of the time.
LegalDrafting standard legal documents by trained professionalsPromisingLarge speed gains, modest quality gains, in studies with law students.
CodingDevelopers on routine or new codePromisingMore work completed, especially by less experienced developers; quality and maintenance costs rise.
CodingExperienced developers on large, familiar codebasesMixedMeasured slower in 2025, probably faster in 2026; the picture is moving.
CodingNon-developers building apps that handle customer data or paymentsPoor fit / cautionPeople wrote less secure code with AI and felt more confident about it.
Spreadsheets and analysisGeneral AI tools analysing data or building and checking spreadsheetsPoor fit / cautionSlower and less accurate in a UK trial; more wrong answers on judgement-heavy problems.
AdviceBrainstorming ideas and optionsPromisingOne person with AI matched a two-person team in a large experiment.
AdviceRelying on AI advice for hard decisionsPoor fit / cautionNo average benefit in a trial measured on real sales and profits; struggling firms did worse.
QuotingAI-generated quotes and quantity estimatesWeakly evidencedNothing independent measures quote accuracy, win rates or margins.
TendersAI-written bid responsesWeakly evidencedNo evidence on win rates; UK guidance warns of plausible but false statements.
TendersAI checking tenders or contracts for complianceWeakly evidencedOnly capability tests, one by a vendor.
Office toolsSlides and scheduling in AI office assistantsInsufficient to classifyVery little evidence; the small amount there is leans negative.

In our classification of these 25 uses, two meet our Established rule and eight are graded Poor fit / caution, one of them with low confidence. These counts depend on how the uses are divided; a different split would give different numbers. The assessments matter more than the totals.

What the combined evidence tells us across tasks

The findings below come from reading the studies together. No single study states them.

Does AI really make work faster, and better?

Across many well-bounded tasks, speed gains are common. Better work is conditional; correct work remains the critical question. A well-bounded task is one with a clear output that the person can check. Speed is not guaranteed: there are real slowdowns, listed below.

Speed gains turn up across many well-bounded tasks:

  • writing;
  • customer support;
  • summarising;
  • legal drafting;
  • search;
  • routine coding.

Better quality is less consistent. It rose for less experienced people and fell for some of the most skilled. In the legal studies, gains in speed were large and gains in quality small.

Correctness is where AI is weakest:

  • In one experiment, people using an AI search tool were about as accurate as others when the AI was right. When it gave a wrong answer, their accuracy fell from 93% to 47%.
  • Management consultants given AI on a problem chosen to be beyond its competence were 19% less likely to reach the correct answer.

Speed is not guaranteed:

  • Experienced software developers took 19% longer with AI in a 2025 study.
  • UK civil servants were slower and less accurate on a spreadsheet task with an AI assistant.
  • In a study of software projects, the jump in output faded within about two months.

“Versus what?”: how to read any claim about AI results

Every claimed gain is a comparison, and the comparison matters as much as the number. When you read any AI result, including ours, ask what AI was compared with:

ComparisonExample
AI versus a personAI-written marketing emails earned about the same profit as staff-written ones.
AI versus no serviceAn AI chatbot raised sales by 16.3%, compared with an automated notice telling shoppers customer service was unavailable. It was not compared with a human adviser. That makes it good evidence for answering enquiries nobody is answering, and no evidence for replacing your staff.
AI versus an older AI systemMeta’s improved ad-writing AI raised clicks by 6.7% compared with Meta’s previous AI, not compared with human copywriters.
AI versus your own previous performanceMost “hours saved” figures are users’ estimates against how long they think the task used to take, with no measured baseline.

Why AI time savings feel bigger than they are

People’s sense of how much time AI saves them can be wrong, and most published time savings are users’ estimates rather than measurements. Don’t ask whether AI feels faster. Measure it. Two examples show why:

  • Software developers (METR). In a 2025 experiment, experienced developers using AI took 19% longer than without it. Before starting, they expected AI to make them 24% faster. Afterwards, they still believed it had made them about 20% faster.
  • UK civil servants (DBT). In the Department for Business and Trade’s evaluation, staff reported in diaries that AI saved them time on data analysis. When people were observed doing an Excel analysis task, those using the AI assistant were slower and less accurate than those without it. The report itself notes the conflict. The observed group was very small, so this is a warning rather than proof, but it is exactly the gap a business should check for.

How to measure it in your own business: see How to tell whether AI is working.

Who benefits most from AI, and who can lose out?

Who gains depends on what the AI is doing. Less experienced people often gain most when AI supplies knowledge they lack. Expertise matters more when the AI’s output itself needs judging. Experts can get worse when AI replaces judgement they already exercise well.

This is SuperIntelligenceGuide’s interpretation of the studies taken together. No individual study draws this conclusion.

The studies seem to disagree about who benefits. Some find newer or weaker workers gain most; others find experienced people gain most. The disagreement largely disappears if you ask what job the AI is doing:

  • When AI supplies knowledge someone lacks, less experienced people often gain most.
    • Newer customer-service agents improved most when AI suggested replies learned from the company’s past conversations.
    • Less experienced developers and weaker writers also gained most.
  • When AI output itself needs judging, expertise becomes more important.
    • Experienced accountants used AI suggestions more selectively, stepping in when the software was unsure, and appeared to gain more.
    • In a business-advice trial, owners of stronger businesses benefited and owners of weaker ones did worse. The researchers trace the gap to how owners chose and applied the advice, not to the advice they received.
  • When AI substitutes for judgement an expert already exercises well, quality can fall.
    • In two large customer-service studies, the best-performing agents got slightly worse with AI help.

If you run the business yourself

If you are already the most experienced person doing the work, don’t assume that evidence showing AI helps less experienced staff applies to you. In two large customer-service studies, at a business-software company and at Alibaba, less experienced agents gained most from AI suggestions, while the best performers got slightly worse. In a small business, the person with the most experience is often the owner.

In our reading of the evidence, a better first use is preparation and repetitive work, such as first drafts, meeting summaries and templates for common questions, rather than handing over judgement you already exercise well. Then measure whether it helps.

What AI capability tests can and can’t tell you

Several sources are tests in which AI models attempt set tasks, such as accounting exercises, tax returns or spreadsheet problems, without people using them in a real business.

  • These tests can’t show that AI works in practice.
  • They can show it doesn’t yet. If the best current models fail most such tasks under controlled conditions, unsupervised use in a small business is unlikely to do better.

We use these tests to support caution, never to support a positive assessment.

The evidence task by task

Can AI write your business letters, reports and emails?

Yes, for first drafts that you check and edit: this is the best-supported use of AI we found. Messages that depend on a personal relationship are different, because people judged the senders of heavily AI-assisted messages as less sincere.

Routine first drafts, edited by a person: Established (good evidence from settings like a small business)

The strongest evidence on this page:

  • Speed and quality. In a controlled experiment with 453 professionals, ChatGPT cut the time for short writing tasks by 40% and raised quality, as graded by professionals in the same occupations, by 18%. People who were weaker writers gained most.
  • In a small firm. A small online wine retailer found AI-written marketing emails earned about the same profit as its staff-written ones. Each staff email took about three hours; checking an AI email took about ten minutes.
  • UK government. Civil servants using an AI office assistant reported the biggest time savings on drafting documents.

Two cautions:

  • Edit properly. A third of the participants in the professionals experiment submitted the AI’s first version without changes.
  • Check the facts. About a fifth of the civil servants said they had seen the AI make things up.
  • Try: standard letters, proposals, policies and reports drafted from your own notes.
  • Measure: editing time after the draft, not just drafting time; mistakes in facts and figures.
  • Don’t trust unsupervised: facts, figures, names and commitments.

Personal or relationship messages written mainly by AI: Poor fit / caution (low confidence)

  • Recipients’ reactions. In an experiment with 1,100 working professionals, a manager’s congratulatory email written with heavy AI help was judged less sincere and caring, and its author less able, even though it still read as professional.
  • No engagement gain. A separate field study found AI rewriting of employees’ emails made no difference to whether recipients opened or replied (an early, unreviewed paper).

This is evidence about perceptions, not lost customers, so our confidence is low. But small firms trade on personal relationships, and nothing shows a benefit.

  • Try: using AI to check tone or tighten something you wrote.
  • Don’t trust unsupervised: writing in your voice to people who know you.
Noy & Zhang (2023), Science: professionals’ writing tasks (S01)

453 college-educated professionals (marketers, consultants, HR staff, managers and others) were randomly given ChatGPT or not for short writing tasks typical of their jobs, in early 2023. With ChatGPT, tasks took 40% less time and quality rose 18%, graded by other professionals. Weaker writers gained most. A third submitted the AI’s first draft unedited. Limits: short, self-contained tasks needing no knowledge of a particular business; effects measured on the day only.

Dubé & Xu (2026), Quantitative Marketing and Economics: email marketing in a small online retailer (S02)

Wine Access, a small US online wine retailer, ran three experiments with about 27,500 customers each, comparing marketing emails written by staff, by AI, and by AI with staff editing. The authors “fail to reject that the LLM and Hybrid cells generate equal gross profits from orders as the Human cell”: in plain terms, they found no measurable difference in profit. Open rates were also the same. After counting staff time, the AI options raised net profit by about 2–9%. Limits: one firm; the tests could detect only large differences.

Department for Business and Trade (2025): Microsoft 365 Copilot evaluation (S03)

About 1,000 DBT staff used Copilot from October to December 2024, about 70% of them volunteers. In diaries, users reported saving 1.3 hours per document drafted and 0.2 hours per email. These figures are their own estimates, not measurements. 22% of respondents said they had seen it make things up. The report concludes: “We did not find robust evidence to suggest that time savings are leading to improved productivity.” Limits: self-reported, volunteers, public sector.

Cardon & Coman (2025), International Journal of Business Communication: how recipients judge AI-assisted emails (S04)

1,100 working professionals read versions of a manager’s congratulatory email described as written with low to high levels of AI help. Medium-to-high AI help lowered how sincere, caring and able the sender seemed, though messages were still seen as professional. Limits: participants reacted to a described scenario rather than real emails.

Ben-Zion & Lazebnik (2026), preprint: AI-rewritten workplace emails (S05)

121 employees in six companies had 16,880 emails rewritten by GPT-5 over three weeks, switching between conditions. AI rewriting made no direct difference to open rates, reply rates or response times. Limits: not yet peer reviewed.

Doshi & Hauser (2024), Science Advances: AI ideas and creative writing (S06)

293 writers wrote short stories with or without AI-generated ideas; 600 people rated them. AI ideas made individual stories more novel and useful (by about 8–9% with five ideas), especially for less creative writers, but made the stories more similar to each other. Limits: fiction, used here only as a sign that AI content tends toward sameness.

Marketing content

For newsletters and product copy, AI-written content performed about as well as staff-written content for much less effort, not better. For ad headlines and promotional push messages, independent tests found no clear sales gain.

Newsletters and product copy: Promising (credible evidence of benefit, but limited or from less similar settings)

  • Newsletters. Experiments at a small online wine retailer found AI-written newsletters performed about as well as staff-written ones, for a fraction of the effort.
  • Product descriptions. On a large international online marketplace, adding AI-written product descriptions alongside the sellers’ own raised sales by 2.1% and cut returns by 3.9%.

The sound conclusion is parity at lower cost, not better marketing. It rests on one firm and one platform, and AI-generated content tends to look alike.

  • Try: regular newsletters and product descriptions, with a person checking every claim.
  • Measure: orders or revenue per send against your previous staff-written sends; unsubscribes.
  • Don’t trust unsupervised: claims about products, prices, offers or regulated features.

Ad headlines and promotional push messages: Mixed (credible evidence points both ways)

The independent evidence and the vendor evidence point in different directions:

  • Independent tests (a large online marketplace). AI promotional push messages produced no clear sales gain, though slightly more people bought. AI-written titles for Google Shopping ads produced a small fall in sales that was too small to be sure of.
  • Vendor evidence. Meta’s own study found its improved ad-writing AI raised clicks by 6.7%, but compared with Meta’s previous AI, not with human copy.
  • Try: generating variants, if you already test one version of an ad against another.
  • Measure: run a split test against your own copy before switching.
Fang et al. (2025–26), field experiments on an online marketplace (S07)

Seven experiments on a large cross-border online retail platform (which the paper does not name), 2023–24, with the AI run centrally by the platform:

  • AI product descriptions added alongside sellers’ own: sales +2.1%, returns −3.9%.
  • AI push marketing messages: sales +1.6%, too small to be sure of, though conversion rose 3.0%.
  • AI Google Shopping ad titles: sales −4.5%, too small to be sure of.

Smaller and newer sellers sometimes showed larger gains, but the differences were mostly not statistically reliable. Limits: sellers did not choose or edit the AI copy; short-run results; one of the authors works with the platform.

Dubé & Xu (2026): see Writing (S08)

Staff-written, AI-written and AI-plus-editing newsletters earned about the same profit; AI cut writing effort from about three hours to about ten minutes per email.

Jiang et al. (2025), Meta: AI-generated Facebook ad text (S09)

About 35,000 advertisers and 640,000 ad variations over 10 weeks. Meta’s improved ad-text AI raised click-through by 6.7%, compared with Meta’s previous AI model. Limits: run and reported by the vendor; does not compare AI with human copy.

Doshi & Hauser (2024): see Writing (S10)

AI help made individual pieces more creative but made content more alike: a risk for marketing that needs to stand out.

Does AI customer service actually work?

AI customer service works for some jobs and not others. The strongest evidence is for AI suggesting replies to less experienced staff in busy support teams. Customer-facing chatbots have good evidence only for answering enquiries that would otherwise go unanswered, with a route to a person. The evidence turns against AI when it replaces people for complaints or sensitive cases, or when it is used to “help” the experienced person who already answers well.

Should AI replace your customer-service staff?

The evidence doesn’t support a general case for replacing good human customer service with AI. Three distinctions matter:

  • Helping newer staff has strong evidence. In large studies, AI suggestions made less experienced agents faster and better, while the best performers got slightly worse.
  • Answering enquiries that would otherwise get no answer has promising evidence. The best-known result, a 16.3% sales gain from an AI chatbot, was measured against an automated notice that customer service was unavailable. It shows AI beat no answer, not that it beat a person.
  • Replacing experienced human service is where the evidence turns against it. A business is liable for what its chatbot tells customers, as Air Canada found. Klarna, which reported in 2024 that its AI assistant handled two-thirds of customer chats, was reported in 2025 as saying a focus on cost had led to lower quality and that customers should always be able to reach a person.

Do customers mind dealing with AI?

The evidence is limited and partly dated, but it points one way: customers can react badly when they realise, or suspect, that they are dealing with AI.

  • In an older experiment, telling customers at the start of a sales call that they were talking to a bot cut purchases by 79.7%.
  • After a chatbot had failed them, customers reacted worse to replies so fast they seemed automated.
  • People judged the senders of heavily AI-assisted emails as less sincere.

AI suggestions for less experienced staff in busy support teams: Established (good evidence from settings like a small business)

Three large studies, in three different companies, found AI suggestions made support staff faster and customers slightly happier, with the biggest gains for newer or weaker staff:

  • A business-software company. Its support agents resolved 15% more issues per hour, and about 30% more for the least experienced.
  • Alibaba (Taobao), a randomised trial. Agents identified customers’ problems faster and got slightly better ratings.
  • A meal-delivery company. Agents responded faster and customer sentiment improved.

Two limits matter for small firms:

  • Volume and turnover. These were big operations with large volumes and staff turnover; most small businesses don’t have a support team like that.
  • Not solved more often. In the Alibaba trial, customers were no less likely to come back about the same problem within three days. The chats were better; the problems were not solved more often.
  • Try: suggested replies and knowledge-base look-ups for newer staff.
  • Measure: repeat contacts, not just speed or ratings.
  • Don’t trust unsupervised: policy, refund or contractual answers sent without staff reading them.

AI help for the experienced person who already answers well: Poor fit / caution (evidence of worse work or costly mistakes)

  • The most skilled got worse. In the business-software and Alibaba studies, the most skilled agents gained little speed and their quality fell slightly. At Alibaba, the top performers’ ratings fell and dissatisfaction rose.
  • In a small firm, that person is often the owner. The person answering customers is usually the most experienced one: exactly the group that lost out.
  • Try: templates for genuinely repetitive questions.
  • Don’t trust unsupervised: your judgement on unusual or upset customers.

A chatbot for enquiries that would otherwise go unanswered, with a route to a person: Promising (credible evidence of benefit, but limited or from less similar settings)

This is the model “versus what?” case. On a large online marketplace, a pre-sale AI chatbot raised sales by 16.3% compared with shoppers who got an automated notice that customer service was unavailable. That is good evidence for answering enquiries nobody is answering, which is a common small-business problem. It says nothing about whether a chatbot is better than you.

Two cautions:

  • Telling customers it’s a bot. In an older field experiment (before today’s AI chatbots), telling customers at the start of a sales call that they were talking to a bot cut purchases by 79.7%.
  • You are liable for what the bot says. A Canadian tribunal held an airline responsible for its chatbot’s wrong answer: “It should be obvious to Air Canada that it is responsible for all the information on its website.”
  • Try: answering common pre-sale questions when nobody is available, with an easy hand-off to a person.
  • Measure: sales from bot conversations, and complaints about wrong answers.
  • Don’t trust unsupervised: prices, refunds, terms, eligibility, or anything a customer could rely on.

A chatbot replacing people for complaints, complex or sensitive cases: Poor fit / caution

  • Liability. The business is liable for what its chatbot says.
  • Failed bot conversations. Customer feelings suffered when a failed chatbot conversation was followed by replies so fast they seemed automated.
  • Klarna. It reported in 2024 that its AI assistant was handling two-thirds of customer chats. In 2025 its chief executive was reported as saying that a focus on cost had led to lower quality, and that customers should always be able to reach a person.
  • UK regulators. The Financial Conduct Authority lists routing sensitive chatbot conversations, such as bereavement, to a person as good practice.
  • Try: triage only, then a person.
  • Don’t trust unsupervised: complaints, vulnerable customers, money owed, legal rights.

AI phone receptionist for trades and small service firms: Weakly evidenced (plausible, but no convincing evidence of results)

The only figures we found are vendors’ case studies, such as a claimed thousand-plus calls answered with none missed, plus widely repeated statistics about missed calls that we could not trace to any survey. The idea is plausible for firms that miss calls while on a job, but it is unmeasured.

  • Try: a short trial, keeping call logs.
  • Measure: calls answered, bookings made and callers who hang up, against the same weeks last year.
  • Don’t trust unsupervised: bookings and prices the system commits you to.
Brynjolfsson, Li & Raymond (2025), Quarterly Journal of Economics: support agents at a software company (S11)

5,172 customer-support agents, mostly in the Philippines, at a large business-software company were given an AI tool, in stages, that suggested replies learned from the company’s own past support conversations (2020–21). Issues resolved per hour rose 15% on average and about 30% for less skilled and less experienced agents. Customer sentiment improved. For the most experienced and skilled agents the authors found “small gains in speed and small declines in quality”. Limits: one large firm; a tool trained on its own records; an AI model from before ChatGPT. Figures are from the version matching the published journal article; an earlier 2023 version reported +14% and +34%.

Ni et al. (2026), preprint: Alibaba customer-service trial (S12)

5,940 Taobao after-sales agents were randomly given AI suggestions (or not) for four weeks in early 2024, across 2.56 million chats. Agents identified customers’ issues 8.2% faster on average; ratings rose slightly (by 0.042 on a five-point scale) and dissatisfaction fell 3.4%. Whether customers came back within three days, the measure of whether the problem was solved, did not change. Top-performing agents got worse on both ratings and dissatisfaction. Limits: one platform; four weeks; not yet peer reviewed.

Zhang & Narayandas (2025), Management Science: AI-assisted chat at a meal-delivery company (S13)

Agents given AI reply suggestions responded faster and customers’ sentiment improved, with newer agents gaining most. After customers had first struggled with a chatbot, very fast AI-assisted replies from human agents made them suspect they were still talking to a bot, which hurt sentiment. Limits: we could read only summaries, so we use the direction of the findings, not the effect sizes.

Fang et al. (2025–26): pre-sale chatbot experiment (S14)

44,614 shoppers were randomly assigned. Those given an AI pre-sale chatbot spent 16.3% more. The comparison group, in the authors’ words: “Consumers in the control group received the platform’s automated response service, which delivered a pre-programmed standardized notification indicating that customer service was unavailable.” Return rates and ratings did not get worse. Limits: AI versus no service, not versus a person.

Luo, Tong, Fang & Qu (2019), Marketing Science: telling customers it’s a bot (S15)

6,255 customers of a large Asian financial-services firm received loan-renewal sales calls from staff or from a voice bot. When not told it was a bot, the bot sold about as well as proficient staff (23.7% vs 25.1% purchased) and far better than inexperienced staff (4.9%). Telling customers at the start that it was a bot cut purchases by 79.7%, from 23.7% to 4.8%. Limits: a voice bot from before today’s AI chatbots; attitudes to disclosure may have changed.

Moffatt v Air Canada, 2024 BCCRT 149 (S16)

A Canadian tribunal held Air Canada liable for wrong bereavement-fare information from its website chatbot: “It should be obvious to Air Canada that it is responsible for all the information on its website.” It added: “I find Air Canada did not take reasonable care to ensure its chatbot was accurate.” Damages were C$650.88. Limits: a Canadian small-claims decision; persuasive, not binding, in the UK.

Financial Conduct Authority: consumer support good practice (S17)

The FCA’s published examples of good practice for regulated firms include automatically routing chatbot conversations about bereavement to a person, and removing obstacles that make it hard for customers to get help. Limits: financial services; guidance, not outcome evidence.

Klarna (2024–25), company statements (S18)

In February 2024 Klarna said its AI assistant handled two-thirds of customer-service chats with customer satisfaction equal to human agents. In May 2025 its chief executive was reported by Bloomberg (as quoted in trade press) as saying a focus on cost had led to lower quality, and that customers should always be able to reach a person. Limits: company self-report; the 2025 remarks are second-hand.

AI phone receptionist case studies (vendor, S19)

One UK vendor’s case-study page describes eight small businesses, including an estate agency, a repair firm and a sign company, with claims such as more than 1,000 calls answered with none missed, and monthly savings of £500–£1,500. No method or baseline is given. Treat these as marketing claims.

How reliable are AI summaries?

Reliable enough to save time on meeting notes and internal documents, if someone who knows the material checks them. In recent evaluations errors were uncommon but could be serious, and an older trial found AI summaries worse than people’s.

Meeting notes and internal summaries, checked by someone who knows the material: Promising (credible evidence of benefit, but limited or from less similar settings)

For:

  • DBT observed task. Civil servants using an AI assistant summarised reports in about 13 minutes against 42 without it, and their summaries were rated more accurate. Only a handful of people were observed.
  • DWP. Staff in the Department for Work and Pensions’ trial rated the accuracy of AI meeting notes good or very good in 85% of cases.
  • Clinical audit. An audit of AI-written clinical notes found made-up content in 1.47% of sentences and omissions in 3.45%. Low, but 44% of the made-up content was rated major.

Against:

  • ASIC (2024). The Australian Securities and Investments Commission found AI summaries of submissions to a parliamentary inquiry “performed lower on all criteria compared to the human summaries”. The AI model was older, and ASIC warned the results “should not be extrapolated more widely”.

Results depend on the model and on someone checking.

  • Try: meeting notes and internal digests reviewed by someone who was there.
  • Measure: spot-check summaries against the original, looking for what’s missing as well as what’s wrong.
  • Don’t trust unsupervised: summaries of contracts, legal or financial documents used instead of reading them.
DBT (2025): observed summarising task (S20)

In a small observed exercise, Copilot users summarised reports in 12 min 37 s against 41 min 34 s for non-users, with higher accuracy (4 vs 2.5 out of 5) and quality scores. Diary users estimated saving 0.8 hours per research summary and 0.7 hours per meeting summary. Limits: very few people observed; the report treats these results as supplementary.

Department for Work and Pensions (2026): Copilot trial evaluation (S21)

3,549 staff had licences from October 2024 to March 2025. Comparing users’ and non-users’ survey answers, users were estimated to save an average of 19 minutes a day across eight routine tasks (self-reported). 85% rated the accuracy of Copilot’s meeting notes good or very good. Limits: not randomly allocated; no measured baseline.

ASIC (2024): AI summarisation trial (S22)

Over five weeks, the Australian regulator tested an AI model on summarising public submissions to a parliamentary inquiry. In ASIC’s words: “The final results showed that the AI summaries performed lower on all criteria compared to the human summaries.” It noted the model’s limited ability to pick up nuance and context, and cautioned: “These point-in-time results relate to the use of certain prompts using a specific LLM, for a particular use-case, and therefore should not be extrapolated more widely.” Limits: an older model. Press reports give detailed scores we have not been able to check against ASIC’s documents, so we don’t use them.

Asgari et al. (2025), npj Digital Medicine: errors in AI clinical notes (S23)

50 doctors checked 12,999 sentences of AI-written consultation notes (GPT-4). 1.47% of sentences contained made-up content and 3.45% of relevant content was missing. 44% of the made-up sentences were rated major. Errors fell as prompts and workflows were refined. Limits: authors work for a company selling such tools; a clinical setting.

Vals Legal AI Report (2025): legal document summarisation test (S24)

In a commercial test of legal AI tools against practising lawyers, the best tool (CoCounsel) scored 77.2% on document summarisation against the lawyers’ 50.3%. Lawyers did better than the tools on other tasks, such as marking up contracts. Limits: vendors chose which tasks to enter.

Not without checking the source. In a controlled experiment, AI search was faster than ordinary search and as accurate when the AI was right; when it was wrong, many people didn’t notice. Independent checks have found frequent errors in AI answers.

Using AI chat or AI search for answers you will act on: Mixed (credible evidence points both ways)

AI search is faster, and fine when it’s right. In a controlled experiment:

  • people choosing between products took 1.6 minutes with an AI chat tool against 3.4 with ordinary search;
  • they were about as accurate when the AI was right (95% vs 92%).

People miss its mistakes:

  • When the AI was wrong, accuracy fell to 47%, against 93% with ordinary search.
  • Highlighting the AI’s less certain statements helped people spot errors.

Independent checks found frequent errors:

  • Columbia’s Tow Center. AI search tools got more than 60% of 1,600 source-identification questions wrong.
  • Public broadcasters across 18 countries. Almost half of AI assistants’ news answers had at least one significant issue.

UK government trials measured only time saved, as reported by users. None checked whether the answers were right.

If you act on an AI answer without checking the source, treat this as Poor fit / caution.

  • Try: a starting point for research, then open and read the sources.
  • Measure: how often a cited source actually says what the answer claims.
  • Don’t trust unsupervised: regulatory, tax, legal or safety answers.
Spatharioti et al. (2025), CHI: AI search vs ordinary search (S25)

US online participants (90 and 120 in two experiments) chose between products using an AI chat tool or traditional search. AI users finished in 1.6 minutes against 3.4 and were about as accurate when the AI was right (95% vs 92%). On a question where the AI was deliberately wrong, accuracy fell to 47% against 93%. Highlighting low-confidence text raised the rate of spotting errors from 26% to 53–58%. Limits: one simple consumer task; the authors work at Microsoft Research.

Tow Center, Columbia Journalism Review (2025): AI search engines and sources (S26)

1,600 queries to eight AI search tools asked them to identify a news article’s headline, publisher, date and link from an excerpt. Over 60% of answers were incorrect, often stated confidently. Limits: a narrow source-finding task.

EBU and BBC (2025): news accuracy of AI assistants (S27)

Journalists from 22 public broadcasters in 18 countries assessed more than 3,000 answers from major AI assistants. Almost half had at least one significant issue; about a third had serious sourcing problems; a fifth had major accuracy issues. Limits: news questions, not business ones.

UK government Copilot trials: searching (S28)

DBT users estimated saving 0.7 hours per search task; in the cross-government trial run by the Government Digital Service, over 70% reported spending less time searching. Neither measured whether the information found was correct.

Can AI do your bookkeeping?

With an accountant checking its suggestions inside accounting software, the evidence is promising, though it doesn’t prove AI causes the gains. On its own, no: in 2025–26 tests, current AI models could not complete accounting and tax tasks reliably.

AI suggestions inside accounting software, checked by an accountant: Promising (credible evidence of benefit, but limited or from less similar settings)

The most directly relevant evidence on this page for small businesses comes from a study of an AI-enabled accounting platform serving 79 small and medium-sized businesses:

  • accountants using the AI supported 55% more clients each week;
  • they closed the monthly books 7.5 days sooner.

Associated with, not proven to cause. The published paper describes these as associations. Accountants who chose to adopt AI may have differed in other ways. Its own experiment found that relying on AI suggestions where the evidence was uncertain increased the risk of errors.

A separate controlled trial with 100 auditors found better and faster work with an AI assistant, but no better compliance with the firm’s own audit guidelines.

  • Try: automatic categorisation and bank-matching suggestions, reviewed before anything is filed.
  • Measure: corrections your accountant makes at year-end; time to close the month.
  • Don’t trust unsupervised: final figures, VAT treatment and anything filed with HMRC.

General AI tools doing reconciliations, month-end or tax on their own: Poor fit / caution (evidence of worse work or costly mistakes)

Independent tests in 2025–26 found current AI models unreliable at accounting work. On 160 accounting tasks, including reconciliations and journal entries, no model got a task fully right in all of eight attempts more than 2.6% of the time (APEX-Accounting).

Tax questions

Don’t rely on a general AI chatbot for tax answers or tax calculations.

  • Tax returns (TaxCalcBench). On 51 US personal tax returns, the best model calculated about a third exactly right.
  • UK consumer questions (Which?). ChatGPT and Copilot accepted a deliberately false ISA allowance of £25,000 (the real limit is £20,000) and advised on investing it. Two tools pointed people to paid tax-refund firms alongside the free government service.
  • Try: explaining entries, or drafting questions for your accountant.
  • Don’t trust unsupervised: reconciliations, tax computations, tax-rule answers.
Choi & Xie (2026), Journal of Accounting Research: AI in accounting practice (S29)

A survey of 277 accountants plus data from an AI-enabled accounting platform serving 79 small and medium-sized businesses (over 200,000 transaction records). The working-paper version reports that AI-using accountants supported 55% more clients each week, moved about 8.5% of their time from data entry to higher-value work, and closed the monthly books 7.5 days sooner, with more detailed ledgers. The published version describes adoption as associated with these gains. It also reports a framed experiment: AI improved transaction classification on average, but relying on AI recommendations that went against the consensus could increase errors. Experienced accountants intervened more when the software was unsure. Limits: adopters chose themselves; one software provider; US.

Jezierski et al. (2025), working paper: auditors with an AI assistant (S30)

100 professional auditors at a global audit firm were randomly given access to the firm’s own AI assistant, or not, for a complex inventory valuation task. AI access improved output quality and cut completion time, most for active users, but did not improve adherence to the firm’s internal audit guidelines. Gains did not differ by experience. Limits: working paper; effect sizes not published in the summary we could read.

APEX-Accounting (2026): accounting task test (S31)

160 realistic accounting tasks (reconciliations, journal entries, variance analysis) in 10 simulated companies. The best model averaged 56.4% on expert marking criteria, and no model got any task fully right in all eight attempts more than 2.6% of the time. Limits: simulated companies; built by Mercor with Ramp, which have commercial interests.

TaxCalcBench (2025): AI and US tax returns (S32)

51 US federal personal tax returns. The best model (Gemini 2.5 Pro) calculated 32% exactly right. Common errors were using the wrong tax tables, calculation mistakes and wrong eligibility decisions. Limits: US tax; authors from a tax-software company.

Which? (2025): AI chatbots put to the test (S33)

In September 2025 Which? put 40 consumer questions to six AI tools. ChatGPT and Copilot missed a deliberate error about the ISA allowance (£25,000 rather than £20,000) and advised on investing it, which Which? said risked breaching HMRC rules. ChatGPT and Perplexity linked to paid tax-refund firms alongside the free government service.

Chartered Accountants Worldwide and Ipsos (2025): accountants’ use of AI (S34)

A survey of 2,718 chartered accountants from 13 institutes worldwide (211 in the UK), autumn 2024. 70% used public AI chatbots at least monthly; just under half said AI already helped them work more effectively; 30% named data security as the main reason they didn’t use AI more. Limits: opinions, not outcomes; mostly outside the UK.

Xero and Intuit: product claims (vendor, S35)

Xero (November 2025): “Our goal is for JAX to work behind the scenes to automatically reconcile more than 80% of your bank statement lines in real time.” This is a stated goal for a feature in testing. Intuit (July 2025): “saving businesses up to 12 hours a month”. Its footnote says 45% of customers reported saving 12 hours in a survey Intuit commissioned. Treat both as marketing claims.

Not for legal research with general chatbots: invented case citations have reached UK courts, and the High Court has said free AI tools can’t do reliable legal research. For trained professionals drafting standard documents, the evidence is promising.

Legal research with general chatbots: Poor fit / caution (evidence of worse work or costly mistakes)

  • The courts. In June 2025 the High Court warned: “Freely available generative artificial intelligence tools, trained on a large language model such as ChatGPT are not capable of conducting reliable legal research.” In one of the two cases before it, 18 of 45 cited authorities did not exist.
  • A tax tribunal. A litigant cited nine tribunal decisions that turned out to have been generated by AI.
  • Specialist tools. Even specialist legal research tools from LexisNexis and Thomson Reuters made things up 17–33% of the time in an independent test.
  • Try: getting oriented before speaking to a solicitor.
  • Don’t trust unsupervised: any legal position, citation or contract term you will rely on.

Drafting standard legal documents by trained professionals: Promising (credible evidence of benefit, but limited or from less similar settings)

Two controlled experiments with law students found large gains in speed:

  • GPT-4 (2024 study). 12–32% less time, depending on the task.
  • Newer tools (2026 study). Productivity gains of roughly 50–130% on five of six tasks.

Quality improved only modestly. A specialist tool that answers from legal sources produced no more made-up content than working without AI; a general AI model produced more.

  • Try: first drafts of standard documents within a professional’s own competence.
  • Measure: the time a reviewer spends correcting them.
  • Don’t trust unsupervised: citations and clauses not checked against authoritative sources.
Magesh et al. (2025), Journal of Empirical Legal Studies: specialist legal research tools (S36)

202 US legal research questions, pre-registered, tested in 2024. Lexis+ AI, Westlaw AI-Assisted Research and Ask Practical Law AI “each hallucinate between 17% and 33% of the time”: fewer than GPT-4, but far from the claims that such tools eliminate made-up answers. Lexis+ AI answered 65% accurately, Westlaw 41%, Practical Law 19%. Limits: US law; tools have been updated since.

Ayinde v Haringey; Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin) (S37)

The Divisional Court (6 June 2025): “Freely available generative artificial intelligence tools, trained on a large language model such as ChatGPT are not capable of conducting reliable legal research.” Lawyers have a duty to check AI research against authoritative sources. In Al-Haroun, 18 of 45 citations did not exist; the lawyers involved were referred to their regulators.

Harber v HMRC [2023] UKFTT 1007 (TC) (S38)

A litigant in person cited nine First-tier Tribunal decisions that the tribunal found were not genuine but had been generated by an AI system. The tribunal accepted the litigant had not known; the appeal was dismissed.

Choi, Monahan & Schwarcz (2024), Minnesota Law Review: GPT-4 and legal drafting (S39)

59 law students completed four drafting tasks with or without GPT-4 (late 2023). Time fell 12–32% depending on the task. Quality improved only slightly and inconsistently (clearly only for contract drafting). The lowest-skilled gained most. Limits: students; US law.

Schwarcz et al. (2026), Journal of Law & Empirical Analysis: newer legal AI tools (S40)

Law students (137 completed at least one task) used a legal-source-based tool (Vincent AI), a general reasoning model (o1-preview) or no AI on six tasks. Productivity gains were roughly 50–130% on five of six tasks. Quality improved modestly. Submissions containing made-up content: 3 with Vincent, 4 without AI, 11 with o1-preview. Limits: students; US law.

The Law Society: generative AI guidance for solicitors (S41)

Advises verifying every citation against reliable, authoritative sources and not putting confidential information into free AI tools; professional duties apply whether or not AI was used.

Coding and software

For businesses that employ developers, AI coding tools have promising evidence for routine work, with real costs in quality and maintenance. For non-developers building apps that handle customer data or payments, the evidence points to caution: security problems are common, and people feel more confident than they should.

Developers on routine or new code: Promising (credible evidence of benefit, but limited or from less similar settings)

Output rose:

  • Three large companies. Randomised roll-outs of GitHub Copilot found developers completed about 26% more tasks. The estimate is imprecise, and less experienced developers gained more.

But volume is not finished work:

  • Open-source projects. An independent study found that after adopting an AI coding editor, output jumped 3–5 times in the first month, then faded within about two months. Code warnings (+30%) and complexity (+41%) stayed higher.
  • Over 500,000 developers. With the newest AI agents, the amount of code committed rose roughly 180–240%, but finished software releases rose about 30%. The bottleneck moved to reviewing and testing.
  • Try: boilerplate, tests and well-defined features.
  • Measure: defects, rework and what actually ships, not lines of code.
  • Don’t trust unsupervised: security-sensitive code; anything merged without review.

Experienced developers on large, familiar codebases: Mixed (credible evidence points both ways)

2025: slower. METR’s randomised study found experienced developers took 19% longer with AI tools, while believing they were faster.

2026: probably faster, but uncertain. METR’s later data suggest a speed-up, but the estimate is not statistically certain. METR considers it an underestimate, because developers avoided doing work without AI. The tools are changing faster than the studies.

  • Measure: actual time on real tasks; self-estimates were wrong in the 2025 study.

Non-developers building apps that handle customer data or payments: Poor fit / caution (evidence of worse work or costly mistakes)

  • Less secure, more confident. In a controlled study, people with an AI assistant wrote significantly less secure code and were more likely to believe it was secure.
  • Public apps. A researcher’s scan of 1,645 apps built on one popular AI app-building platform found exposed data endpoints in 170 of them, about 10%.
  • No study shows non-developers can ship safe apps this way.
  • Try: internal prototypes that never touch customer or payment data.
  • Measure: get a security review before anything goes live.
  • Don’t trust unsupervised: logins, payments, personal data.
Cui et al. (2026), Management Science: GitHub Copilot in three companies (S42)

4,867 developers at Microsoft, Accenture and a Fortune 100 company, 2022–23. Access to GitHub Copilot raised completed tasks (pull requests) by 26% on average, with wide uncertainty. Less experienced developers adopted it more and gained more. At one company, the share of successful software builds fell. Limits: large firms; two authors work at Microsoft; counts of tasks are not the same as value.

METR (2025): experienced open-source developers (S43)

16 experienced developers worked on 246 real issues in large projects they knew well, with AI allowed or not, at random. “When developers are allowed to use AI tools, they take 19% longer to complete issues”. They had expected to be 24% faster and afterwards believed they had been 20% faster. Tools: Cursor with Claude 3.5/3.7 Sonnet. Limits: small sample; a very expert setting.

METR (2026): design update (S44)

With 57 developers from August 2025, METR estimated an 18% speed-up for developers returning from the first study and 4% for new ones; neither was statistically certain. METR considers these likely underestimates, because developers avoided tasks or the study rather than work without AI, and it is redesigning the study.

He et al. (2025): AI coding editor and open-source projects (S45)

806 open-source projects that adopted the Cursor editor, compared with 1,380 similar projects. Lines of code added rose 3–5 times in the first month, but the gains faded after about two months; code warnings rose about 30% and complexity about 41%, and stayed up. Limits: open-source projects, not businesses.

Demirer, Musolff & Yang (2026), NBER: code written vs software shipped (S46)

More than 500,000 GitHub developers across three generations of AI coding tools. With agent-style tools, code committed rose roughly 180–240%, but finished releases rose only about 30%; review and testing became the bottleneck. Limits: uses GitHub data (owned by Microsoft); the paper was revised in September 2026.

Perry et al. (2023): AI assistants and code security (S47)

47 participants did security-related programming tasks with or without an AI assistant. Those with AI wrote significantly less secure code and were more likely to believe it was secure. Limits: small study; an older model.

Anthropic (2026): AI help and learning to code (S48)

52 mostly junior engineers learned an unfamiliar programming library with or without AI help. The AI group was not significantly faster and scored 50% vs 67% on a later skills quiz, with the biggest gap on debugging. Limits: run by an AI company; small study.

Cloud Security Alliance (2026): apps built with AI app builders (S49)

Reports a researcher’s 2025 scan (CVE-2025-48757) of 1,645 apps built on Lovable: 303 exposed endpoints across 170 apps, caused by missing access controls. Limits: a scan, not a controlled study.

Can AI analyse your spreadsheets reliably?

Not reliably, on current evidence. UK civil servants were slower and less accurate on an Excel task with an AI assistant, management consultants were less likely to get a data-and-judgement problem right, and the best AI model in a 2026 test completed about a third of business spreadsheet tasks.

General AI tools analysing data or building and checking spreadsheets: Poor fit / caution (evidence of worse work or costly mistakes)

The evidence points one way, from UK office workers, elite consultants, researchers and tests:

  • UK civil servants. On an observed Excel task, those using an AI assistant were slower and less accurate, though their colleagues’ diaries claimed time savings.
  • Management consultants. On a problem combining spreadsheet data with interview notes, chosen to be beyond the AI’s competence, those with AI were “19% less likely to produce correct solutions”.
  • Researchers checking published studies. Teams with AI help were no more successful at reproducing published results than teams without, and found fewer major errors.
  • Spreadsheet test (June 2026). The best AI model completed 35% of business spreadsheet tasks and 12% of debugging tasks.

Model capability is improving quickly, so this is one of the assessments most likely to change.

  • Try: writing formulas you then test; explaining an existing spreadsheet.
  • Measure: check results against a known answer before trusting the method.
  • Don’t trust unsupervised: analysis behind pricing, cash-flow or lending decisions.
DBT (2025): observed Excel task (S50)

In a small observed exercise, Copilot users took 25 min 01 s against 20 min 33 s and scored lower on accuracy (1.5 vs 2.7 out of 5) and quality. The report: “M365 Copilot users completed Excel data analysis more slowly and to a worse quality and accuracy than non-users”. It notes this conflicts with the time savings reported in diaries. Limits: very few people observed.

Dell’Acqua et al. (2026), Organization Science: consultants and the “jagged frontier” (S51)

758 Boston Consulting Group consultants, GPT-4, 2023:

  • On tasks within AI’s competence, those with AI completed 12.2% more tasks, 25.1% faster, at higher quality.
  • On a business problem chosen to be beyond it, “subjects using AI were 19% less likely to produce correct solutions compared with those without AI”. Correct answers fell from 84.5% to 70.6% (AI only) and 60% (AI plus guidance).

Limits: an elite sample; a 2023 model.

Brodeur et al. (2025): researchers reproducing studies with and without AI (S52)

288 researchers in 103 teams tried to reproduce published quantitative studies. Teams with AI help matched human-only teams’ success rates; teams led by AI did much worse. Human-only teams found more major errors than AI-assisted teams (1.36 vs 0.63 per team). Limits: academic work, not business analysis.

SpreadsheetBench 2 (2026): business spreadsheet test (S53)

321 business spreadsheet tasks, including financial models, templates and debugging. The best model (Claude Opus 4.6) completed 34.89% overall and 12.00% of debugging tasks; a common failure was not inspecting the spreadsheet properly before changing it. Limits: no human comparison.

Can AI advice make business decisions worse?

AI advice can make business decisions worse. The only trial of AI business advice measured on real sales and profits found no average benefit, and owners of struggling businesses did worse. Brainstorming ideas and options is a different, better-supported use.

Brainstorming ideas and options: Promising (credible evidence of benefit, but limited or from less similar settings)

In a pre-registered experiment with 776 Procter & Gamble staff working on real product challenges, one person with AI produced solutions as good as a two-person team without it, in about 16% less time. Staff from outside the relevant function reached the level of teams that included a specialist. This is relevant to small firms without colleagues to think with. It is a single one-day exercise in a very large company, judged by experts rather than by sales.

  • Try: generating options, ideas and objections to your plan.
  • Measure: which ideas survive contact with customers.
  • Don’t trust unsupervised: choosing between the options for you.

Relying on AI advice for hard decisions: Poor fit / caution (evidence of worse work or costly mistakes)

The only field trial of AI business advice measured on real sales and profits:

  • Who. 640 entrepreneurs in Kenya.
  • Average effect. None.
  • Stronger businesses improved by about 15%.
  • Weaker businesses did about 8% worse.
  • Why, according to the researchers. The difference came from how owners chose and applied the advice, not from the advice they were given.

Supporting evidence:

  • Consultants. Given AI on a judgement-heavy business problem, they were less likely to get it right.
  • A laboratory study. Simply knowing advice came from AI made people rely on it too much, even against their own judgement.
  • Try: a second opinion after you have formed your own view.
  • Measure: decisions made with AI help and how they turned out.
  • Don’t trust unsupervised: diagnosis of what is wrong with the business.
Dell’Acqua et al. (2025), “The Cybernetic Teammate”: P&G product development (S55)

776 Procter & Gamble professionals in a one-day product-development exercise, randomly assigned to work alone or in pairs, with or without GPT-4. Individuals with AI performed as well as teams without it (+0.37 vs +0.24 standard deviations above individuals without AI) and took 16.4% less time. Staff whose usual job was outside product development reached the level of teams with a specialist. Limits: one day; one large firm; expert ratings, not market results.

Otis et al. (2026), Management Science: AI business mentor in Kenya (S54)

640 Kenyan small-business owners were randomly given a GPT-4 business mentor on WhatsApp, or standard business guides, in 2023. There was no average effect on an index of revenue and profit. High performers gained just over 15%; low performers did about 8% worse. The authors attribute the gap to how owners selected and implemented advice, not to the questions asked or advice received. Limits: Kenya; self-reported revenue and profit; two-month follow-up.

Dell’Acqua et al. (2026): see Spreadsheets (S56)

On a judgement-heavy business problem beyond AI’s competence, consultants with AI were 19% less likely to reach the correct answer.

Klingbeil, Grützner & Schreck (2024), Computers in Human Behavior: over-reliance on AI advice (S57)

In an incentivised laboratory experiment (319 participants), knowing that advice came from AI led people to over-rely on it, even when it contradicted other information and their own assessment, leading to worse outcomes. Limits: not a business setting.

Is there evidence for AI quoting and estimating?

No independent evidence on results. We found nothing measuring whether AI quoting improves accuracy, win rates or margins, only a small pilot, a capability test and vendors’ claims.

AI-generated quotes and quantity estimates: Weakly evidenced (plausible, but no convincing evidence of results)

We found no independent study measuring whether AI quoting improves accuracy, win rates or margins for trades, construction or manufacturing firms. What exists:

  • One pilot. A study covered five of 86 cost items on one bridge project. Its total was close, but individual items were out by up to 17.5%.
  • One drawings test. No AI model reached 80% on reading drawings and measuring quantities, in a test its authors call relatively straightforward.
  • Vendors’ claims. One UK quoting app, for example, says its quotes are “typically within 10–15% of a manually-produced quote”.

The surveyors’ professional body now regulates this. RICS’s professional standard, in force since March 2026, requires a named, qualified surveyor to make a written judgement on the reliability of material AI outputs, and requires AI use to be disclosed to clients.

  • Try: drafting quote wording or checklists from your own prices.
  • Measure: compare AI quantities with your own on past jobs before using it on live ones.
  • Don’t trust unsupervised: quantities, prices and the final figure.
Ghasemi & Dai (2024), Smart Construction: GPT-4 cost estimating pilot (S58)

GPT-4, given a custom cost database, priced five of 86 items on one US bridge repair project. The total came within about 0.2% of the reference estimate, but individual items were out by 1.6% to 17.5%, and answers varied with how questions were asked. Limits: a single case covering a small part of one project.

CEQuest (2025): AI and construction drawings (S59)

164 multiple-choice questions on reading construction drawings and measuring quantities. The best model (GPT-4.1) scored 75.4%; the authors note that all models scored below 80% on what they describe as a relatively straightforward dataset. Limits: a test, not real estimating work.

RICS (2025–26): Responsible use of AI in surveying practice (S60)

A mandatory professional standard for RICS members, published November 2025 and in force from 9 March 2026. It requires governance and risk management for AI, due diligence when buying AI systems, a written decision on the reliability of material AI outputs by a named qualified surveyor, and disclosure of AI use in terms of engagement.

AI quoting software: vendor claims (S61)

One UK trades quoting app claims quotes in 2–5 minutes rather than 1–3 hours, and: “Quotes are typically within 10–15% of a manually-produced quote, with accuracy improving when you provide photos and detailed job descriptions.” No method or source is given.

Is there evidence for AI tender and bid writing?

No evidence on results: no study measures whether AI-written bids win more or less often. UK procurement guidance allows suppliers to use AI but warns that AI-written content can be plausible but false.

AI-written bid responses: Weakly evidenced (plausible, but no convincing evidence of results)

  • No study measures results. Nothing shows whether AI-written bids win more or less often, or how evaluators score them.
  • UK procurement policy. It says “suppliers’ use of AI is not prohibited during the commercial process”, but warns that AI-written content may include “statements, facts or references” that “appear plausible, but are in fact false”.
  • Buyers’ lawyers. They predict more bids without more capacity to evaluate them.
  • Try: structuring responses and checking them against the question set.
  • Measure: scores and feedback compared with your previous bids.
  • Don’t trust unsupervised: facts, case studies and figures about your own firm.

AI checking tenders or contracts for compliance: Weakly evidenced

The evidence is capability tests outside UK procurement. The most often cited, on contract review, was written by a legal-technology company. Plausible, but unmeasured.

  • Try: a second check after your own read-through.
  • Don’t trust unsupervised: compliance sign-off.
Cabinet Office, Procurement Policy Note 017 (2025) (S62)

Applies to procurements starting from 24 February 2025 (replacing PPN 02/24). Says “suppliers’ use of AI is not prohibited during the commercial process”, and warns: “Content created with the support of LLMs may include inaccurate or misleading statements; where statements, facts or references appear plausible, but are in fact false.” Buyers should expect more clarification questions and responses.

Martin et al. (2024), “Better Call GPT”: AI contract review (S63)

Claims AI models match or exceed human accuracy on contract review at a small fraction of the cost. Limits: all authors work for a legal-technology company; contract review, not tender compliance.

Browne Jacobson (2025), Local Government Lawyer: AI and procurement (S64)

Lawyers advising public buyers predict a surge in tender responses without a matching increase in evaluation capacity, and responses that look similar. A prediction, not a finding.

Slides and scheduling in AI office assistants

There is too little evidence to judge, and the little there is leans negative.

Insufficient to classify (too little evidence to judge; tentatively negative)

  • Slides. In a small observed DBT exercise, civil servants made PowerPoint slides more than seven minutes faster with an AI assistant, but the slides were worse.
  • Scheduling. DBT users reported that scheduling took longer with it.
  • The wider cross-government trial reported overall time savings, as estimated by users.

These are heavily promoted features with almost no reliable evidence either way.

  • Try: if you use them, time them and check whether you redo the output.
  • Don’t trust unsupervised: client-facing slides.
DBT (2025): PowerPoint and scheduling (S65)

Observed task: Copilot users produced slides over seven minutes faster (10 min 47 s vs 18 min 30 s) but with lower accuracy (1.5 vs 5 out of 5) and quality. Diary users reported that scheduling took 0.6 hours longer per task and generating images 0.5 hours longer. Limits: very small observed sample; self-reported diary figures.

Government Digital Service (2025): cross-government Copilot experiment (S66)

20,000 licences across 12 public bodies, late 2024; 7,115 survey responses. Users estimated saving 26 minutes a day on average. There was no comparison group, and users also reported incorrect information.

How we assessed the evidence

What we included

We started from the tasks small businesses are most often told AI can help with. For each, we searched for published evidence of results, positive, null or negative:

  • academic studies;
  • government evaluations;
  • court and tribunal decisions;
  • regulators’ and professional bodies’ rules and guidance;
  • independent tests of AI tools;
  • vendors’ own claims, clearly labelled.

We re-opened every source we rely on, and checked every figure against the original document where we could reach it. Where we could read only a summary, we say so. Where a figure appears only in press reports, we either leave it out or label it. Direct quotations were each extracted twice, independently, and used only where both extractions matched word for word.

Types of evidence

Each source is labelled with one code:

CodeTypePlain meaning
CSIndependent causal studyPeople or businesses were assigned to use AI or not, usually at random (a randomised controlled trial), or a staged roll-out allowed a fair comparison. The strongest evidence that AI caused the difference.
IOIndependent observational studyCompares people or firms who chose to use AI with those who didn’t. Can show association, not cause: the adopters may differ in other ways.
PCProgramme or government evaluationFor example, the UK government’s Copilot trials. Often based on what users report.
TBTrade body, regulator, court or official guidanceRules, decisions and warnings; usually not evidence of results.
PSSurvey of firms or professionalsWhat people say they do or think.
SRCompany self-reportA company’s own account of its results.
VRVendor-reportedClaims or studies by companies selling AI. Never the sole basis for a positive assessment.
TCCapability test (benchmark)AI models attempting set tasks, without people using them in real work. Can support caution, not a positive assessment.
OSOfficial statisticsNot used for task-level assessments on this page.

Self-reported evidence, meaning people’s own estimates of time saved or benefit, never establishes an assessment on its own. In several studies it disagreed with what was measured.

How close the evidence is to a UK small business

DistanceMeaning
D1UK setting, or the same off-the-shelf tool on the same task
D2Same task, but a large organisation or another high-income country
D3A laboratory, online panel, students, or a lower-income country
D4A capability test only, with no people doing real work

We also record the AI model and the date. Many causal studies used models from 2020–24, and capability has moved since.

The assessment ladder

AssessmentRule
EstablishedAt least one causal study in a comparable task; the same direction of result in at least two independent sources; and a stated argument for why it applies to small businesses.
PromisingCausal evidence in a less comparable setting, or several credible, consistent implementations, including self-reported ones if labelled.
EmergingReal implementations and credible capability; results unknown.
Weakly evidencedPlausible, but no convincing evidence of results in comparable businesses.
Poor fit / cautionEvidence of harm or worse work, a high risk of confident error, or prerequisites most small businesses lack.

Assessment states outside the ladder

StateRule
MixedCredible evidence points in both directions and cannot yet be reconciled.
Insufficient to classifyToo little evidence of any kind.

These are not steps between other levels. “Mixed” does not mean “halfway between Promising and Caution”.

Rules we apply

  • Every assessment has a “because” naming the evidence it rests on.
  • Null and negative results are reported as prominently as positive ones.
  • Every gain states what it was compared with.
  • Self-reported time savings never establish an assessment alone; capability tests never support a positive one.
  • Lower-independence evidence (company and vendor claims) is kept and labelled, never relied on alone for a positive assessment.
  • “There isn’t enough evidence to recommend prioritising this” is an expected outcome.
  • Every assessment carries a last-checked date. We review the page at least annually, and sooner when a major study appears.

What we don’t do

  • We don’t combine studies into a single average effect.
  • We don’t estimate return on investment.
  • We don’t rank or recommend products.
  • Assessments are editorial judgements made against stated rules, not scientific scores.
  • The cross-task interpretations in “What the combined evidence tells us” are ours and are labelled as such.

Download the data

The full evidence behind this page is available as two spreadsheet files (CSV, which opens in Excel, Google Sheets or Numbers):

  • Assessments (ai-task-evidence-assessments.csv): the 25 uses, each with its assessment, the reasoning, the evidence types relied on, the argument for applying it to small businesses, limitations, and what to try, measure and not trust.
  • Sources (ai-task-evidence-sources.csv): the 66 source records (S01–S66) behind the assessments, with who was studied, which AI tool, what was found, design, how close the setting is to a UK small business, limitations, and how each source was checked.

Last checked: 4 October 2026. Version 1.0.

Reusing the data

Our licence. SuperIntelligenceGuide’s compilation, classifications, assessments and original analytical text are licensed under Creative Commons Attribution 4.0 (CC BY 4.0). Please credit “SuperIntelligenceGuide.co.uk, Where does AI actually help a small business? What the evidence shows”, with a link to this page.

What the licence does not cover. Verbatim quotations (the verbatim_quote column), source titles and other third-party material remain the copyright of their owners. They are included under the quotation exception in UK copyright law. If you reuse them, you are responsible for meeting its conditions yourself: fair dealing, no more than necessary, and acknowledgement.

Attribution statements:

  • Contains public sector information licensed under the Open Government Licence v3.0.
  • Contains information licensed under the Open Justice Licence v2.0.

Some sources restrict reproduction of their own material, for example consumer-test scores and ratings. We summarise those sources’ findings in our own words and do not reproduce their scores.

Data dictionary

Assessments file

ColumnMeaning
task_idTask group, T1–T13
taskThe type of work
useThe specific use assessed
assessment_kind“Tier” (on the evidence ladder) or “Assessment state (not a tier)”
tierEstablished, Promising, Emerging, Weakly evidenced, Poor fit / caution; or the states Mixed, Insufficient to classify
confidence_noteAny qualification on confidence or application
becauseOur justification, naming the evidence
source_codesEvidence types relied on (see the codes above)
transfer_argumentWhy the evidence does or doesn’t carry over to UK small businesses
evidence_limitationsMain weaknesses of the evidence
reasonable_to_tryWhat a business could reasonably try
measure_carefullyWhat to measure
do_not_trust_unsupervisedWhat not to leave to AI unchecked
last_checkedDate last verified

Sources file

ColumnMeaning
source_idStable identifier, S01–S66. One source can appear under more than one task.
task_idThe task the record informs
sourceCitation
urlThe address we checked
codeEvidence type (see the codes above)
population_settingWho and where
tool_testedWhich AI, and when
measured_outcomeMain results, in our words
quality_error_outcomeEffects on quality, errors or accuracy
experience_levelDifferences by experience or skill
designHow the study was done
setting_distanceD1–D4 (see above)
limitationsMain weaknesses
direction+ positive, 0 no clear effect, − negative, or a combination; “caution” for guidance without outcome data
verbatim_quoteShort quotations, each checked twice against the source. Not covered by our licence.
quote_checkResult of that check
verificationPrimary full text, primary abstract only, or primary in part (with flagged figures from secondary reports)
last_checkedDate last verified

Sources

59 sources, grouped by task in the order they appear above. The record IDs (S01–S66) match the downloadable source table; a source that informs more than one task has more than one ID.

Writing and marketing

Customer contact

Summarising and search

Bookkeeping and accounting

Legal

Coding

Spreadsheets, analysis and advice

  • S51, S56: Dell’Acqua, F. et al. (2026). Navigating the jagged technological frontier: field experimental evidence of the effects of artificial intelligence on knowledge worker productivity and quality. Organization Science 37(2): 403–423. https://doi.org/10.1287/orsc.2025.21838
  • S52: Brodeur, A. et al. (2025). Comparing human-only, AI-assisted, and AI-led teams on assessing research reproducibility in quantitative social science. IZA Discussion Paper 17645. https://docs.iza.org/dp17645.pdf
  • S53: Zhu et al. (2026). SpreadsheetBench 2. arXiv:2606.29955. https://arxiv.org/abs/2606.29955
  • S54: Otis, N., Clarke, R., Delecourt, S., Holtz, D. & Koning, R. (2026). The uneven impact of generative artificial intelligence on entrepreneurial performance: evidence from a field experiment in Kenya. Management Science (Articles in Advance). https://doi.org/10.1287/mnsc.2024.06909
  • S55: Dell’Acqua, F. et al. (2025). The cybernetic teammate: a field experiment on generative AI reshaping teamwork and expertise. NBER Working Paper 33641. https://www.nber.org/papers/w33641
  • S57: Klingbeil, A., Grützner, C. & Schreck, P. (2024). Trust and reliance on AI: an experimental study on the extent and costs of overreliance on AI. Computers in Human Behavior 160: 108352. https://doi.org/10.1016/j.chb.2024.108352

Quoting and tenders