How any company can use artificial intelligence (AI)
Artificial intelligence is not a single product. It is a set of tools that improve specific, everyday processes. Below are the processes almost every business runs, whatever the sector, and how AI can improve each one: honest numbers, real sources and clear notes on what to watch out for before you build anything.
Sales and lead handling
Sales and lead handling is everything that happens between a prospect showing interest and a salesperson talking to a qualified buyer. Enquiries arrive from every channel: web forms, email, phone, chat, social media, trade shows. Each one has to be recorded cleanly in the CRM, qualified (is this a real, funded, in-market buyer?), routed to the right person and followed up until they buy or clearly say no.
Every business that sells does this, and most do it patchily. Leads land faster than people can react, follow-up depends on who happens to be busy that week, and good prospects go cold simply because nobody came back to them in time. What turns interest into revenue is speed and consistency, not heroics.
The evidence here is unusually strong. A widely cited 2011 Harvard Business Review study audited 2,241 US companies and found that firms attempting contact within an hour of a web lead arriving were nearly seven times more likely to have a meaningful conversation with a decision-maker than those that waited just one hour longer, and more than 60 times more likely than those that waited a day or more. In the same audit, 23% of companies never responded at all.
This is precisely where AI earns its keep. It never sleeps, never forgets a follow-up, and reads and writes faster than any team. It does not replace a salesperson's judgement, relationships or ability to close; it removes the latency and the dropped threads that lose deals before a human is ever involved.
Instant 24/7 first response and conversational qualification
An AI assistant (a large language model such as Claude behind your website chat, an email autoresponder or a messaging channel) replies to a new enquiry within seconds, at any hour. Rather than pushing a rigid form, it holds a natural conversation: it answers the prospect's questions, asks a few qualifying ones (budget, timeline, company size, use case) and either books a slot in a rep's diary or hands over to a human the moment the lead is hot. Everything it learns is written back into the CRM, so the rep starts with full context.
A prospect fills in the contact form on a Leeds-based B2B software firm's website at 9pm. Within about 30 seconds an AI reply greets them by name, asks what problem they are trying to solve and roughly how many users are involved, checks whether they own the buying decision, and offers three diary slots for a demo. By morning the rep has a booked, pre-qualified meeting rather than a cold form to chase.
It closes the speed-to-lead gap that quietly drains pipeline. The 2011 HBR audit found that responding within the hour makes a lead nearly seven times more likely to be qualified, and the earlier MIT/InsideSales research (led by Dr James Oldroyd of MIT Sloan, across more than 15,000 leads at six companies) found that calling within 5 minutes rather than 30 made contact roughly 100 times more likely and qualification about 21 times more likely. No team reliably hits a 5-minute window around the clock; an AI first responder can.
Lead scoring and prioritisation
A model ranks incoming leads by how likely they are to convert, so your team spends its limited hours on the best opportunities first. Classic predictive scoring learns from historical CRM outcomes plus firmographic and behavioural signals. A language model adds a second layer: it reads the unstructured context (the email thread, the meeting notes, the pages the lead viewed) and explains in plain English why a lead looks strong or weak, which makes the score auditable rather than a black box.
An inbound queue of 200 leads a week is sorted automatically, so the 25 prospects who visited the pricing page twice and mentioned a Q4 deadline rise to the top, each with a short note such as "strong fit, active buying timeline, decision-maker title". Reps work the top of the list first instead of triaging 200 records by hand.
Scarce selling time is concentrated on the leads most likely to convert, fewer leads rot untouched, and reps get a reason for the ordering that they can trust and challenge. Vendors advertise figures such as roughly 90% scoring accuracy and 25% higher conversion, but those are vendor-reported and depend almost entirely on the depth and cleanliness of your own historical data, so read them as plausible upside rather than a promise.
Automatic CRM capture and enrichment
AI extracts structured fields from messy inputs (an inbound email, a call transcript, a business card photographed at a trade stand) and fills in the CRM record on its own: name, company, role, need, next step. It can also de-duplicate records, standardise formats and append missing firmographics, so records are clean enough to act on and to score against.
After a discovery call, the AI reads the transcript and updates the deal record: contact details captured, pain points logged, budget noted, next action set to "send proposal by Friday". The rep approves it in one click instead of typing notes for a quarter of an hour, or skipping them entirely, as so often happens.
Salespeople famously avoid CRM admin, so records go stale and pipeline reporting becomes fiction. McKinsey highlights automated CRM updates and meeting summaries as high-value, low-risk generative AI uses in sales, and cleaner data directly improves the scoring and routing that depend on it, which makes this foundational rather than cosmetic.
Persistent, personalised follow-up
The AI drafts a tailored sequence of follow-ups that reference what the specific prospect actually said, rather than a generic template, and sends them with human approval (or automatically for low-stakes touches). It manages the cadence, knows when to nudge and when to stop, and re-engages leads that have gone quiet.
A lead asked for pricing and then went silent for a week. Instead of a limp "just checking in", the AI sends something specific: "You mentioned needing this live before your March launch; here is a one-pager on timelines." It spaces three touches over two weeks and flags any reply for a human to take over.
Most deals are lost to absent follow-up, not to an outright no. Consistent, relevant follow-up recovers revenue that was already half-earned, without the rep having to hold every thread in their head. In effect it scales the persistence of a disciplined salesperson across the whole pipeline, including the long tail a busy rep would otherwise drop.
An honest reading of what can be built for sales today, and where the claims outrun the evidence:
- The workhorses can be built now: instant first response with conversational qualification, CRM data capture and enrichment, follow-up drafting with human approval, and meeting summarisation all rest on mature, well-understood capabilities. McKinsey singles out exactly this cluster (automated CRM updates, meeting summaries, drafted outreach) as offering measurable productivity gains with limited downside in B2B sales.McKinsey, An unconstrained future: how generative AI could reshape B2B sales
- The problem being solved is real and rigorously documented. The 2011 Harvard Business Review audit of 2,241 companies and the earlier MIT/InsideSales lead-response study together show that response speed measured in minutes, not hours, decides whether a lead is ever qualified, and that a large share of firms never respond at all.Harvard Business Review, The Short Life of Online Sales Leads (2011)MIT / InsideSales Lead Response Management Study (Dr James Oldroyd, MIT Sloan)
- Treat the headline numbers as direction, not guarantee. The vendor scoring-accuracy and conversion-uplift figures quoted above come from the people selling the tools, measured on other companies' data; your results depend almost entirely on the depth and cleanliness of your own CRM history. Baseline first, then measure the uplift on your own pipeline before believing any figure.
- Fully autonomous "AI SDR" agents that prospect, negotiate and close with nobody in the loop are still more pitch than practice. Conversational qualification and meeting booking work well in production; unsupervised outbound at scale still risks tone-deaf or non-compliant messaging and real brand damage. The sensible posture for 2026: AI does the reading, drafting, scoring and logging, and your people own judgement, relationship and the close.
For a UK business the watchdogs here are the ICO and the CMA, not Brussels, and the practical risks are as much about your data as about the law:
- Scoring and routing are only as good as your CRM history. Messy, duplicated or biased data teaches the model the wrong patterns, and if past conversions reflect where reps chose to spend their attention rather than genuine buyer fit, the model can learn to down-score whole regions or segments in a self-reinforcing loop that hides good leads. The ICO's guidance on AI and data protection expects you to address fairness and the sources of bias across the AI lifecycle, with a data protection impact assessment where the processing is likely to be high risk.ICO, Guidance on AI and data protection
- A model can quote a price or a policy that does not exist with complete confidence, so pricing, discounts and terms must come from a trusted system of record, never from the model's memory, and a human should approve anything that commits the business. The cautionary tale is Moffatt v Air Canada, where a Canadian tribunal held the airline liable for a refund policy its chatbot invented; that ruling is illustrative, but the domestic teeth are real: since 6 April 2025 the Digital Markets, Competition and Consumers Act 2024 lets the CMA fine misleading commercial practices directly, at up to 10% of worldwide turnover.American Bar Association, BC tribunal confirms companies remain liable for AI chatbot information (Moffatt v Air Canada)GOV.UK / CMA, Unfair commercial practices guidance (CMA207)
- Automatically scoring individuals engages UK data protection law. Since 5 February 2026 the Data (Use and Access) Act 2025 has replaced UK GDPR Article 22 with Articles 22A to 22D: solely automated decisions with significant effects on a person are now generally permitted for non-special-category data, but only with safeguards, which means telling the person, letting them make representations, and giving them meaningful human intervention and a route to contest the decision. Getting the processing principles or the automated decision-making rules wrong exposes you to ICO fines of up to £17.5 million or 4% of worldwide turnover.legislation.gov.uk, Data (Use and Access) Act 2025, section 80 (new UK GDPR Articles 22A to 22D)legislation.gov.uk, UK GDPR Article 83 (maximum fines)
- The wider rulebook in brief: there is no UK AI act, only existing regulators applying existing law, chiefly the ICO for anything touching personal data and consent for automated outreach; and the EU AI Act still reaches you extraterritorially the moment you place a system on the EU market or its outputs are used there.GOV.UK, AI regulation: a pro-innovation approach (white paper)EU AI Act, Article 2 (Scope)
Marketing and communication
Marketing and communication is how your company finds an audience, earns its attention and turns that attention into demand. In practice it covers writing content (blog articles, web pages, product descriptions), running campaigns (email, ads, social media) and staying visible where people search.
Every business does this in some form, whether it is a founder posting on LinkedIn between client calls or a 50-person team running campaigns across six channels. The work leans heavily on language, turns repetitive at the edges (the same message rewritten for different channels and segments) and is slow to produce in volume.
That is precisely the shape of work AI handles well today. It drafts, rewrites, summarises, classifies and personalises text and images quickly, so a small team can produce and test far more variations while people keep control of strategy, brand voice and the facts.
The realistic gain is a force multiplier on the routine 60-70% of the workload, not full automation. AI takes on first drafts, format conversions and variant generation; you keep the brief, the judgement and the final sign-off.
Content drafting and repurposing at scale
You give the model a brief, your brand-voice guide and a few past examples, and it drafts blog posts, web copy, product descriptions and ad variations. One source idea then becomes a newsletter, several social posts and a video script. Long-context models such as Claude can read an entire style guide and campaign brief in one go and hold that context across a whole working session, so the output stays close to your voice rather than drifting generic. The practical workflow is draft then edit: the model removes the blank-page problem and delivers a usable 70% first draft, and a person shapes the angle, checks the facts and polishes.
A B2B software team turns a single webinar transcript into a blog article, a LinkedIn carousel, three short posts and a follow-up email in an afternoon rather than spreading the work across a week. At the larger end of the ecosystem, beauty retailer Sephora worked with its agency Monks to use generative AI for producing and editing multiple cuts of one social film campaign across placements and phases, compressing the creative-iteration loop.
The gain sits in the speed and volume of first drafts and format conversion, not in replacing the writer. Self-reported surveys of marketers claim large multiples in content output and that most marketers now create content faster with AI (aggregated figures cite roughly 93% creating content faster), but read these as directional self-reports rather than measured productivity. The durable, repeatable benefit is fewer revision rounds and far less time spent staring at a blank page.
Email campaign optimisation
You feed the model your past campaigns together with their open and click data and ask it to surface which subject-line and structure patterns correlated with the strongest results, then draft fresh variants for A/B testing. It also writes segment-specific versions of the same email (one tone for new leads, another for existing customers) faster than anyone could by hand, while the marketer still decides which variants actually go out and validates the reading of the data.
A retailer loads its last 20 newsletters plus open and click metrics into the model, receives a ranked summary of the patterns behind the best performers, and generates five new subject lines and two body variants per segment for the next send. The model flags a pattern; a person confirms it is causal rather than seasonal before acting on it.
More tests and better-targeted sends generally support higher engagement, though the size of the lift depends entirely on list quality and the offer itself. The well-evidenced upstream benefit is personalisation: McKinsey attributes a 5-15% revenue lift and 10-30% better marketing-spend efficiency to personalisation done well at scale. The honest mechanism is that AI makes it cheap to produce and test more relevant variants; it does not guarantee any particular open rate.
Search visibility and generative engine optimisation (GEO)
AI helps cluster keywords, draft and structure articles around search intent, and add the structured data, clear citations and statistics that both traditional search and AI answer engines tend to favour. The newer discipline is GEO: raising the chance that your brand is the source ChatGPT, Google AI Overviews and Claude cite when they answer a question in your category, by publishing authoritative, well-sourced, quotable content. The AI assists with structuring and drafting; the underlying facts and authority still have to be real, because answer engines reward verifiable, cited claims.
A local services firm uses AI to map the questions its customers genuinely ask, drafts an FAQ-structured page with cited facts, then checks whether AI assistants now mention the brand when asked for recommendations in its category, and iterates on the gaps it finds.
Visibility is shifting from ranking on page one to being in the answer. Bain & Company found that about 80% of search users now rely on AI-generated summaries at least 40% of the time, and a large share of searches end without a click through to any website. Optimising to be cited therefore protects how you get discovered in future. The honest caveat is measurement: GEO attribution and tracking are still immature, so treat it as a hedge worth building rather than a solved, fully measurable channel.
Audience segmentation and personalisation
Instead of writing for a handful of broad audience buckets, AI lets your team generate many tailored message variants mapped to much finer-grained segments and behaviours, with copy, offer and imagery adjusted per group. The model does the variant-writing at volume; the marketer sets the segmentation strategy, the guardrails and which segments are worth the effort in the first place.
A European telecom moved from 4 macro-segments to roughly 150 personalised segments using a generative-AI messaging engine trained on non-personally-identifiable data and, as reported by McKinsey, saw a 40% lift in response rates alongside a 25% reduction in deployment costs.
Personalisation at scale is consistently tied to revenue lift (McKinsey: 5-15%) and lower acquisition cost (up to roughly 50% lower CAC in McKinsey's research), and faster-growing companies derive about 40% more of their revenue from personalisation than slower-growing peers. The mechanism is simple: the right message reaches the right person without a linear increase in human effort. The discipline is that more variants only help if the segmentation and the data behind them are sound.
Where the technology stands: what is solid, what is early, block by block.
- First-draft generation, repurposing one piece of content into many formats, summarising and analysing past campaign data, producing large numbers of personalised or A/B variants, and structured brainstorming all rest on mature capability and can be built dependably now. Self-reported adoption among marketers is high, with surveys placing usage or planned usage above 80%, though the exact figures vary by survey and should be read as directional.CoSchedule, State of AI in Marketing Report 2025Typeface, How Enterprise Marketing Teams Use Generative AI
- The time savings are credibly documented: roughly 6 to 13 hours per week per marketer, sourced from at least two independent surveys (ActiveCampaign/Talker Research, 1,000 marketers, 2025; HubSpot), concentrated exactly in drafting and planning work. McKinsey's personalisation figures (5-15% revenue lift, 10-30% spend efficiency, up to about 50% lower customer-acquisition cost) come from primary McKinsey research and are repeatedly corroborated. These are evidence the build is worth making, not guarantees of your outcome.ActiveCampaign / Talker Research, 13 Hours Back Each WeekMcKinsey, Unlocking the next frontier of personalized marketingMcKinsey, What is personalization?
- Be sceptical of the headline multiples. Figures such as 3.2x ROI or 10x content output originate in surveys run by the companies selling the tools, and set-and-forget autonomous campaign agents are not yet trustworthy in production. Take vendor numbers as a claimed ceiling rather than a floor you can count on, then benchmark against your own baseline before and after; that measurement, not the brochure, tells you what the build is worth.BizIQ, AI in Marketing Statistics 2026 (aggregated survey figures)
- GEO, optimising to be cited by AI answer engines, is real and growing: Bain found about 80% of search users rely on AI summaries at least 40% of the time. But the playbook is new and measurement and attribution remain immature, so build it as insurance for future discovery rather than a channel you can already report on precisely.Bain & Company, Consumer reliance on AI search resultsa16z, How Generative Engine Optimization (GEO) Rewrites the Rules of Search
Published marketing is where AI mistakes become public, so the guardrails matter as much as the speed.
- Hallucination is the central risk: a model can confidently state a wrong price, invent a product feature, misquote a source or fabricate a statistic, and once that ships in an ad or a landing page it damages brand trust and creates misleading-advertising exposure. Every factual claim, price, statistic and product detail in AI-drafted content needs human verification before publication, and AI output steered without a real brand voice tends towards generic sameness that quietly dilutes what makes you distinctive.Springer, AI Hallucinations in Marketing: Risks, Impacts, MitigationNeuralTrust, The Risk of AI Hallucinations: How to Protect Your Brand
- UK advertising rules apply in full to AI-generated ads. CAP guidance (29 May 2025) confirms the CAP and BCAP Codes contain no AI-specific rules, but the existing rules on misleadingness, offence and social responsibility apply regardless of how an ad was generated, edited or targeted. There is no blanket legal duty to disclose AI use in an ad; disclosure is needed where the audience would otherwise be misled, for example AI imagery showing product results that are not real, and disclosing AI use cannot cure a claim that is misleading in substance.ASA/CAP, Disclosure of AI in advertising (29 May 2025)
- Consumer-law enforcement has sharpened. Under the Digital Markets, Competition and Consumers Act 2024, in force for unfair commercial practices since 6 April 2025, the CMA can decide breaches itself and fine up to 10% of annual worldwide turnover. Fake reviews (including AI-generated ones) and drip pricing are now explicitly banned, so AI-assisted review programmes and AI-driven pricing displays need compliance review before launch. And although the UK has no general AI statute, a UK business whose AI outputs are used in the EU, or that markets AI-powered products into the EU, can still be caught by the EU AI Act's extraterritorial scope.GOV.UK/CMA, Unfair commercial practices guidance (CMA207)EU AI Act, Article 2 (Scope)
- Customer data in marketing AI sits under the UK GDPR, enforced by the ICO with fines up to £17.5 million or 4% of worldwide turnover. Feeding customer lists, behavioural data or CRM records into third-party AI tools requires a lawful basis, transparency about the processing and care over where the data is hosted; the ICO's Guidance on AI and data protection covers fairness, transparency and what an AI DPIA must address. Two practical limits complete the picture: if error rates are high, human review time can eat the hours AI saved, and GEO leaves you dependent on third-party answer engines whose citation behaviour you neither control nor reliably measure. Strategy, positioning, segmentation logic and final judgement should stay human.ICO, Guidance on AI and data protectionlegislation.gov.uk, UK GDPR Article 83 (fines)
Customer enquiries and support
Every business answers customer questions, whether it calls that a support desk or just the inbox. A customer emails, rings, opens a chat or raises a ticket, and someone has to work out what the issue is, find the right information, take whatever action is needed (a refund, a reset, a rebooking) and reply clearly.
The workload splits neatly in two. At the bottom it is repetitive: the same 20 to 30 questions, over and over, all week. At the top it demands real judgement: upset customers, edge cases, money on the line.
This is the gap AI actually closes. Modern language models can read a customer's message in plain English, pull the correct answer from your own help articles and order data, and then either draft a reply for a member of your team to approve or close simple cases entirely on their own.
The payoff is well documented: faster responses and a lower cost per routine contact, which frees your people for the difficult, genuinely human cases. The risk is equally well documented: a model can answer confidently and be wrong. The strongest deployments therefore keep a person in the loop wherever money, safety or an unhappy customer is involved.
A front-line AI agent that resolves routine enquiries on its own
An AI agent sits on your chat, email or help centre, reads the customer's question, retrieves the answer from your documented policies and the customer's account data, and replies directly. For simple, well-documented issues (where is my order, how do I reset my password, what is your returns policy) it can close the case end to end without anyone on your team touching it. You set the guardrails: which topics it may answer, when it must hand over to a person, and which actions it is allowed to take, for example looking up an order versus issuing a refund.
Klarna reported that its AI assistant handled 2.3 million conversations in its first month live, roughly two thirds of its support chat volume and the equivalent of about 700 full-time agents, across 23 markets and more than 35 languages. Intercom reported its Fin agent passing 40 million cumulative resolutions by late 2025, with a resolution rate of around 67% over a rolling 30-day window across its customer base. Both are cited not as products to buy off the shelf, but as ecosystem proof that the approach works at scale.
Instant answers around the clock, and a meaningful share of routine tickets closed with no staff time at all. The hard part is quality on anything non-routine: Klarna itself later walked the strategy back publicly, which is precisely why tight scoping and clear handover rules matter more than any headline deflection figure.
Agent assist: drafted replies your team approves and sends
Instead of replying to the customer, the AI sits beside the human agent and proposes a draft answer for every incoming message, drawing on your knowledge base and the account context. The agent edits it and sends it. Your team stays in control of every word that goes out, but the blank page work of composing each reply and hunting down the relevant policy disappears.
A field study published by the NBER (Brynjolfsson, Li and Raymond) followed 5,179 support agents at a large business software firm who were given a generative AI assistant suggesting a draft reply during each customer chat. Separately, French rail operator SNCF uses Claude to give roughly 150 support agents real-time draft responses and knowledge retrieval. The NBER study predates the current generation of models and was not Claude based; it is cited for the productivity effect, not the vendor.
In the NBER study, access to the assistant lifted issues resolved per hour by 14% on average, with a gain of about 34% for newer and less experienced agents and little change for the most experienced, alongside better customer sentiment and lower staff turnover. In practice, new hires reach the output of an experienced agent far sooner, because the tool encodes what your best people already do.
Automatic triage, priority and routing of tickets
Before anyone on your team opens a ticket, AI reads it and tags it: what it is about, what language it is written in, and how the customer is feeling. It then routes the ticket to the right team and sets its priority, so a furious complaint about a failed delivery jumps the queue while a routine billing query waits its turn, and a specialist question skips the generalist inbox altogether.
Zendesk now ships this kind of intent, language and sentiment triage as a standard platform feature, one sign the capability is mature. Accuracy and time saved depend heavily on your own ticket history and taxonomy, so treat any single benchmark with caution and measure on your own data.
Faster first responses, fewer tickets sitting in the wrong queue, and high emotion or high value cases surfaced to senior staff immediately rather than waiting their turn. Even a small per ticket saving on sorting compounds heavily at volume, and consistent priority rules cut the worst case wait for your angriest customers.
Knowledge retrieval and instant wrap-up for agents
AI indexes your scattered knowledge (help articles, past tickets, internal wikis, product documentation) so an agent can ask a question in plain English and get a synthesised, cited answer in seconds instead of trawling several systems by hand. The same engine can power self service search for customers. And at the end of a chat or call, the AI writes the summary, logs the outcome and drafts the internal note and any follow-up email, so nobody spends the last minute of every interaction typing notes.
Lyft built a customer care assistant on Claude (via Amazon Bedrock) that answers common rider and driver questions and routes harder cases to human specialists with an AI generated summary attached; it reported an 87% reduction in average resolution time, with over half of requests resolved in under three minutes. Wrap-up summarisation is among the lowest risk uses of all, because the human has already resolved the case and is only editing a recap.
Agents stop hunting across systems and give more consistent, correct answers; onboarding speeds up because knowledge becomes queryable rather than memorised. The wrap-up minutes recovered on every single interaction add up to a real fraction of an agent's day, and cleaner handover summaries mean customers rarely have to repeat their whole story to the next person.
All of this can be built for a UK business today, and the evidence is strongest exactly where a person stays in the loop.
- Agent assist carries the most solid evidence in this whole guide: a peer reviewed NBER field study of 5,179 support agents found a 14% average productivity gain, around 34% for newer agents, together with better customer sentiment and lower staff turnover.NBER, Generative AI at Work (Brynjolfsson, Li, Raymond): 5,179 agents, 14% average productivity gain, 34% for novices
- Triage, knowledge retrieval and after contact summarisation are mature, low risk builds, because a human still owns the decision or the case is already resolved; they usually pay back within weeks; the Lyft deployment cited above shows how large the gains can be when retrieval and summarisation are done well.Lyft blog, Lyft and Anthropic team up (Claude via Amazon Bedrock; 87% reduction in average resolution time)
- Fully autonomous resolution is real and operating at serious scale, as the Klarna and Intercom deployments cited above show. It is also the highest variance use, and the distinction that matters is between deflection (the AI handled the contact) and resolution (the customer's problem was actually solved).Klarna press release, AI assistant handles two thirds of customer service chats in its first monthIntercom, From resolutions to outcomes (Fin: 40M+ resolutions, around 67% rolling 30 day resolution rate)
- Read the headline figures as a compass, not a contract. They come from other markets and, often, from the companies selling the tools. Klarna's own chief executive later said the company had gone too far, that chasing cost had let quality slip, and began rehiring people for complex, premium service. Pilot against your own ticket history and judge on genuine resolution before you scale anything.TechCrunch, Klarna CEO says company will use humans to offer VIP customer service
The central risk is a confident wrong answer: a model can state a policy that has never existed. For a UK business, what catches that is not an AI statute but the law you already work under.
- The UK is not covered by the EU AI Act, and as of mid-2026 there is no general UK AI law. Government policy is the pro-innovation approach: existing regulators (the ICO, CMA, ASA and FCA) apply existing law to AI within their own remits. So there is no blanket legal duty to announce a chatbot, but honesty about AI is established good practice, and being unclear that a customer is talking to a machine can itself mislead: ASA and CAP guidance is explicit that disclosing AI use cannot cure a misleading claim. One caution for exporters: if you serve EU customers, or your AI's output is used in the EU, the EU AI Act can still reach you under its extraterritorial scope.GOV.UK, AI regulation: a pro-innovation approach (white paper)ASA/CAP, Disclosure of AI in advertising (29 May 2025)EU AI Act, Article 2 (Scope): extraterritorial reach over third country providers and deployers
- You answer for what your AI tells customers. In Moffatt v Air Canada, a Canadian tribunal held the airline liable for negligent misrepresentation after its chatbot invented a bereavement refund policy, rejecting outright the argument that the bot was a separate entity; the sum was small, the liability principle is what travels. In the UK the equivalent exposure runs through consumer law: under the Digital Markets, Competition and Consumers Act 2024 regime in force since 6 April 2025, the CMA can itself fine misleading commercial practices up to 10% of worldwide turnover, and a chatbot quoting prices or policies that do not exist sits squarely in that territory.American Bar Association, BC Tribunal confirms companies remain liable for AI chatbot information (Moffatt v Air Canada)GOV.UK/CMA, Unfair commercial practices guidance (CMA207): direct fines up to 10% of global turnover
- Personal and payment data flows through every one of these systems, so the UK GDPR and the Data Protection Act 2018 apply in full, enforced by the ICO with fines of up to £17.5 million or 4% of worldwide turnover. The ICO's Guidance on AI and data protection sets the expectations: fairness across the AI lifecycle, transparency with the people whose data you process, a lawful basis for the processing, and a data protection impact assessment where the risk is high. Review your supplier's security, data retention and hosting arrangements before launch, not after.legislation.gov.uk, UK GDPR Article 83 (fines up to 17.5 million pounds or 4% of worldwide turnover)ICO, Guidance on AI and data protection
- Four working rules from the deployments that go well. Keep a person in the loop on anything involving money, legal or safety content, account changes or an upset customer, and route those cases out automatically using sentiment and topic detection. Ground every answer in your own approved content and restrict which actions the AI may take, rather than letting it improvise. Test against real ticket history before launch, log everything for audit, and track whether the customer's problem was actually solved, not just how many contacts were deflected. And do not cut the team on day one: redeploy the people the AI frees up onto the complex, empathy heavy cases it handles poorly.
Appointments and scheduling
Every business runs on a diary somewhere. A dental practice fills its chairs, a letting agent books viewings, a heating firm routes engineers between jobs, and every office hunts for a slot that suits four busy calendars. The work looks trivial and rarely is: it is endless back-and-forth, time-zone arithmetic, last-minute changes and the quiet cost of appointments nobody turns up to.
The bottleneck is rarely the decision itself. Picking a slot takes seconds; the hours disappear into the communication around it, chasing confirmations, absorbing cancellations, having the same phone conversation forty times a day. That is exactly the kind of work modern language models handle well. An AI agent can read a vague request like "sometime Thursday afternoon, but not before two", check live availability through the calendar's API, propose valid options, write the booking and rearrange it when plans change.
No-shows deserve their own line in the accounts. For a clinic, a salon or an MOT garage, every empty slot is capacity you have already paid for. AI helps on two fronts: a risk model flags the bookings most likely to be missed, and a conversational reminder makes confirming or moving an appointment a one-tap job instead of a phone call the customer keeps putting off.
The honest condition: the value is real only when the agent is wired into the actual system of record, whether that is your practice management software or a shared team calendar, and when a person still signs off the cases that carry real stakes. Built that way, scheduling is one of the most dependable places to put AI to work first.
Self-service booking in chat and on the web
A customer or an employee types a plain request, for instance "book me a 30-minute consultation next week, mornings if possible". An agent built on a language model works out the intent, the time zone and the constraints, then calls your calendar or booking system as a tool: it reads real availability, proposes valid slots and writes the confirmed event with the right title, attendees and video link. It runs a loop of thinking, acting and checking the result until the booking lands. Because it understands natural language, it copes with fuzzy phrasing ("after lunch", "not Fridays") that rigid web forms reject, while staying pinned to slots the diary genuinely offers.
A patient messages a physiotherapy clinic's website at 9pm: "I need to see someone about my shoulder, ideally early next week." The agent checks the practice diary, offers Monday 8.40am or Tuesday 11.20am, books the chosen slot and sends a confirmation, with no staff involved. Mainstream scheduling platforms now bundle booking agents built on exactly this pattern, a fair signal that the approach has matured.
Bookings happen around the clock with no phone tag, and your team stops fielding routine scheduling calls. Requests that arrive after hours, which would otherwise die in voicemail or bounce off a closed reception, get captured instead.
A voice agent on the phone line
An AI voice agent answers inbound calls, transcribes and understands speech, checks the live diary and books, moves or cancels appointments while updating the scheduling system in real time. It can take several calls in parallel, quote opening hours, read the confirmed details back to the caller and pass anything complex or clinical to a human. This matters because a large share of appointments, especially in healthcare and local services, still arrive by phone.
A dental practice in Manchester puts a voice agent on its after-hours and overflow line. Routine hygiene appointments go straight into the practice management system; clinical questions are queued for the reception team the next morning. Commercial platforms for building this kind of voice receptionist are now an established category, so the capability is no longer experimental.
Fewer missed calls, shorter hold times and a front desk that is interrupted far less often. Demand that used to land in voicemail gets captured. The one condition: the agent needs a clean, reliable escalation path for anything it cannot handle on its own.
No-show prediction and intelligent reminders
A model scores each upcoming appointment for no-show risk using history: past attendance, how far ahead the booking was made, appointment type, day and time, distance from the premises. High-risk bookings get targeted reminders with one-tap rescheduling, and a slot that is likely to open up can be offered to a waiting list. A language layer runs the reminder conversation (confirm, cancel, move) in plain English over SMS or chat. Worth knowing: the messaging and the risk model are often separate components, and many real deployments earn most of their result from two-way reminders alone.
A peer-reviewed before-and-after study in primary care in the United Arab Emirates (JMIR Formative Research, 2025, covering 135,393 appointments) reported a 50.7% reduction in no-shows after a prediction model of roughly 86% accuracy fed a real-time dashboard that coordinators used to manage high-risk bookings. Separately, a vendor case study of El Rio Health, a community health centre in Arizona, reported a 32% drop in no-shows and roughly $100,000 a month in recovered revenue from automated voice and SMS reminders with two-way rescheduling; that one is self-reported by the vendor, not an independent trial.
Missed appointments are estimated to cost the US healthcare system in the order of $150 billion a year, an industry figure that is widely cited but not peer-reviewed. Reminder and prediction programmes are commonly reported to cut no-shows by roughly 15 to 30%, sometimes more, and every recovered slot is capacity you had already paid for. Read the headline single-site numbers as illustrative ceilings, not a promise.
Dispatch and rota optimisation with a plain-English front end
For field service teams, clinics with rooms and equipment, or shift-based rotas, optimisation software assigns the right person or asset to each job and re-plans when reality moves: a job overruns, an engineer calls in sick, traffic wrecks a route. The assignment weighs location, skills, parts availability, travel time and workload balance. The genuinely new piece is a language model on top, so a dispatcher can ask "who can cover the 3pm emergency in Salford?" and get an explained recommendation instead of raw solver output. The hard optimisation underneath is classical operations research, not a language model.
A heating firm's system re-plans the day when an engineer goes off sick: it finds the nearest qualified engineer with capacity, proposes which lower-priority boiler service to move, and shows the dispatcher its reasoning for approval rather than applying the change silently.
Better utilisation, less time on the road, faster response to emergencies and fewer manual reshuffles. One field-service vendor describes weighing more than 50 variables per assignment, faster than a person can in the moment; treat that as the vendor's own number, and remember the gains rest entirely on clean, current data about jobs, skills and locations.
Judged as build-it-now engineering rather than futurology, most of this block stands on solid ground: one-to-one self-service booking, automated reminders with one-tap rescheduling and no-show risk scoring all run on well-defined calendar APIs and narrow conversations, and they can be built against your existing systems today.
- The proven core is backed by more than vendor decks: the UAE primary-care study cited above is peer reviewed and ran across a six-figure appointment volume, evidence the pattern works at scale, not just in a demo.Real-time analytics and AI for managing no-show appointments in primary care (JMIR Formative Research, 2025)Same UAE study, full text mirror (NCBI/PMC)
- Voice booking for routine appointments has crossed into everyday commercial use in clinics, dental practices and service businesses, and multi-party meeting coordination is documented in both research prototypes and production agents. The fully autonomous version, negotiating with external strangers end to end, is still the aspirational tier: most working setups keep a human copied in or confirming.AI voice agents for appointment scheduling in clinics (Retell AI, vendor landscape)ScheduleMe: multi-agent calendar assistant architecture (arXiv)
- One health warning on the numbers: the most quotable figures in this space, the El Rio Health no-show and recovered-revenue results cited above and the industry estimate of what missed appointments cost the US system, are respectively a single vendor-reported case study and an unaudited figure. They tell you the direction of travel, not what your diary will deliver; pilot against your own no-show rate before you bank either.El Rio Health automated reminders case study, 32% no-show reduction (Emerging Global, vendor-reported)Missed appointments cost the US healthcare system $150B a year (HCI Innovation Group, industry estimate)
- Dispatch and rota optimisation is mature technology that predates the current AI wave: classical solvers have carried the heavy mathematics for years. What is genuinely new, and worth building now, is the conversational layer that lets a dispatcher query and understand the plan in plain English. Keep that layer advisory and leave the final call with the solver plus a human.The guide to AI in field service management (IBM)AI-powered scheduling for field service (FieldCamp, vendor source)
An agent that misreads a date or drops a time zone will do it with complete confidence, so the protection is engineering discipline, not optimism. These are the guardrails, and the UK rules, to design around.
- Constrain the agent to reality: it should act through the real calendar API (read first, then write), echo the exact slot back for confirmation and never generate availability freely. Independent evaluations show language models hallucinate at meaningful rates, and time zones, the British Summer Time switch and recurring events remain classic failure points. A brilliant agent on a stale calendar will still double-book, and a person should approve the high-stakes cases: clinical triage, external or VIP meetings, conflicting bookings.Guide to hallucinations in large language models (Lakera)
- Booking data is personal data under the UK GDPR and the Data Protection Act 2018, enforced by the ICO, and an appointment often reveals health information: a physiotherapy or dental booking says something about a person's health, which brings the stricter special category rules into play. Get the lawful basis right, run a data protection impact assessment where AI processes personal data at scale, and follow the ICO's Guidance on AI and data protection; the penalties behind it are the ICO maximums quoted throughout this guide.ICO, Guidance on AI and data protectionUK GDPR Article 83, penalties (legislation.gov.uk)
- If a no-show risk score starts deciding by itself who gets a slot, or an optimiser reshuffles staff rotas with no human involved, you are in automated decision-making territory. Since 5 February 2026, the UK GDPR's Articles 22A to 22D (inserted by the Data (Use and Access) Act 2025) permit solely automated significant decisions only with safeguards: informing the person, letting them make representations, contest the decision and obtain human intervention, with a stricter regime where health data is involved. For rotas there is a human reason too: silent algorithmic reshuffling erodes staff trust even when it is mathematically optimal, so propose changes for approval rather than imposing them.Data (Use and Access) Act 2025, section 80 (legislation.gov.uk)ICO, The Data (Use and Access) Act 2025: what it means for organisations
- The UK has no AI act: the pro-innovation framework leaves enforcement to existing regulators in their own lanes, so there is no blanket statutory duty to announce that a caller is speaking to an AI. Concealing it is still poor practice, because the ICO expects transparency around AI and consumer rules catch misleading behaviour. And the exemption stops at the border: if you take bookings from customers in the EU, or your agent's output is used there, the EU AI Act applies to you extraterritorially, transparency obligations included.GOV.UK, AI regulation: a pro-innovation approach (white paper)EU AI Act, Article 2 (scope and extraterritorial reach)
Invoicing, payments and bookkeeping
Every business, whatever it sells, raises invoices, receives supplier bills, collects money, pays its own suppliers and records all of it in the books. The cycle runs from invoicing and credit control through purchase ledger and bank reconciliation to bookkeeping proper, where every transaction is coded to the right account for VAT and reporting.
The work is high volume, repetitive, rule-bound and unforgiving of mistakes, which is exactly why AI changes it. Modern AI reads messy documents, drafts chasing emails, codes transactions and explains discrepancies in plain English, shifting hours of manual keying and chasing towards a review-and-approve model.
UK businesses start with an advantage here. Making Tax Digital already requires VAT-registered firms to keep digital records and file through software, so much of the data foundation an AI build needs is in place before you begin.
The realistic pattern is not a finance team replaced. It is a finance team that handles exceptions and approvals while routine extraction, matching and coding run in the background, with a human gate before any money moves or anything is filed with HMRC.
Invoice and receipt capture (intelligent document processing)
Instead of a person reading a supplier PDF, a scanned receipt or an emailed bill and typing the supplier, date, amounts, VAT and line items into the accounting system, an AI model reads the document directly. Large language models combined with intelligent document processing interpret what a document means rather than only its fixed layout, so they handle invoices from thousands of different suppliers without anyone building a template per vendor, which was the main limitation of older template OCR. Extracted fields flow into the purchase ledger, and anything the model is unsure about is flagged for a person to approve.
A facilities management firm in Leeds receives 400 supplier invoices a month in dozens of formats. The AI pulls the totals, VAT, purchase order number and line items from each one, fills the purchase ledger itself and routes only the minority it cannot read confidently, such as smudged scans or unusual layouts, to a clerk to confirm. Mainstream accounting platforms now pre-fill bills and receipts this way, a sign the capability is mature rather than experimental; a system built on your own accounts and rules does the same without changing your ledger software.
Industry benchmarks report AI invoice capture reaching roughly 95 to 99% field-level accuracy on common header fields, against 85 to 95% for older template OCR, and end-to-end processing time falling from over two weeks to a few days for the best-run teams. These are vendor and analyst figures and they vary widely with the mix of documents you receive.
Three-way matching and transaction coding
The AI compares each incoming invoice against its purchase order and the goods-received record, confirming quantities and prices agree before payment is allowed. For the books, it learns from how you have coded transactions in the past and proposes the right expense or income account for each new bank or card transaction, so routine coding largely runs itself. Exceptions go to a person rather than the whole population, and every correction a person makes teaches the model.
A builders' merchant's card feed produces hundreds of transactions a week. The AI codes each one from prior patterns: fuel to travel, a software subscription to IT costs, a supplier payment matched against its open invoice. The big accounting packages work on the same learn-from-history principle, proposing a treatment for each new transaction for you to confirm or adjust.
Vendors and industry write-ups report automated three-way match rates of 85 to 95% on clean, everyday cases, alongside material time savings; platform vendors report most of their users saving time with these features, a vendor-reported figure. The durable win is quieter: fewer duplicate and overpayments, because mismatches are caught before money leaves.
Bank reconciliation and discrepancy explanation
Reconciliation means matching the money in the bank to the invoices and payments it represents. AI takes on the messy middle: partial payments, one payment covering several invoices, remittance details buried in an email or PDF, and timing gaps. A language model reads a remittance advice or a customer's email, works out which invoices a payment covers, and rather than just flagging a discrepancy, states in plain English why two figures differ, for example a deducted credit note or a short payment for a damaged delivery, leaving a person to approve the allocation.
A customer settles five invoices with a single BACS payment, less a credit note for a damaged delivery. The AI parses the remittance email, allocates the payment across the right invoices, identifies the reason for the short payment and drafts a note the bookkeeper can approve in seconds. Platform vendors such as BILL now ship the same capability as a built-in reconciliation agent, a sign the approach has gone mainstream.
It clears the exception backlog that normally swallows month-end and adds finance capacity without adding headcount. BILL reported that in early rollout the share of transactions processed entirely by its AI rose 533%, at about 92% accuracy; that is a vendor-reported early figure measured on its own platform, not an independent benchmark.
Credit control: intelligent chasing and collections drafting
The AI watches which invoices are overdue, flags who is likely to pay late based on their history, and drafts professional, personalised reminders timed to each customer's behaviour. It reads the replies that come back, whether a dispute, a promise to pay or a query about an invoice number, infers the intent and drafts an appropriate response, escalating the tone gradually from a gentle nudge to a formal notice. A person approves before anything goes out, especially to key accounts.
A B2B services firm has 150 open invoices. The AI groups customers by payment history, chases habitual late payers earlier and sends a courteous note to reliable ones. When a client emails 'we never received invoice 1042', it retrieves the invoice, drafts a reply with it attached and updates the record, ready for a person to send.
Faster collection and lower debtor days without anyone manually working through every line of the aged debtors report. Some vendors claim reductions of up to about 80% in collections call-centre costs for voice AI deployments, but those are marketing figures for specific set-ups and should not be read as a general benchmark.
A candid picture of what can be built solidly today, what is still ambitious and where to keep expectations grounded:
- AI document capture for invoices and receipts, learned transaction coding, three-way matching and drafted payment reminders all rest on mature, production-proven technology, so a build in this area starts from solid ground rather than research territory.Parseur, AI Invoice Processing Benchmarks 2026Artsyl, Automated Invoice Processing 2025-2026: AI with Human Oversight
- Reconciliation of clean, everyday cases can be built dependably today, and AI that explains a discrepancy in plain English, such as a short payment caused by a deducted credit note, is real and steadily improving.Ledge, AI reconciliation: real-world use casesBILL, New AI agents for touchless transactions (Reconciliation Agent)
- Fully autonomous bookkeeping with no human involvement sits at the early, ambitious end. 'AI accountant' products appeared in 2026 marketed as running the whole bookkeeping process with little or no human input; treat those claims with scepticism until they are validated independently on your own ledger.Accounting Today, Pilot launches fully autonomous AI bookkeeper (Feb 2026)
- One caution that covers every figure above: accuracy and cost benchmarks in this space are published mostly by the people selling the software, measured on their data rather than yours. Read them as a direction of travel, then pilot on a sample of your own invoices and let your own ledger decide.
Finance is unforgiving territory for automation: a wrong figure is not a typo, it is a misstatement, a duplicate payment or a VAT error.
- General-purpose language models can hallucinate or misread on financial tasks, so AI output should never trigger an irreversible or material action, such as paying a supplier, filing a VAT return or releasing a refund, without a human approval gate, confidence thresholds and a full audit trail. Segregation of duties still applies: the AI proposes, a person approves and a separate control reconciles.Baytech Consulting, Hidden dangers of AI hallucinations in financial services
- There is no general UK AI statute to comply with; the UK regulates AI through existing regulators, and for finance data that means the ICO and the UK GDPR. Supplier contacts, customer records and payment histories are personal data, the ICO publishes dedicated guidance on AI and data protection, and fines reach £17.5 million or 4% of worldwide annual turnover, so ground any AI build in a lawful basis and, where the risk warrants it, a data protection impact assessment.ICO, Guidance on AI and data protectionlegislation.gov.uk, UK GDPR Article 83 (penalties)
- Making Tax Digital already binds you: every VAT-registered business must keep digital records and file VAT returns through compatible software with unbroken digital links, and from 6 April 2026 sole traders and landlords with qualifying income over £50,000 join MTD for Income Tax. Any AI layer over your books must preserve that digital chain, never replace it with copy-and-paste.GOV.UK, VAT record keeping (Making Tax Digital for VAT)GOV.UK, Making Tax Digital for Income Tax for sole traders and landlords
- E-invoicing is coming, not current: the UK has no general B2B e-invoicing mandate today, but at Budget 2025 the government announced that all VAT invoices must be issued as e-invoices from April 2029, covering B2B and B2G with B2C excluded, and the implementation roadmap is due at Budget 2026. If you are building invoicing automation now, design it around structured invoice data so the 2029 switch is a configuration change, not a rebuild.GOV.UK, E-invoicing consultation response (26 November 2025)
Documents, contracts and data extraction
Every business, whatever it sells, runs on paperwork that arrives in formats meant for human eyes: supplier invoices, signed contracts, purchase orders, application forms, scanned PDFs and emails with attachments. Someone has to read each one, find the fields that matter, the amounts, the dates, the parties, the clauses, and type them into an accounting package, a CRM or a spreadsheet. It is slow, it costs money and it invites mistakes.
Modern AI reads a document much the way a person does. Large language models (LLMs) such as Claude, including the vision capable ones, interpret a document's structure and meaning and return clean, structured data: the invoice total, the renewal date, the payment terms, ready to drop into your systems. The practical difference from legacy optical character recognition (OCR) is comprehension. Classic OCR turns pixels into raw text and usually needs a template configured for each layout, whereas a language model can read a document it has never seen and still find the right fields, because it reasons from context rather than fixed coordinates.
There is a trade off, and it is worth being straight about it. The output of a language model is probabilistic, so it has to be validated rather than taken on sight. The mature pattern pairs automated extraction with confidence based routing: what passes the checks flows through, and anything uncertain or high stakes goes to a person for sign-off.
For a UK business there is an extra reason to care. Making Tax Digital already expects VAT records to live as digital data, moved between systems by digital links rather than re-keyed by hand, and mandatory e-invoicing for all VAT invoices has been announced for 2029. Structured document data is not just an efficiency play here; it is the direction the rules are moving.
Invoice and receipt capture for accounts payable
You give the model a PDF or a photo of an invoice and ask for the fields you need: supplier name, invoice number, date, line items, net, VAT, total, bank details. It reads the layout both visually and semantically, so it copes even though every supplier formats invoices differently, and it returns structured data that flows straight into your accounting or ERP system. Extractions that pass validation, the line items sum to the total, the supplier matches a known vendor, can post automatically; anything uncertain is flagged for a clerk.
A finance team receiving a few hundred supplier invoices a month stops keying them in by hand. Invoices land in the inbox, the AI extracts the fields, matches each one against the purchase order and posts the clean ones automatically, while the exceptions queue up for a person to review.
Manual data entry is the main bottleneck and cost in accounts payable. Ardent Partners' State of ePayables 2024 puts the average fully loaded cost of processing a single invoice at $12.88, against $2.78 for best-in-class operations, a gap of roughly 78%, and the average cycle time at 14.6 days against 3.1. The same research finds only about 30% of invoices are processed touchlessly industry wide, rising to around 49% among the best teams, so the headroom for automation is large and documented.
Contract review and clause extraction
The model reads a contract and pulls out the terms that matter: the parties, the effective date, the term, renewal and termination notice periods, liability caps, payment terms, governing law, and any clause that is non standard or risky. It can compare a contract against your own template and highlight the deviations, or answer a plain question such as when the agreement auto renews and how much notice you must give.
JPMorgan's COiN (Contract Intelligence) system was reported in 2017 to interpret commercial loan agreements that had previously absorbed about 360,000 hours of lawyer and loan officer time a year, extracting roughly 150 attributes from 12,000 agreements in seconds; the bank also said it cut servicing mistakes that came from people misreading contracts, without publishing a figure. A smaller firm runs the lighter version: drop a supplier contract in and get a one page summary of the key dates, obligations and risks.
A first pass that took hours of skilled reading takes minutes, renewal and notice deadlines stop slipping past unnoticed, and unfavourable or unusual terms are flagged for a person to confirm. Because contract decisions carry legal and financial weight, the AI accelerates the review; it never finalises it unattended.
Email triage and attachment extraction
Incoming email is classified by intent, an order, a complaint, an invoice, a request for quotation, and any attached documents are read and extracted automatically. The model routes the message, pre-fills a ticket or a record, and drafts a contextual reply, so the person handling it reviews and approves rather than starting from a blank page.
A shared inbox where orders, invoices and support requests all land together: the AI labels each thread, lifts the order details from the attached purchase order into the system, and prepares a draft response for the right colleague to check and send.
Less time reading and sorting an overflowing inbox, a lower chance of a message being misfiled or missed, and faster replies, because the relevant data and a draft are already prepared. Classification confidence does useful work here: obvious cases route automatically, ambiguous ones queue for a person.
Plain-language search across your document archive
Instead of opening files one by one, you index your contracts, manuals, policies and past invoices, then ask questions in plain English. The system retrieves the relevant passages and the model answers with the specific figure or clause, plus a pointer to the source document. The pattern is called retrieval-augmented generation (RAG): it grounds the answer in your real paperwork and lets the reader verify it against the cited source.
A director asks which supplier contracts allow a price increase this quarter and what notice they require, and gets an answer with citations pulled from the signed PDFs themselves, instead of emailing colleagues to dig through shared drives.
Institutional knowledge becomes searchable in plain language, which cuts the hours staff spend hunting for information and the dependence on the one person who knows where everything lives. The citations are not decoration: retrieval occasionally surfaces the wrong passage, and a model can phrase a wrong answer with complete confidence, so every answer must be checkable at the source.
Here is where the technology genuinely stands today, and where a level head still helps:
- Independent benchmarking backs the core claim. In AIMultiple's invoice OCR benchmark, leading language models handled a wide spread of invoice formats without any per supplier template setup, and a recent generation Claude Sonnet model showed the highest overall accuracy and the strongest resilience across the full range of document qualities. That is the practical edge over legacy OCR, which needs configuring for every layout, and it is exactly the kind of capability that can be built for your document flow now.AIMultiple Research, Invoice OCR Benchmark: extraction accuracy of LLMs vs OCRs
- The time savings are proven at serious scale: the JPMorgan COiN deployment cited above shows what contract extraction delivers when the volume is extreme, and vendors commonly report field level accuracy in the high nineties on clean, printed documents. A smaller organisation will not run at that volume, but the mechanics that make it work are the same ones a bespoke build uses.Bloomberg, JPMorgan software does in seconds what took lawyers 360,000 hours (COiN)ABA Journal, JPMorgan Chase uses tech to save 360,000 hours of annual work by lawyers (COiN)
- Email classification, question answering over your own archive and contract summarisation are mature enough to build now with a human review step. The wider pattern, pairing automated extraction with validation and exception routing, is standard, documented practice in intelligent document processing, not an experiment.AWS, What is Intelligent Document Processing (IDP)?
- One honest brake on the enthusiasm: promises of fully autonomous, 99%-plus accuracy across every document type do not survive contact with real paperwork. Accuracy drops visibly on poor scans, handwriting, unusual layouts and detailed line item tables, and the headline figures tend to come from other markets and from those selling the tooling. Take them as a direction to test, not a result to expect, and judge the build on your own numbers: fields corrected per hundred invoices, exceptions caught at intake, hours actually returned to the team.
Four limits worth knowing before you start, none of them a reason not to:
- Hallucination is measurable, not theoretical. A large 2026 study of document question answering, covering more than 172 billion tokens across 35 models, found that even the best performing models fabricated answers at roughly 1.19% in the best case with a 32,000 token context, and that the rate climbed as the context grew longer. The practical response is straightforward: run the extraction more than once and auto accept only values that agree, apply validation rules such as whether the arithmetic adds up and the date is plausible, and match against a source of truth like the purchase order.arXiv, How much do LLMs hallucinate in document Q&A scenarios? A 172-billion-token study
- Human sign-off is non-negotiable for anything financial, legal or regulatory. The reliable design is straight-through processing for high confidence, clearly formatted documents, with automatic routing of low confidence or unusual ones to a person who approves the exceptions. In the UK there is a sharper edge to this: Making Tax Digital for VAT already requires every VAT registered business to keep digital records and move data between systems by digital links, and responsibility for the accuracy of what reaches your VAT records stays with you, whoever, or whatever, did the typing.GOV.UK, VAT record keeping and Making Tax Digital for VAT
- On e-invoicing, know where the UK actually is. There is no general B2B e-invoicing mandate in force today, a genuine difference from much of the EU, but the government announced at Budget 2025 that all VAT invoices must be issued as e-invoices from April 2029, covering B2B and business to government, with the implementation roadmap due at Budget 2026. Coming, not current: there is nothing to buy yet, and anyone selling 2029 certainty now is ahead of the facts, but getting your document data structured today makes the eventual switch a formality.GOV.UK, Promoting electronic invoicing across UK businesses and the public sector: consultation response
- Contracts, HR paperwork and customer correspondence hold personal data, so the UK GDPR applies, enforced by the ICO with fines of up to £17.5 million or 4% of worldwide annual turnover at the higher maximum. That means controlling where documents are sent and processed, how long they are retained, and following the ICO's guidance on AI and data protection. One design rule on top: avoid AI loops that check themselves with no external ground truth, because a generating model and a checking model can share the same blind spot and reinforce the same mistake. AI for speed and coverage, people for judgement and sign-off on what matters.legislation.gov.uk, UK GDPR Article 83 (penalties)ICO, Guidance on AI and data protection
Procurement and suppliers
Procurement is how your company buys what it needs to operate: finding and qualifying suppliers, requesting and approving purchases, comparing quotes, raising purchase orders, negotiating prices and terms, signing contracts, and then managing each supplier relationship over time, from deliveries and invoices to performance and risk. Every business does it, whether that is a sole trader ordering software licences or a manufacturer dealing with thousands of vendors.
It deserves attention because bought-in goods and services are usually one of the largest cost lines in the accounts. A supplier who delivers late, ignores the agreed terms or quietly slides into financial trouble can bring your operation to a halt.
The day-to-day reality is also stubbornly manual. Quotes sit buried in email threads, invoices arrive in a dozen layouts, contracts get signed and never reread, and the long tail of small purchases is managed by nobody at all.
That mix of language, documents and judgement plays directly to what modern AI assistants (large language models such as Claude) do well: they take over the repetitive reading, matching and drafting, and surface the decisions that genuinely need a person.
Matching invoices to purchase orders and contracts
A language model reads every incoming invoice, whatever the format or layout, extracts the line items and checks them one by one against the purchase order and the underlying contract: quantities, unit prices, VAT, delivery charges, discounts and payment terms. Invoices that agree in full pass straight through. Only the ones with discrepancies reach a person, and they arrive with the mismatch already explained in plain English.
An electrical wholesaler in the Midlands receives 250 supplier invoices a month in a dozen different layouts. The assistant clears the ones that match the order and the contracted price, and flags one where a supplier has billed an old, higher unit price from before the renegotiation, so the overcharge is caught before payment rather than after it.
McKinsey describes a global manufacturer that used generative AI to automate invoice processing, cutting errors by about 80% and roughly halving processing time, and a global pharmaceutical company whose invoice-to-contract reconciliation proof of concept, built in four weeks, surfaced more than $10 million of value leakage to claw back through renegotiation. Your team stops keying invoices and concentrates on the minority that genuinely need review.
Reading contracts and flagging risky clauses
The assistant ingests supplier contracts, often long PDFs nobody rereads, and pulls out the terms that matter: renewal and termination dates, price and indexation clauses, payment terms, liability caps, auto-renewals, SLAs and penalty provisions. It then compares those terms against your standard playbook and flags anything off-market or risky, in language a manager can act on straight away.
Before signing, a manager pastes a supplier's 30-page service agreement and asks the assistant to compare it with last year's version and the company's standard terms. It surfaces a new auto-renewal clause, a change that removes the liability cap and a price rise tied to an index, so the manager can push back before signature instead of discovering it all at renewal.
Hours of careful reading become minutes, and off-market or auto-renewing terms get caught before they cost money. You also end up with a searchable inventory of obligations, so renewal dates and price clauses stop slipping past unnoticed. Gartner predicts half of procurement contract management will be AI-enabled by 2027, a sign of how mainstream this is becoming.
Guided purchase requests in plain English
Instead of forcing staff to learn a complicated ERP form, an AI assistant, often embedded in Teams or a web form, lets them describe what they need in ordinary language. It turns the request into a clean, system-ready requisition, steers it towards a preferred supplier, attaches the right budget codes, answers policy questions such as the approval threshold, and starts the correct sign-off workflow.
An office manager types that she needs five standby laptops for the new starters arriving next month. The assistant proposes the approved model and supplier, fills in the requisition, notes that the amount needs a director's sign-off and sets the approval in motion.
Off-contract (maverick) spending and approval delays both fall, because buyers are steered to compliant channels at the moment they ask. Gartner predicts that by 2027, 70% of procurement intake requests will be assisted by AI and generative AI, a signal that this is becoming the default way requests enter the system.
Drafting and triaging supplier correspondence
The assistant drafts the routine traffic that fills a buyer's inbox: requests for quotes, order confirmations, delivery-status chasers, escalations for late deliveries and onboarding document requests. It reads inbound supplier emails, summarises long threads, extracts commitments such as a promised dispatch date, pulls the quotes received into a single view so they can be weighed on like-for-like terms, and proposes a reply in the right tone and language for each supplier.
A buyer asks the assistant to chase eight suppliers for overdue order confirmations. It drafts eight polite, personalised emails referencing each order number and due date, ready to review and send, then later condenses the replies into one status view.
Hours of repetitive writing and reading come back, correspondence stays consistent and multilingual, and the chance of a thread or a commitment falling through the cracks drops sharply. That matters most in small teams, where one person often handles every supplier conversation alongside everything else.
A realistic picture of what can be built for you today: the proven, the emerging and the still-experimental:
- The document-heavy work is buildable now with confidence: invoice extraction and matching, contract clause extraction, supplier-risk triage and plain-English purchase intake are all running at real companies. Deloitte's Global CPO Survey found the large majority of procurement chiefs planning or assessing generative AI (92% in its 2024 reading), and McKinsey estimates the next wave of automation could make procurement operations 25 to 40% more efficient.DeloitteMcKinsey
- The direction of travel is clear from the analysts: Gartner predicts half of procurement contract management will be AI-enabled by 2027, and that 70% of procurement intake requests will be assisted by AI and generative AI by the same year.Gartner
- Autonomous AI negotiation is real but narrower. Walmart has used an LLM-based negotiation agent since 2023 for its long tail of smaller suppliers, reporting around a 3% average saving and payment terms extended by an average of 35 days, with roughly 75% of the suppliers involved saying they preferred negotiating with the AI. That evidence covers routine, well-bounded terms, not strategic, high-value deals, where a human buyer still leads.PYMNTS
- One honest health warning on the numbers: the headline results (the 3% saving, the 80% error reduction, the $10 million of recovered leakage) come from individual case studies and from consultant or vendor reporting. Read them as an indication of what is possible, not a promise, and measure any build against your own invoices, contracts and baseline before you scale it.
Procurement sits where money, contracts and supplier data meet, so a few disciplines are non-negotiable:
- Keep a person in the loop wherever money or legal commitment is at stake: the AI drafts the contract, recommends the negotiated deal and flags the invoice mismatch, but a human approves the spend, signs the contract and releases the payment. The UK deliberately has no general AI statute; it applies a pro-innovation, principles-based approach through existing regulators, which means responsibility for what an automated system communicates or commits to sits squarely with your business. And if your AI outputs are used in the EU, or you place AI systems on the EU market, the EU AI Act reaches you even though it does not apply in the UK itself.GOV.UKEU AI Act, Article 2 (Scope)
- Treat supplier contracts and pricing as confidential data. Do not paste them into consumer AI tools that may retain or train on the content; use an enterprise deployment with clear data controls. Where supplier files contain personal data, contact names, emails, signatures, the UK GDPR applies, enforced by the ICO with fines of up to £17.5 million or 4% of worldwide annual turnover, and the ICO publishes dedicated guidance on using AI with personal data responsibly.legislation.gov.ukICO
- Verify, do not trust. A language model can misread an unusual invoice layout or attribute a clause to a contract that does not contain it, so every extracted figure and flagged risk must be checkable against the source document. On the tax side, all VAT-registered businesses already keep digital records and file through software under Making Tax Digital, so AI-extracted invoice data must reconcile with what goes to HMRC; and since the government has announced that all VAT invoices must be issued as e-invoices from April 2029, any invoice pipeline you build now should be designed with that mandate in mind.GOV.UKGOV.UK
- Matching and risk scoring are only as good as your supplier master data and contract archive, which in many companies are messy, so budget for clean-up as part of any build. Set clear guardrails on anything automated: an unattended negotiation or auto-approval can push payment terms in ways that strain smaller suppliers or lock in unfair terms, so define walk-away limits and review outcomes regularly. And remember what AI does not replace: your judgement on supplier selection, ethics and long-term strategy. It removes the clerical load precisely so your people can spend more time on those decisions.
HR and recruitment
HR and recruitment is everything people-related in your business: attracting candidates and sifting through them, coordinating interviews, settling new starters in, and fielding the daily stream of questions about policy, annual leave, pay and benefits.
It is a process built on documents (CVs, job adverts, contracts, the staff handbook) and on constant back-and-forth (scheduling, follow-ups, reminders). Much of it is repetitive at the edges, yet the decisions at its core demand judgement.
That mix is exactly the load AI can carry. Language models are strong precisely where this process is weakest: they read and compare unstructured documents, draft text in your tone, run scheduling logistics, and answer questions from your own internal documents rather than the open internet.
The sensitive part is that hiring is regulated and ethically loaded. In the UK that regulation comes through existing law, the Equality Act 2010 and the ICO's data protection regime, rather than an AI statute, but it bites just as hard. So the safest, largest wins come from letting AI read, summarise, draft, schedule and answer, while a person keeps clear authority over who is hired, promoted or turned down.
CV screening support and evidence-backed shortlists
The AI reads every CV, whatever the layout, against the requirements of the role and produces a structured summary per candidate: relevant skills, years of experience, gaps, and a plain-English rationale for how well they fit. Instead of a black-box accept-or-reject score, a well-configured model extracts and normalises the messy detail and surfaces the evidence, which the recruiter then reviews. The recruiter still decides; the AI removes the hours of reading. Crucially, the system is built so the model cannot see or infer protected characteristics, ideally with the name, photo and address stripped before screening.
A 40-person firm in Manchester advertises one role and receives 300 applications. Overnight, the model reads all 300 CVs, marks against each candidate which must-have skills it found and which it did not, brings the strongest 25 profiles forward with a three-line justification apiece, and files the rest in a list the recruiter can revisit. The morning starts with 25 evidence-backed summaries instead of 300 raw PDFs, and any candidate the model passed over is one click away.
The most tedious, slowest part of high-volume recruitment disappears, and the recruiter gets consistent, comparable summaries backed by quoted evidence. The gain is real only while a human reviews the shortlist: automatic rejection without review is exactly where bias and Equality Act exposure concentrate, which is why well-built systems never turn a candidate down without a person confirming the decision.
Job adverts and candidate communication, drafted in your tone
From a short brief, the AI drafts job adverts, structured interview questions, polite rejection notes and offer emails, matching your company's voice and reusing adverts that have performed before. Drafting is generative and low-risk, because a person edits before anything is sent, which is why it is usually the first place HR teams adopt AI: the recruiter corrects a good first draft rather than facing a blank page. The model can also be asked to flag biased or exclusionary phrasing before an advert goes out.
A hiring manager types "senior backend engineer, Python, five years, remote-friendly, our usual benefits." The model returns a complete, inclusive job advert, an interview scorecard and three screening questions. HR adjusts two lines and publishes in minutes rather than an hour.
Recruiter time shifts from typing to the work that actually needs a person: conversations with candidates and calibrating the panel. Tone and inclusive language stay consistent across every advert and every rejection note, which matters when hundreds of applicants form their impression of your business from those messages.
Interview scheduling and coordination
An AI scheduling agent reads the interviewers' calendars, proposes times to candidates in natural language by email or chat, books the slot, sends reminders and reschedules when a conflict appears, involving a person only for exceptions. It turns the multi-message "what time suits you?" loop into a task that runs itself, and writes the outcome back into your applicant tracking system and calendars.
Once a candidate clears screening, the agent emails them slots that fit three interviewers' diaries, books the chosen one, creates the calendar invite with the video link and sends a reminder the day before. Edge cases, such as a candidate who needs an unusual adjustment or a panel that will not converge, are escalated to the coordinator.
Coordination that can swallow anywhere from half an hour to two hours per candidate largely runs without anyone touching it, and automatic reminders cut no-shows. Vendor claims about exact figures are directional rather than audited, so the honest measure is your own diary: hours spent chasing availability before and after.
Self-service answers on policy, leave and pay
Employees and new starters ask everyday HR questions in plain language and get answers drawn from your actual staff handbook, leave rules and benefits documents. The system retrieves the relevant policy passage and answers with a citation, so people can verify it and HR can trust the model is not improvising policy. Anything sensitive or ambiguous is handed cleanly to a named person, and the assistant is available outside HR hours and in several languages, which helps small HR teams supporting distributed staff.
An employee asks "how many days of annual leave can I carry over into next year?" The assistant retrieves the carry-over clause from the handbook, gives the specific answer and links the source paragraph. A question that touches a grievance, a harassment concern or an individual pay dispute is routed straight to a person and never answered by the assistant.
The daily volume of repetitive queries drops and employees get instant, consistent answers around the clock. Grounding every reply in a retrieved, cited company document is what prevents the model from confidently inventing policy, which is the main failure mode of this use case, and it is a design requirement, not an optional extra.
Where this stands today: what can be built with confidence, where the early advantage lies, and where honesty is due.
- The assist-and-organise layer is proven and buildable now. Drafting job adverts and candidate emails, summarising CVs, conversational interview scheduling and retrieval-based Q&A over your own handbook all rest on mechanics current models handle reliably: reading documents, drafting text, retrieving from a knowledge base and calling calendar and email tools. Workplace AI use in general is already broad; McKinsey reports about 76% of employees surveyed had used AI in some capacity at work in 2025, up from roughly 30% in 2023.McKinsey, Superagency in the workplace: AI in the workplace, a report for 2025
- Embedded use inside HR itself is still thin, which is the opportunity. McKinsey's European HR research found only around 19% of core HR processes enhanced with generative AI, with many efforts still at pilot stage. A business that wires AI properly into screening support, scheduling and policy Q&A is ahead of most of its market, not catching up to it.McKinsey, HR Monitor 2025 (gen AI adoption across HR processes)
- Hold the headline numbers loosely. The widely quoted claim that 83% of companies would use AI in hiring by 2025 comes from a late-2024 survey of 948 business leaders about intent, not from measured deployment, and the popular figures on time-to-hire falling by a third to a half are reported by the vendors selling the tools. Take them as direction, then judge the build on your own numbers: hours of CV reading saved, days from application to interview, queries answered without HR touching them.ResumeBuilder.com, 7 in 10 companies will use AI in the hiring process in 2025 (survey of 948 business leaders)
- What is not ready: fully automated scoring and rejection. University of Washington researchers testing three language models across more than 550 real CVs found the models favoured white-associated names about 85% of the time, and male-associated names 52% of the time against 11% for female-associated names. A separate study by the same university (528 participants) found people tend to mirror a biased AI's hiring recommendations rather than correct them, so a nominal human in the loop is not automatically a safeguard; the human must genuinely decide.University of Washington News, AI tools show biases in ranking job applicants' names (October 2024)University of Washington, People mirror AI systems' hiring biases, study finds (November 2025)
The UK has no AI act; it regulates AI in hiring through existing law, and that law is specific. Three regimes matter here: the Equality Act 2010, the ICO's data protection enforcement, and the UK GDPR's new automated decision-making rules in force since 5 February 2026.
- Under the Equality Act 2010, the employer carries liability for discriminatory outcomes of an AI screening tool it uses, even when a vendor built the tool. Direct discrimination is covered by section 13, and a facially neutral algorithm that disadvantages a group sharing a protected characteristic, such as age, disability, race or sex, engages indirect discrimination under section 19 unless it can be justified as a proportionate means of achieving a legitimate aim. This is why the model is kept away from protected traits and a person keeps deciding.legislation.gov.uk, Equality Act 2010
- The ICO has looked at recruitment AI directly and did not like everything it found. Its November 2024 audit of AI sourcing and screening tools uncovered tools filtering candidates by protected characteristics, tools inferring gender and ethnicity from names, excessive data collection and indefinite retention, and it issued almost 300 recommendations. The practical asks apply to any tool you deploy: process candidate information fairly, collect only what you need, run a data protection impact assessment, demand bias-testing evidence from the vendor, and tell candidates clearly how AI uses their information. The ICO's broader guidance on AI and data protection covers fairness, transparency and accountability across the whole AI lifecycle.ICO, AI tools in recruitment audit outcomes report (November 2024)ICO, Guidance on AI and data protection
- Since 5 February 2026 the UK GDPR has a new automated decision-making regime: the Data (Use and Access) Act 2025 replaced Article 22 with Articles 22A to 22D. A solely automated rejection of a job applicant is a significant decision, and it is now generally permitted for non-special-category data, a genuine divergence from the EU's stricter default, but only with the Article 22C safeguards in place: the candidate must be informed, able to make representations, able to obtain meaningful human intervention and able to contest the decision. Meaningful human involvement must be genuine, not a rubber stamp. Permitted is not the same as advisable: given the documented bias evidence and Equality Act liability, the defensible position is to take the speed the law allows in the routine layer and still keep a human decision on every rejection.legislation.gov.uk, Data (Use and Access) Act 2025, section 80 (new Articles 22A to 22D UK GDPR)ICO, The Data (Use and Access) Act 2025: what it means for organisations
- CVs and HR records are personal data under the UK GDPR, and ICO fines run to £17.5 million or 4% of worldwide annual turnover at the higher maximum. So the unglamorous groundwork is part of the build, not an afterthought: a written processor contract with any AI provider, a clear retention schedule for candidate data, and internal Q&A assistants grounded in your real documents with citations, with grievances and individual pay disputes always routed to a person.legislation.gov.uk, UK GDPR, Article 83 (administrative fines)
Internal knowledge and operations
Internal knowledge and operations is the everyday machinery of how your company actually runs: the SOPs, policies, wikis, resolved tickets and contracts, plus the unwritten know-how that lives only in people's heads, scattered across email, shared drives, chat threads and business systems, and the routine workflows that shuttle information between them. Every organisation has it, because every organisation has questions that need a reliable answer (what is our refund policy? how do we take on a new supplier?) and steps that must be repeated the same way every time.
It matters because that knowledge is usually trapped. A 2012 McKinsey Global Institute study put the average knowledge worker at nearly 20% of the working week searching for internal information or tracking down the colleague who holds it. The figure predates today's tools, but the underlying problem, knowledge that is scattered, unsearchable and dependent on one person, is still the norm in most organisations.
AI changes the economics here in three ways. It reads messy, unstructured content. It answers questions in plain language, with links back to the source so people can verify. And, increasingly, it carries out multi-step actions across the systems where the work actually happens, with a person supervising the steps that carry consequences.
Internal Q&A over your own documents (a grounded assistant)
You connect an AI assistant to the knowledge sources you already have: your wiki, shared drives, policy PDFs, resolved support tickets. When someone asks a question, the system first retrieves the most relevant passages, then the model writes a plain-language answer grounded in them, ideally with links back to the source so the person can check it. The pattern is called Retrieval-Augmented Generation (RAG): the assistant answers from your own curated content rather than from whatever it memorised in training, which keeps answers company-specific, current and auditable.
Uber built an internal Slack copilot called Genie that answers engineers' questions by retrieving from internal sources: their engineering wiki, an internal Stack Overflow and requirement documents. Per Uber's engineering blog, since launching in September 2023 it expanded to 154 Slack channels and answered over 70,000 questions, against a baseline of roughly 45,000 questions a month across those channels, saving an estimated 13,000 engineering hours that on-call staff would otherwise have spent fielding repetitive queries. Uber also reports a 48.9% helpfulness rate, a useful reminder that even a successful production assistant does not answer everything well and still needs a human fallback.
Cuts the time your people lose hunting for answers and takes repetitive questions off senior or on-call staff, who are otherwise a bottleneck. Because answers cite their sources, the assistant doubles as a way to find the authoritative document, not just a paraphrase of it.
Turning raw material into clean, structured SOPs
Instead of starting from a blank page, you give the AI the raw inputs you already have: a meeting transcript where someone explained the process, an email thread, an old checklist, a screen recording. You ask for a structured standard operating procedure with numbered steps, roles, inputs and edge cases, and a human owner reviews and approves it. The AI also keeps procedures consistent in tone and format, and rewrites them when the process changes, so documentation does not rot the moment the process shifts.
An operations manager records a 20-minute call walking through how month-end invoicing works, then asks the model to turn the transcript into a step-by-step SOP with a checklist and a list of exceptions. The same approach drafts onboarding guides, policy summaries and FAQ entries from scattered notes, with the owner correcting anything wrong before it is published.
Documentation actually gets written, because the effort drops from hours to minutes and the activation energy that usually leaves processes undocumented disappears. That attacks the single point of failure directly: critical process knowledge stops living only in one person's head.
Cross-system workflow automation with connected tools (MCP and agents)
Through standardised connectors, most prominently the Model Context Protocol (MCP), the AI can be given controlled, permission-scoped access to your live business systems: CRM, ticketing, calendar, email, accounting. It then reads from one system and acts in another as a multi-step workflow: look up a record, draft a document, create a task, update a status. High-risk steps, anything that writes to a system of record or sends an external message, sit behind human approval, and every action is logged.
A new sales enquiry arrives. The assistant pulls the record from the CRM, checks the calendar for availability, drafts a tailored follow-up email and creates a follow-up task, then pauses and waits for a person to approve before anything actually goes out. In a software organisation, MCP-connected agents query ticket systems and error logs to help on-call engineers during an incident.
Removes the copy-paste-and-context-switch tax of moving information between tools by hand, and lets routine operational chains run with a person supervising rather than executing every keystroke.
Faster onboarding and capturing unwritten know-how
A grounded Q&A assistant lets new starters ask questions in natural language and get cited answers from the same internal sources a veteran would point them to, at any hour, without interrupting a colleague. There is a useful side effect too: building the assistant forces scattered tribal knowledge into a documented, searchable form, because content that is missing or wrong becomes visible the first time someone asks about it.
A new hire in operations asks how a refund above £500 is handled, or who signs off a new supplier, and gets a cited answer drawn from internal policy instead of waiting hours for a busy teammate to reply on chat. Where the assistant cannot answer, that gap flags a missing or outdated document for the knowledge owner to fix.
Shortens ramp-up time and eases the load on experienced staff, while making the organisation more resilient when a key person is on leave, unavailable or moves on.
A realistic picture of what can be built for your organisation today, and where the marketing runs ahead of the evidence:
- Grounded internal Q&A is the most proven piece: the Uber Genie deployment cited above has run in production since September 2023 at a scale no pilot matches. A system of this kind can be built now on your own wikis, drives and resolved tickets, with answers that cite their sources.Genie: Uber's Gen AI On-Call Copilot (Uber Engineering Blog)Enhanced Agentic-RAG (Uber Engineering Blog, 2025)
- Drafting and summarisation with a human editor (SOPs, minutes, status reports, FAQ answers) are dependable and low risk precisely because a person reviews before anything is used. GitLab, in a pilot reported via Anthropic, saw 98% of surveyed participants satisfied and self-reported productivity gains of 25 to 50% on repetitive drafting; self-reported pilot figures, not an audited benchmark, but they show where the time savings concentrate.GitLab boosts productivity across teams with Claude (Anthropic customer story)
- Cross-system automation has matured unusually fast: MCP went from Anthropic's November 2024 launch to a cross-vendor de facto standard within about a year, adopted by OpenAI, Google, Microsoft and AWS and donated to a Linux Foundation effort in late 2025. Connectors to CRM, ticketing, email and calendar are real and usable today; the dependable pattern is the agent drafts and a person approves, especially for anything that writes to a system of record.One Year of MCP: November 2025 Spec Release (Model Context Protocol Blog)Donating the Model Context Protocol and establishing the Agentic AI Foundation (Anthropic)
- Where to stay sober: nobody gets a flawless, self-maintaining knowledge brain by pointing AI at a folder. Even Uber's successful deployment, by the helpfulness rate cited above, answers barely half of questions well, so the realistic goal is meaningful deflection with a human fallback, not every question answered. Treat published gains as somebody else's result, not a promise of yours, and measure deflection and answer quality on your own questions before scaling.Genie: Uber's Gen AI On-Call Copilot (Uber Engineering Blog)
In the UK the constraints come less from AI-specific law, which does not yet exist, and more from data protection and plain operational discipline.
- The commonest real-world failure is retrieval, not invention: if the right document was never found (missing, out of date, badly tagged, or siloed in a system the assistant cannot see), the AI cannot answer correctly however capable the model is. Stale knowledge in means confident-sounding wrong answers out, and an answer can be faithful to its source yet still incomplete, so track which questions the assistant fails or declines rather than assuming it works.The Ugly Truth About Enterprise RAG Evaluation (retrieval failure and faithful-but-useless answers)
- UK data protection applies in full. There is no UK AI act, but the ICO already regulates AI through the UK GDPR and its dedicated Guidance on AI and data protection, and an assistant that reads across systems must inherit and respect existing access permissions so it never surfaces salary data or confidential files to someone who should not see them. The fines behind that are the ICO maximums quoted throughout this guide.ICO, Guidance on AI and data protectionlegislation.gov.uk, UK GDPR Article 83 (fines)
- The wider rulebook in one line: the UK has no general AI statute, existing regulators police AI in their own lanes, and the EU AI Act still reaches you if your system or its outputs are used in the EU.GOV.UK, AI regulation: a pro-innovation approach (white paper)EU AI Act, Article 2 (Scope)
- Agentic actions need governance, and the whole system needs an owner. Anything that changes records or sends messages should require approval or run with tightly limited scopes and an audit log, and someone must be accountable for keeping the source content current and periodically evaluating answer quality. Without that, accuracy quietly degrades and the assistant slowly loses the team's trust.
Reporting and analytics
Every business runs on questions: what sold, what it cost, which customers are slipping away, whether the cash will stretch to the end of the quarter. The answers usually exist somewhere across the sales system, the accounts package and a stack of spreadsheets. Getting them out is the slow part.
In most firms that translation work sits with one or two people. Someone knows where the right numbers live, writes the query or the formula, and explains what the result means. Every routine question queues behind them, and when they are on holiday, reporting stops.
AI built on models like Claude targets exactly that layer. You ask in plain English, the system drafts the query, runs it, checks the result and replies in words, with a chart. It can also write the commentary on a dashboard, flag figures that break pattern, and assemble the monthly pack for a person to review rather than build.
One honest condition, well documented in production deployments: this is reliable only when the AI is given a curated map of what your data means. The model brings fluent language and reasoning; your organisation has to bring trustworthy definitions. Data quality sets the ceiling, not model cleverness.
Ask your data in plain English
You type a question the way you would ask a colleague: "which customers did we lose last quarter, and what did they have in common?". The AI reads a description of your data model (table and column names, definitions, vetted example queries), writes the SQL the database needs, runs it, inspects the result and corrects itself if something errors, then answers in words plus a chart. It is doing translation, not magic: its accuracy rises and falls with the quality of the data map it is given, not with raw model size.
Anthropic's internal data team reports that Claude-powered agents now handle about 95% of incoming business analytics queries at roughly 95% aggregate accuracy, freeing analysts for forecasting and causal work. Uber built a similar internal tool and reports cutting the time to produce a reliable query from about 10 minutes to about 3, an estimated saving of roughly 140,000 query-writing hours per month. In practice, a sales director can ask "top ten accounts by revenue growth this financial year" and have the answer in seconds, instead of raising a request and waiting for whoever owns the spreadsheet.
Self-service: non-technical staff get answers in seconds, and the analyst or the owner stops being the bottleneck for every routine question. The essential caveat: accuracy depends almost entirely on the data context provided, so this earns production trust only after definitions are curated and verified (see the maturity notes below).
Dashboards that explain themselves in words
Instead of a chart someone has to interpret, the AI reads the underlying numbers and writes a short plain-language summary: what moved, by how much, against last period and against target, and the most likely contributing factors visible in the data. The same layer can watch metrics over time and flag values that break the normal pattern, with a written explanation of what was detected and why it matters. It describes what the numbers show; it does not by itself prove causation, so the wording stays at the level the data supports.
A weekly trading summary writes itself: "Revenue is up 8% week on week, driven mainly by the South East (+22%); Scotland is down 5% for the third week running, concentrated in two trade accounts." Gartner predicts that by 2027, 75% of new analytics content will be contextualised through generative AI rather than presented as raw charts.
Faster comprehension and fewer missed signals; managers who never study charts closely still get the point. It also answers the familiar complaint of dashboards nobody opens, because the insight is pushed to people instead of waiting for them to dig.
Routine reports assembled for review, not built from scratch
Recurring reports (the monthly board pack, the weekly KPI digest, the regional sales summary) are drafted by an AI agent that pulls the numbers, applies your house format, writes the commentary and prepares the document or email. A person reviews and approves rather than assembling from zero each cycle, so the format and the underlying definitions stay consistent from one period to the next.
A monthly management pack that used to take an analyst the best part of a day drafts itself: tables populated, variances commented, executive summary written. The reviewer spends half an hour checking and adjusting tone instead of a day assembling. Time savings vary by report; the durable gain is moving the human from assembly to review.
Recurring time savings and consistency across periods, and the report still goes out on time when the one person who knows how to build it is away. In a small team, that directly reduces single-point-of-failure risk.
Forecasts and what-if scenarios, explained in plain language
Proper statistical or machine-learning models produce the forecast itself (demand, cash flow, headcount). The AI's job is to make it usable: explain the drivers, turn what-if questions into adjusted scenarios, and write the assumptions out in plain English. The language model is the interface and the explainer sitting on top of a vetted forecasting method, not necessarily the forecaster, which keeps the numbers grounded.
You ask "what happens to cash if our two biggest customers pay 30 days late" and get a re-run scenario with a written explanation of the impact and the assumptions used, instead of waiting for finance to rebuild the model by hand.
Planning conversations become interactive and self-serve; decision-makers explore options directly rather than queuing requests to a specialist. The underlying model stays the source of truth, and a forecast is read as direction with stated assumptions, never as a promise.
A clear-eyed view of what can be built solidly now, and where the honest limits sit:
- The best-documented finding in this field is that accuracy comes from context and verification, not from model power. In Anthropic's own deployment, agents with raw data access scored no better than around 21% on internal evaluations; adding structured skills that route the model to trusted definitions lifted accuracy consistently above 95% in aggregate, and to around 99% in some domains. That is the difference between a toy and a tool, and it is entirely buildable now.Anthropic, How Anthropic enables self-service data analytics with Claude
- The honest counterweight is the gap between demos and enterprise reality. On the older, tidy Spider 1.0 benchmark, text-to-SQL looks near-solved, above 90% accuracy on small databases averaging fewer than 10 tables. On Spider 2.0, which uses massive real schemas averaging hundreds of columns, multiple SQL dialects and external business logic, the leading models succeed only about 21% of the time, with BIRD, a midpoint of more realistic databases, around 73%. Pilot on your own schema before you trust anything.Spider 2.0, Evaluating Language Models on Real-World Enterprise Text-to-SQL WorkflowsTowards Data Science, Why 90% Accuracy in Text-to-SQL is 100% Useless
- Production systems reflect that gap and show what closing it takes. Uber's internal QueryGPT measured only about 50% table-overlap with ground truth on its own evaluation set, which is why heavy context engineering, a prompt enhancer and curated example sets were needed before the tool was trustworthy. Once that groundwork was done, query time fell from about 10 minutes to about 3, saving an estimated 140,000 query-writing hours per month. The work has shifted from buying a smarter model to investing in clean definitions and a semantic layer the AI can rely on.Uber Engineering, QueryGPT: Natural Language to SQL Using Generative AI
- On adoption, a Gartner survey of 403 analytics and AI leaders (October to December 2024) found over 50% already using AI for automated insights and natural-language queries, and Gartner predicts 75% of new analytics content will be contextualised through generative AI by 2027. Read those figures as direction of travel, not as a guarantee for your firm: analyst projections and headline deployment numbers come from other organisations and from parties with something to sell, so measure the gain on your own reports before you bank it.Gartner, Top Data and Analytics Predictions (survey of 403 leaders, Oct to Dec 2024)Gartner, 75% of analytics content to use GenAI for contextual intelligence by 2027
Keep a person accountable for anything that drives a real decision. The UK has no general AI statute; the rules come from existing regulators, and they still apply in full:
- Confident wrong answers are the core risk: a language model will happily return a plausible figure from the wrong column, or with a silent join error, and you cannot tell a right answer from a wrong one just by looking. Definitions, tests and spot-checks matter more than the model, and rigour has a price: in Anthropic's deployment, adding adversarial self-review improved accuracy by about 6% but cost roughly 32% more tokens and 72% more latency. The practical rule: let the AI draft, explain and accelerate, but have a named person own and approve the numbers that go to clients, the board or the tax authority.Anthropic, How Anthropic enables self-service data analytics with Claude
- Connecting AI to customer and financial data is governed in the UK by the UK GDPR and the Data Protection Act 2018, enforced by the ICO, whose Guidance on AI and data protection covers fairness, transparency, lawfulness, accountability and when a data protection impact assessment is required. You need a lawful basis, access controls and an audit trail before an AI reads personal data, and the higher maximum penalty is £17.5 million or 4% of total worldwide annual turnover.ICO, Guidance on AI and data protectionlegislation.gov.uk, UK GDPR Article 83 (penalties)
- If analytics output feeds automated decisions about individuals, staff scoring, customer credit terms, pricing tied to a person, a specific regime applies. Since 5 February 2026, new UK GDPR Articles 22A to 22D (inserted by the Data (Use and Access) Act 2025) generally permit solely automated significant decisions, but only with safeguards: the person must be informed, able to make representations, able to obtain meaningful human intervention and able to contest the decision. Meaningful involvement has to be genuine, not a rubber stamp.legislation.gov.uk, Data (Use and Access) Act 2025, section 80ICO, The Data (Use and Access) Act 2025: what it means for organisations
- Numbers that flow into tax reporting carry their own obligations. Making Tax Digital for VAT already applies to all VAT-registered businesses (digital records, returns filed through compatible software), Making Tax Digital for Income Tax starts in April 2026 for sole traders and landlords with qualifying income over £50,000, and the government has announced that from April 2029 all VAT invoices are to be issued as e-invoices, with the implementation roadmap due at Budget 2026. AI can prepare and reconcile the figures, but what is filed with HMRC should always pass through human review and approval.GOV.UK, VAT record keeping and Making Tax Digital for VATGOV.UK, Making Tax Digital for Income Tax for sole traders and landlordsGOV.UK, e-invoicing consultation response (26 November 2025)
We build them, on Claude
These AI flows do not stay on paper. Svennis Cloud Solutions builds and integrates them into your systems, with a team of certified Claude architects, on Anthropic technology, from the first WhatsApp message to the finished invoice.
See how this could work in your business
Describe what your company does in one sentence and Claude will show you, live, where AI could genuinely help your business.
Sources
- 1. McKinsey, An unconstrained future: how generative AI could reshape B2B sales
- 2. Harvard Business Review, The Short Life of Online Sales Leads (2011)
- 3. MIT / InsideSales Lead Response Management Study (Dr James Oldroyd, MIT Sloan)
- 4. ICO, Guidance on AI and data protection
- 5. American Bar Association, BC tribunal confirms companies remain liable for AI chatbot information (Moffatt v Air Canada)
- 6. GOV.UK / CMA, Unfair commercial practices guidance (CMA207)
- 7. legislation.gov.uk, Data (Use and Access) Act 2025, section 80 (new UK GDPR Articles 22A to 22D)
- 8. legislation.gov.uk, UK GDPR Article 83 (maximum fines)
- 9. GOV.UK, AI regulation: a pro-innovation approach (white paper)
- 10. EU AI Act, Article 2 (Scope)
- 11. CoSchedule, State of AI in Marketing Report 2025
- 12. Typeface, How Enterprise Marketing Teams Use Generative AI
- 13. ActiveCampaign / Talker Research, 13 Hours Back Each Week
- 14. McKinsey, Unlocking the next frontier of personalized marketing
- 15. McKinsey, What is personalization?
- 16. BizIQ, AI in Marketing Statistics 2026 (aggregated survey figures)
- 17. Bain & Company, Consumer reliance on AI search results
- 18. a16z, How Generative Engine Optimization (GEO) Rewrites the Rules of Search
- 19. Springer, AI Hallucinations in Marketing: Risks, Impacts, Mitigation
- 20. NeuralTrust, The Risk of AI Hallucinations: How to Protect Your Brand
- 21. ASA/CAP, Disclosure of AI in advertising (29 May 2025)
- 22. NBER, Generative AI at Work (Brynjolfsson, Li, Raymond): 5,179 agents, 14% average productivity gain, 34% for novices
- 23. Lyft blog, Lyft and Anthropic team up (Claude via Amazon Bedrock; 87% reduction in average resolution time)
- 24. Klarna press release, AI assistant handles two thirds of customer service chats in its first month
- 25. Intercom, From resolutions to outcomes (Fin: 40M+ resolutions, around 67% rolling 30 day resolution rate)
- 26. TechCrunch, Klarna CEO says company will use humans to offer VIP customer service
- 27. Real-time analytics and AI for managing no-show appointments in primary care (JMIR Formative Research, 2025)
- 28. Same UAE study, full text mirror (NCBI/PMC)
- 29. AI voice agents for appointment scheduling in clinics (Retell AI, vendor landscape)
- 30. ScheduleMe: multi-agent calendar assistant architecture (arXiv)
- 31. El Rio Health automated reminders case study, 32% no-show reduction (Emerging Global, vendor-reported)
- 32. Missed appointments cost the US healthcare system $150B a year (HCI Innovation Group, industry estimate)
- 33. The guide to AI in field service management (IBM)
- 34. AI-powered scheduling for field service (FieldCamp, vendor source)
- 35. Guide to hallucinations in large language models (Lakera)
- 36. ICO, The Data (Use and Access) Act 2025: what it means for organisations
- 37. Parseur, AI Invoice Processing Benchmarks 2026
- 38. Artsyl, Automated Invoice Processing 2025-2026: AI with Human Oversight
- 39. Ledge, AI reconciliation: real-world use cases
- 40. BILL, New AI agents for touchless transactions (Reconciliation Agent)
- 41. Accounting Today, Pilot launches fully autonomous AI bookkeeper (Feb 2026)
- 42. Baytech Consulting, Hidden dangers of AI hallucinations in financial services
- 43. GOV.UK, VAT record keeping (Making Tax Digital for VAT)
- 44. GOV.UK, Making Tax Digital for Income Tax for sole traders and landlords
- 45. GOV.UK, E-invoicing consultation response (26 November 2025)
- 46. AIMultiple Research, Invoice OCR Benchmark: extraction accuracy of LLMs vs OCRs
- 47. Bloomberg, JPMorgan software does in seconds what took lawyers 360,000 hours (COiN)
- 48. ABA Journal, JPMorgan Chase uses tech to save 360,000 hours of annual work by lawyers (COiN)
- 49. AWS, What is Intelligent Document Processing (IDP)?
- 50. arXiv, How much do LLMs hallucinate in document Q&A scenarios? A 172-billion-token study
- 51. Deloitte - 2025 Global Chief Procurement Officer Survey (PDF)
- 52. McKinsey - Transforming procurement functions for an AI-driven world
- 53. Gartner - Half of procurement contract management will be AI-enabled by 2027
- 54. PYMNTS - Walmart reportedly finds 75% of vendors prefer negotiating with chatbot
- 55. McKinsey, Superagency in the workplace: AI in the workplace, a report for 2025
- 56. McKinsey, HR Monitor 2025 (gen AI adoption across HR processes)
- 57. ResumeBuilder.com, 7 in 10 companies will use AI in the hiring process in 2025 (survey of 948 business leaders)
- 58. University of Washington News, AI tools show biases in ranking job applicants' names (October 2024)
- 59. University of Washington, People mirror AI systems' hiring biases, study finds (November 2025)
- 60. legislation.gov.uk, Equality Act 2010
- 61. ICO, AI tools in recruitment audit outcomes report (November 2024)
- 62. Genie: Uber's Gen AI On-Call Copilot (Uber Engineering Blog)
- 63. Enhanced Agentic-RAG (Uber Engineering Blog, 2025)
- 64. GitLab boosts productivity across teams with Claude (Anthropic customer story)
- 65. One Year of MCP: November 2025 Spec Release (Model Context Protocol Blog)
- 66. Donating the Model Context Protocol and establishing the Agentic AI Foundation (Anthropic)
- 67. The Ugly Truth About Enterprise RAG Evaluation (retrieval failure and faithful-but-useless answers)
- 68. Anthropic, How Anthropic enables self-service data analytics with Claude
- 69. Spider 2.0, Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
- 70. Towards Data Science, Why 90% Accuracy in Text-to-SQL is 100% Useless
- 71. Uber Engineering, QueryGPT: Natural Language to SQL Using Generative AI
- 72. Gartner, Top Data and Analytics Predictions (survey of 403 leaders, Oct to Dec 2024)
- 73. Gartner, 75% of analytics content to use GenAI for contextual intelligence by 2027