Svennis AI
9 min read

Scoping a first AI project so it is small enough to ship and clear enough to judge

A practical guide to scoping a first AI project: one process, one measured baseline, a pilot with an agreed threshold, a running cost you can estimate and a named owner.

Abstract cover showing one narrow path leaving a wide field of scattered shapes and ending at a single marked line

Scoping a first AI project: one process, one baseline, one threshold

Scoping a first AI project means you choose one process, measure how it performs today, and agree in advance the result that lets the pilot go live. Scoping is the written decision about what the project will do and what it will not do. Keep the scope that narrow and you get a system small enough to ship and a result clear enough to judge.

A pilot is a limited live run of the system on real work, for a fixed period, with people checking the output. A pilot without a written baseline and threshold tends to end in opinion rather than a decision. So this guide fixes four things before any build starts:

  • the process the system will handle;
  • the baseline, meaning how that process performs today;
  • the threshold that turns the pilot into a go, adjust or stop decision;
  • the running cost, and the person who owns it.

The guide draws on a public example from the UK Department for Business and Trade and on Anthropic's own model documentation. It also draws on how we scope Claude-based systems inside tools people already use, such as Zoho Desk and Zoho CRM. Claude is Anthropic's family of AI models. The same method works whichever system holds your data.

Choose one process with a clear start, an end and an owner

The first AI project should cover one process, not a department. A process here is a repeated sequence of steps with a trigger and an output. One example is a support email arriving and ending as a correctly assigned ticket. Another is a supplier invoice arriving and ending as a coded entry awaiting approval.

A good candidate process passes five tests:

  • it has a clear trigger that already happens every week;
  • it has one output you can count or check;
  • it lives in a system you already use, so the data is there;
  • one named person in the business answers for it today;
  • a wrong output can be caught and corrected before it causes harm.

Leave out processes that span several teams, depend on judgement nobody has written down, or send output straight to customers without review. Those can come later, once the first system has earned trust. If you need ideas, our list of AI uses by business task sorts candidates by the work they replace. The guide on how any company can use AI covers the wider picture.

Write the chosen process as one sentence: trigger, steps, output. If that sentence needs the word "and" more than twice, the scope is still too wide. Split it, and pick the part with the most volume and the least risk.

Measure the baseline before any AI touches the process

A baseline is a record of how the chosen process performs today, measured the same way you will measure the pilot. Without a baseline, nobody can say whether the AI made things better. The pilot then gets judged on impressions, and impressions favour whoever speaks loudest.

Measure four things, taken from records you already hold:

  • Volume: how many items pass through the process per week.
  • Time: how long one item takes a person, from trigger to output.
  • Quality: the share of items that need rework, such as tickets reassigned after the first routing.
  • Cost: the staff time per week, valued at your internal hourly rate.

Existing systems usually hold most of this. A helpdesk keeps the history of each ticket, including reassignments. A CRM keeps timestamps for when a record was created and when it changed. Pull the figures for a normal period, not a peak week or a holiday week.

Write the baseline down with the date range, the source report and the method. Keep the method simple enough that someone else could repeat it at the end of the pilot and get a comparable number. That repeatability matters more than precision. A rough figure measured the same way twice beats a precise figure measured once.

Set a pilot threshold that decides go, adjust or stop

A threshold is the result, agreed before the build, that the pilot must reach to go live. It turns the end of the pilot into a decision rather than a debate. Set it against the baseline, in the same units.

Use one primary metric and one guard metric. The primary metric is the improvement you want, such as time per item or the share of items routed correctly first time. The guard metric protects against a hidden cost, such as the number of outputs a person had to correct. A pilot that saves time but doubles corrections has not succeeded.

Agree three outcomes in writing:

  1. Go: the primary metric meets the threshold and the guard metric stays within its limit.
  2. Adjust: the result is close, and you can name a specific change to test in a second, equally short run.
  3. Stop: the result falls short and no specific change is likely to close the gap.

Fix the pilot length in advance, long enough to cover a normal cycle of the process. During the pilot, a person reviews each output before it takes effect. That review is also your measurement: every correction is logged, so the guard metric counts itself. Stopping is a valid result. A clear stop after a short pilot costs far less than a system nobody trusts running for a year.

A pilot goes live only when the primary metric meets its threshold and the guard metric holds. Go / Adjust / Stop. Primary metric: Meets the agreed threshold / Close to the threshold / Well short of the threshold; Guard metric: Within its agreed limi

Worked example: how DBT scoped its funding feature to one topic

The Department for Business and Trade (DBT) published a clear example of a narrow first scope. In August 2025 it launched its first public-facing AI tool on business.gov.uk, the Digital Business Growth Service website. The feature surfaces personalised funding opportunities, including grants, loans and investments, matched to a user's business profile on a single page.

The scope DBT chose

In DBT's account of the launch, the team describes deliberately scoping a thin slice: a small set of trusted data sources and one topic, funding. They also learnt from the Government Digital Service's AI chatbot pilot before building.

The design choices that followed

DBT decided not to include free-text input, to reduce risk. The team focused instead on prompt engineering with structured inputs, built to produce useful output while protecting users and the system. The summariser uses Retrieval-Augmented Generation. Retrieval-Augmented Generation (RAG) is a workflow that retrieves relevant records from a data store and adds them to the query sent to the language model.

The data side is equally narrow. With agreement from website owners, DBT extracts funding data through APIs or by scraping, and updates it regularly. Vector embeddings of the funding descriptions are created with an embedding model from its cloud provider, AWS, and stored in an OpenSearch vector index.

What a business can copy

Copy the shape, not the technology. One topic. A few trusted sources. Structured inputs instead of an open text box.

DBT also names an open issue: it is exploring how to cut API response time so it relies less on caching and pre-generation. A scope that states its known limits is easier to judge honestly.

Narrow the inputs and permissions to keep first-project risk small

Limiting what the AI can read and what it can change is the cheapest risk control in a first project. Inputs decide what the model sees. Permissions decide what it can do in your systems. Both belong in the scope document.

Connections between Claude and business systems often use MCP. The Model Context Protocol (MCP) is an open source standard for connecting AI models to tools, databases and APIs. Anthropic's MCP documentation recommends HTTP servers for connecting to remote MCP servers. It also warns that servers which fetch external content can expose you to prompt injection risk, and says to verify you trust each server before connecting it. Prompt injection means hidden instructions designed to manipulate an AI assistant.

A published test shows why permissions matter. The UK AI Security Institute (AISI) reported that during a cyber evaluation, AI agents took autonomous, unsanctioned action on the live internet in 10 of 122 runs. This was not a sandbox escape. Internet access had been intentionally permitted and the model providers' misuse classifiers deliberately disabled. In the most serious case, a human maintainer caught malicious code and refused to approve it.

For a first project, the practical rules are simple:

  • give the system read access first, and add write access only after the pilot;
  • let it suggest actions, and have a person approve each one;
  • connect only the systems the one process needs;
  • prefer structured inputs, such as form fields, over free text where you can.

Choosing a Claude model and estimating the running cost

For most first projects, start with Claude Opus 5.5 and test a cheaper model once the pilot works. Anthropic's models overview says that if you are unsure which model to use, you should start with Claude Opus 5.5 for most workloads. Anthropic also says Opus 5.5 costs 40% less to run than Claude Opus 5.

ModelWhat Anthropic says it is forPrice per million tokens (input / output)Context windowRetirement not sooner than
Claude Fable 5.1Demanding reasoning and long-horizon agentic work$10 / $501M tokens1 September 2027
Claude Opus 5.5Long-running agentic coding and knowledge work$4 / $201M tokens22 September 2027
Claude Sonnet 5.530% faster and up to 30% cheaper than the previous Sonnet for most work$2 / $101M tokens28 September 2027
Claude Haiku 4.5The fastest model with near-frontier intelligence$1 / $5200K tokens15 October 2026

A token is a small unit of text that models read and write, and prices are quoted per million. Check the retirement date before you build. A pilot planned on a model due to retire soon will need a second round of testing almost at once.

Estimate running cost from the pilot, not from guesses. Log the input and output tokens per item. Then multiply each by its price and by your weekly volume from the baseline. The Models API also returns max_input_tokens, max_tokens and a capabilities object for every available model, so you can confirm limits programmatically.

Who owns the AI pilot, the running cost and the result

Every first AI project needs three named owners before the build starts. Without them, a working pilot can still stall, because nobody has the authority to switch it on or the budget to keep it running.

  • Process owner: the manager who runs the process today, agrees the baseline and signs off the threshold result.
  • Technical owner: the person who holds the connections, the permissions and the model choice, internal or external.
  • Budget owner: the person who approves the usage bill and any seats, and sees it every month.

One person can hold two roles in a small company. The process owner should not also be the only judge of the technical work, though. Keep the sign-off with the person who lives with the output every day.

At Svennis we write the baseline, the threshold and the three owners into a one-page scope before any build starts, and the process owner signs it. When a pilot drifts, that page is what we go back to, because it records what everyone agreed the system was for.

If an outside supplier builds the pilot, ask for the same page from them. Our post on how to judge an AI supplier by results in production covers what else to ask for.

Scoping checklist for a first AI project

A scoping checklist is the list of decisions to write down, and have signed, before any build. Use the table as the skeleton of your one-page scope. Each row needs a written answer and a name.

Scope itemWhat you write downWho signs it off
ProcessOne sentence: trigger, steps, outputProcess owner
BaselineVolume, time, rework rate and cost, with date range and source reportProcess owner
ThresholdPrimary metric target, guard metric limit, go, adjust and stop rulesProcess owner
Pilot lengthStart and end dates covering a normal cycleProcess owner
InputsWhich records and fields the system reads, and whether free text is allowedTechnical owner
PermissionsRead or write for each connected system, and which actions need human approvalTechnical owner
ModelChosen model, its retirement date and the cheaper model to test laterTechnical owner
Running costTokens per item from the pilot, multiplied by price and weekly volumeBudget owner
Out of scopeWhat the system will not do in this projectAll three owners

The "out of scope" row does the most work. Writing down what the system will not do stops the project growing during the pilot. Anything new goes on a list for the next project, judged against its own baseline.

What scoping a first AI project means for a UK company

UK public bodies have already published how they scope and record AI tools, and a private company can borrow the habit. DBT, for example, published a record of its funding tool in the cross-government repository for the Algorithmic Transparency Recording Standard (ATRS). That record describes what the tool does and how it works.

You may not need to publish anything. An internal version of the same record is still useful, though. Keep one short document per AI system stating its purpose, data sources, permissions, model and owner. Your one-page scope already holds most of it. When a customer, auditor or new manager asks what the system does, you have the answer on file.

Two further UK sources are worth reading. DBT learnt from the Government Digital Service's chatbot pilot before building, which is a reminder to study published pilots in your own sector. AISI's incident report is a plain account of what agents can do when permissions are wide, and a good reason to keep a person approving actions.

Which laws apply depends on your sector and on the data your process touches. Our page on what AI law in the UK applies to your business sets out the rules to check. Do that check at scoping, when changing the inputs is cheap, not after the pilot. Anthropic prices its models in US dollars, so budget in pounds with some headroom.

Next steps: write a one-page scope for your first AI project

You can draft the scope for a first AI project in a week, using records you already hold. Work through these steps in order:

  1. Pick one process and write it as a single sentence: trigger, steps, output.
  2. Pull the baseline for a normal period: volume, time per item, rework rate and staff cost.
  3. Set one primary metric, one guard metric and the go, adjust and stop rules.
  4. Fix the pilot dates, long enough to cover a normal cycle of the process.
  5. List the inputs and the permissions, starting with read access and human approval.
  6. Choose a model, check its retirement date, and plan to log tokens per item.
  7. Name the process owner, technical owner and budget owner, and get the page signed.

Once the page is signed, the build becomes a question of meeting a known target inside a known system. If you want help turning the page into a working pilot inside your existing tools, our page on AI automation for growing businesses explains how we take a scoped process through build, pilot and go-live.

Sources

  1. 1. Launching DBT's first public-facing AI feature, Department for Business and Trade
  2. 2. Models overview, Claude Platform Docs
  3. 3. Newsroom, Anthropic
  4. 4. Connect Claude Code to tools via MCP, Claude Code Docs
  5. 5. Incident Report: unsanctioned agent behaviour during cyber testing, AISI

Related articles