AI automation pilot checklist

Babagana Zannah
Leads engineering teams; puts AI to work inside one of the UK's largest companies

A useful AI automation pilot tests one bounded workflow against a recorded baseline, explicit failure cases and named human controls. Limit data and permissions, collect comparable evidence, and decide in advance what would justify expansion, repair or stopping.

An AI automation pilot should answer a decision, not produce an impressive demonstration. The practical question is whether one defined workflow can operate within an agreed boundary, with evidence that people can understand and controls that still work when the happy path breaks.

Use this checklist to plan a small UK business pilot. It is not legal, regulatory, cyber-security or data-protection advice, and it cannot determine whether a particular AI use is lawful or appropriate. Bring in competent specialists where the workflow affects people, regulated work, confidential information or important decisions.

1. Write the decision the pilot must support

Start with one sentence: “At the end of this pilot, we will decide whether to stop, repair or expand this workflow.” Name the decision owner and date.

Then define the proposed operating boundary. State what the system may read, propose, classify or write; what it must never do; and what remains a human decision. If the team cannot describe that boundary without referring to a product demo, the pilot is not ready.

The UK government’s Responsible AI Toolkit groups guidance for organisations developing and deploying AI responsibly. Treat responsible use as part of the project design, not a final approval label.

2. Select one workflow, not a department

A manageable first pilot has:

  • a repeatable trigger and finish;
  • a current source of truth;
  • a named process owner;
  • observable inputs and outcomes;
  • known exceptions;
  • reversible or containable actions; and
  • enough volume to produce representative evidence.

“Improve operations with AI” is not a pilot scope. “Draft an internal summary from an approved, non-client test set for human review” is closer to one. Exclude adjacent work explicitly so enthusiasm does not expand permissions mid-pilot.

3. Record the current baseline

Measure the workflow before changing it. Define units consistently: cases completed, elapsed time, active staff minutes, rework, exceptions, complaints, missed service targets and other outcomes relevant to the process.

Do not compare the pilot’s best week with an undocumented memory of the old process. Record the period, sample and known differences. Separate staff time from elapsed time, and separate measured results from forecasts.

4. Map the data and authority boundary

List each input, where it comes from, who owns it, whether it contains personal or confidential information, and whether it may be sent to the proposed service. Record retention, access, deletion and onward-use rules.

The ICO’s AI accountability and governance guidance explains the importance of involving information-governance roles early and documenting risks, mitigations and residual risk. Check the live guidance for the actual processing; a pilot label does not remove data-protection obligations.

Use synthetic or deliberately minimised inputs where they can answer the test. Do not upload real records merely to make a demo look realistic.

5. Name owners and stop conditions

Assign responsibility for business scope, data, technical configuration, output review, security, exceptions and the final decision. One person may hold several roles, but an unattended shared inbox is not an owner.

Write stop conditions before launch. Examples include:

  • wrong identity or source;
  • unexpected sensitive information;
  • output outside the approved task;
  • repeated unsupported answers;
  • permission or write ambiguity;
  • a material supplier change;
  • an incident or complaint; and
  • reviewers unable to keep up.

Each stop needs a safe state, evidence record and person authorised to restart.

6. Build representative test cases

Include ordinary cases, boundary cases and expected failures. Test missing fields, contradictory instructions, duplicates, inaccessible systems, unusually long content, malicious or irrelevant text, service timeouts and a reviewer who disagrees.

Specify the expected action for each case. “The AI handles it” is not an expected result. Use observable states such as propose, abstain, route, reject, request clarification or wait for approval.

7. Limit access and actions

Give the pilot the minimum access needed for the shortest useful period. Separate read, draft and write permissions. Use a test environment or isolated dataset where possible, and avoid long-lived credentials.

The NCSC’s secure-deployment guidance covers access controls, protection of models and data, incident procedures, evaluation, known limitations and secure user guidance. Apply established cyber-security practice as well as AI-specific controls.

8. Run shadow-first where practical

In shadow mode, the system produces a proposed result without taking the live action. A reviewer compares it with the real process and records agreement, disagreement, missed context and time required.

Shadow results can reveal whether the task definition, data or review interface is weak. They do not prove that live writes are safe, because production permissions and failure consequences differ. Move to a bounded live step only through an explicit decision.

9. Measure the whole operating loop

Count correction, monitoring and exception effort—not only seconds saved on generation. Record:

  • eligible and excluded cases;
  • correct proposals and meaningful abstentions;
  • errors by consequence, not just count;
  • reviewer changes and reasons;
  • unresolved and delayed exceptions;
  • staff time across setup, review and recovery;
  • supplier and infrastructure cost; and
  • incidents, complaints and near misses.

The government’s Introduction to AI assurance describes assurance techniques such as impact assessment, audit and performance testing, supported by governance, responsibility, escalation and quality assurance. Choose evidence that matches the actual risk and decision.

10. Make a documented go, repair or stop decision

Use thresholds agreed before the results were known. A pilot may be technically accurate yet fail because review costs are too high, exceptions are unsafe, users cannot challenge it or the data boundary is unacceptable.

Record one outcome:

  • Stop: the use is unsuitable or risk cannot be reduced within the agreed boundary.
  • Repair: a specific process, control or evidence gap must be resolved before another bounded test.
  • Expand cautiously: the evidence supports a defined next scope with new limits and monitoring.

Do not turn a pilot into permanent production by leaving it running. Production needs an owner, change process, monitoring, incident route, periodic review and an exit plan.

For choosing a candidate workflow, see How can AI help my small business?. For budgeting the complete operating loop, use How much does AI automation cost?.

Sources and review status

Author: Babagana Zannah. Published 13 August 2026 and last updated 13 August 2026. Sources checked 13 August 2026. Next editorial review due 13 November 2026.

These sources were re-opened on the source-check date. Recheck them sooner when law, official guidance, the supplier or the proposed processing changes.

Questions owners ask

How long should an AI automation pilot run?

Long enough to cover representative normal work and important exceptions, not an arbitrary number of days. Set the volume, scenarios and decision date before starting, then extend only for a documented evidence gap.

Which workflow should be used for a first AI pilot?

Prefer a repetitive, well-understood and reversible workflow with a reliable source, clear owner and low consequence of delay or error. Avoid starting with safety-critical, high-impact or poorly defined decisions.

Does a successful demo prove an AI workflow is ready?

No. A demo usually shows a prepared happy path. A pilot should also test incorrect, incomplete, adversarial and ambiguous inputs, system failures, human escalation and recovery under representative operating conditions.

What evidence should an AI pilot retain?

Keep the approved scope, baseline definition, test cases, version and configuration, input classifications, outputs, human decisions, exceptions, incidents, time and cost measures, and the final decision with its owner.

Not ready to talk? Take the scorecard.

The MTD Client-Chasing Readiness Scorecard gives you an indicative fit tier, likely bottleneck and sensible next step. No sign-up is needed to see the result.

Take the readiness scorecard

Your answers are assessed in your browser. The result is indicative, not tax, accounting, legal or regulated advice.

Got a repetitive job in mind?

Tell us about it on a fit call. If an AI helper isn't the right answer, we'll say so — and point you at the simpler option.

Book an AI workflow assessment