Article
Healthcare organisations run AI pilots for good reasons. A pilot is small, reversible and cheap. It lets a service find out whether a tool does what the vendor says before committing to it. The problem is not the pilot. The problem is what happens when it succeeds: the organisation concludes that the tool is safe and effective, and scales it, without noticing that the conditions that made the pilot safe were features of the pilot, not of the tool.
What makes a pilot safe
A pilot is run by enthusiasts. The clinicians who volunteer for it are interested, attentive and forgiving of the tool's early failures. They read every output because the tool is new and they are curious. The data it is run on is often chosen, or at least cleaned, so that the tool has the best chance of showing what it can do. The project lead watches the results daily. Oversight is dense because a small number of people are paying close attention to a small number of cases.
None of that survives scale. At scale the tool is used by staff who did not choose it, on data nobody curated, under time pressure, with the project lead moved on to the next thing. The attention that made the pilot safe is gone, and nothing has been put in its place.
A pilot succeeds under conditions the programme will never have. The question is what replaces them.
What a programme has that a pilot does not
The transition from pilot to programme is a governance transition, not a technical one. A programme has what a pilot could do without:
- A stated intended purpose, fixed in writing, that limits what the tool is used for once the enthusiasts are no longer there to exercise judgement.
- A named accountable owner who will still be accountable in eighteen months.
- Training designed for staff who did not volunteer and may be sceptical, covering what the tool does not do as well as what it does.
- Oversight designed into the workflow rather than supplied by curiosity: who reviews what, when, and where it is recorded.
- Monitoring of the tool's performance over time and across populations, so that drift is seen by someone.
- An incident route that staff know about and use.
- A review cycle and an exit: a date on which someone asks whether the tool is still appropriate, and a plan for withdrawing it if not.
Each of these is a thing the pilot achieved informally through the attention of a few people. The programme has to achieve it formally, because attention does not scale.
Decide the criteria before the pilot starts
The most useful thing an organisation can do is to write down, before the pilot begins, two sets of criteria. First, what result would count as success, in terms the organisation cares about rather than the vendor's metrics. Second, what governance must exist before the tool is scaled, regardless of the result. The second list is the programme, described in advance. Without it, a successful pilot creates pressure to scale immediately, and the governance is left to be assembled afterwards, under that pressure, by people who now have a stake in the tool continuing.
The permanent pilot
There is a second failure, quieter than premature scaling. Some tools remain in pilot indefinitely, in use across the organisation but never formally adopted, because adoption would trigger governance obligations that nobody wants to own. The word pilot becomes a way of avoiding accountability. An inspector, an investigator or a coroner will not recognise the distinction. A tool in use is a tool in use, and the organisation is responsible for it whatever it is called internally.
The rule is simple to state. A pilot has a defined end date, a defined scope and a decision waiting at the end of it. Anything without those three things is not a pilot. It is a programme without governance.
This article sets out Novatib's advisory position. It is not legal or regulatory advice.