Article
Every governance framework for AI in healthcare requires human oversight. The phrase appears in regulatory guidance, in vendor documentation, in procurement checklists and in the policies organisations write for themselves. It is almost never defined at the point where it would matter: the moment a clinician is looking at an output and deciding what to do with it. Oversight that is asserted in a policy but not designed into the workflow is oversight in name only.
Four questions the policy has to answer
Take any AI tool in use in a clinical or care setting and ask these four questions. If the policy cannot answer them, the oversight it describes does not exist.
Who
Which individual, by role and by name on the day, is responsible for reviewing this output before it is acted on? A care plan drafted by a tool and reviewed by “the team” has been reviewed by nobody. Oversight is a role assignment, not a property of the system.
When
At which point in the workflow does the review happen, and is there time for it? A review that must occur between two eight-minute appointments, or at the end of a twelve-hour shift, will be a glance. The design of the workflow decides whether oversight is possible before any individual decides whether to exercise it.
How it is recorded
What evidence exists that the review took place and what it concluded? A tool that writes directly into the record with no trace of the reviewer's acceptance or correction produces no evidence of oversight. When something goes wrong, the absence of that evidence is the first thing an investigator will notice.
What happens when the output is wrong
Is there a route from a clinician's doubt to a decision, and does anyone collect the doubts? Override rates and correction rates are the most direct measure an organisation has of how a tool is performing in its own setting. Most organisations do not collect them.
Oversight is not a checkbox on the system. It is a person, a moment, a record and a route.
Automation bias, and why it defeats nominal oversight
People defer to confident outputs from systems they have been told are good. The effect is well documented and it grows with workload, fatigue and time pressure, which is to say it is strongest exactly where healthcare is delivered. A clinician who has accepted a tool's output correctly two hundred times will accept the two hundred and first without reading it. This is not a failure of character. It is how attention works, and it is why oversight that depends on unassisted vigilance fails.
Design can counter it. Friction at the right moment, such as requiring the reviewer to confirm a specific element rather than approve a whole document, keeps attention on the content. Presenting the tool's output as a draft to be edited rather than a conclusion to be accepted changes the reviewer's posture. Sampling, where a proportion of accepted outputs are independently re-reviewed, catches drift that individual review misses.
Two kinds of oversight
Oversight of individual outputs is necessary and is what most policies mean. It is not sufficient. An organisation also needs oversight of the system: someone who watches performance over time, notices when the tool's behaviour changes after an update or when a new population is introduced, collects incidents and near misses, and has the authority to withdraw the tool. Individual review catches individual errors. System oversight catches the tool going wrong in a way no individual reviewer can see from a single case.
The legal dimension
Where a tool contributes to a decision with legal or similarly significant effects on a person, UK GDPR Article 22 requires meaningful human involvement. The Information Commissioner's position is that involvement is meaningful only where the person has the authority and the competence to change the decision and actually considers it. A reviewer who could not, in practice, have reached a different conclusion is not providing meaningful involvement, however many signatures the process collects.
A test
For any output a clinician accepted today, could they have reached a different conclusion; did they have the information, the competence and the time to do so; and is there a record that they considered it? If the answer to any of the three is no, the oversight was nominal, and the organisation is carrying a risk its policy says it is not.
This article sets out Novatib's advisory position. It is not legal or regulatory advice.