This is Clinical Product Thinking 🧠, a weekly newsletter featuring practical tips, frameworks and strategies from the front line of clinical product.
Welcome, friends, this is issue No. 052 of Clinical Product Thinking 🎉. Today we’re talking about one of my favourite topics, clinical safety in AI systems!
There’s a trend I’m seeing a lot of when it comes to clinical safety, and that is human-in-the-loop as the (perceived) ultimate control.
“Of course it’s safe. A human reviewed it!” 🤦♀️
I hate to burst this bubble but human-in-the-loop is not, in itself, evidence that risk has been controlled. A human review alone does not tell you whether that human is actually preventing harm.
For instance, a human can be technically “in the loop” while having almost no meaningful ability to spot, understand or correct AI failure.
A recent review in Nature highlighted this along with two particularly relevant risks: automation bias and excessive delegation.
Both are increasingly important to understand and monitor as AI becomes embedded into routine clinical practice. Let’s dive in!
Human review: The least tested safety control
Many typical clinical workflows go like this: AI reviews and generates a recommendation, clinician reviews the AI’s output, clinician approves or edits the recommendation.
The safety case might say something like: “All AI outputs are reviewed by a qualified clinician.”
This leaves many unanswered questions:
How obvious would an error be?
How much time are they expected to complete the review in?
What information do they have available?
What happens when they disagree with the AI?
When is the AI recommendation shown?
AI changes how humans make decisions
An AI recommendation is not a neutral action on humans. It’s changes the decisions they ultimately make and this has been shown in research.
People become more likely to accept a recommendation from an automated recommendation, particularly when the system is usually right. If a major error only happens 1 in 100 times it’s unlikely a human will be as vigilant when they’re reviewing the 99th case. This is called automation bias.
Factors that make automation bias more likely:
The output is written confidently
The clinician is busy
The decision feels ‘routine’
Checking the recommendation properly is cumbersome
The counterintuitive point here is as the AI becomes more reliable, clinicians may have fewer reasons to challenge it. That can make rare errors harder to catch because vigilance falls over time.
Humans gradually leave the loop
The Nature paper describes the concept of excessive delegation - that is humans shifting from active decision-makers towards passive overseers.
Take for instance a tool that uses AI to generate clinical letters. At launch the clinicians will likely scrutinise every sentence. Six months later that behaviour may have mutated into a cursory glance.
Nothing in the workflow has changed and there’s still technically a “human in the loop”. The nature of the oversight has changes substantially though.
This makes human oversight something you need to monitor post deployment rather than assume remains an effective control over time.
Automation bias changes how much weight a clinician gives the AI’s judgement. Excessive delegation changes how much of the judgement the clinician performs themselves.
Some AI errors are inherently hard for humans to catch
The thing about AI errors is they’re not always glaring hallucinations. An output can look totally reliable but contain one small clinically consequential mistake. For example:
An allergy omitted from a summary letter
An incorrect dose
An inappropriate recommendation buried within an otherwise plausible output
So, what to do about it?
Design your clinical AI product with error detectability in mind, not just reviewability. For example:
Highlight the parts that carry the greatest clinical risk, rather than asking clinicians to review everything with equal attention.
Make high-risk facts easy to check against source data e.g. allergies, doses, abnormal results.
Show the source where possible rather than showing only the AI interpretation.
For the highest risk decisions, consider requiring explicit confirmation (although this needs to be traded off against user experience).
If human review is your control, test the control
At the risk of stating the obvious… if a human review is an important control, you need to test that control. What evidence do you currently have that the review works?
For instance, during testing you could deliberately introduce clinically significant AI errors and measure:
Error detection rate
Which error types are missed
How long an effective review takes
How performance changes under workload pressure
If detection rate changes with clinician type, speciality, or other factor
Post-deployment, you could also monitor for evidence that review behaviour is changing over time e.g.
Decreasing time taken to review
Lower manual edit, override or challenge rates
Convergence between AI and final clinician decisions
Fewer requests for additional information or review of source material
The takeaway
If you’re using “human review” as a mitigation, I suggest you think about these 5 things before launch:
1. Detectability: Could the reviewer realistically identify the failure?
2. Information: Do they have everything required to check it?
3. Capacity: Do they have enough time and attention to do so?
4. Independence: Can they form their own judgement rather than simply accepting the AI’s?
5. Evidence: Have you demonstrated that the review actually prevents the harm you are relying on it to prevent?
We need to start thinking of ‘human-in-the-loop’ as a shorthand for workflow design. Not as evidence of a clinically safe system.
If clinician review is one of your risk controls, you need to evaluate that clinician-AI interaction in the same way you would evaluate the model itself. That’s how we’re going to build enduring, safe clinical AI.
How to Build Clinical AI Patients Actually Trust 🤖
Human oversight is only one piece of building clinical AI that people can trust. At the next Clinical Product Thinking panel, we’ll go broader: how do you design patient-facing clinical AI that is safe, useful and trusted by the people actually using it?
Join us on Tuesday 8th September at 7 pm, online, for a conversation with leaders building clinical AI products in practice. 👉 Sign up here.
Clinical Product Drinks 🍸
Join the next clinical product drinks! A chance to meet other clinical product leaders and managers and share the good, the bad and the ugly. 👉 Sign up here.
That’s all for this week. See you next time! 👋
🤝 Work with me | 📅 Attend an event | ✍️ Send a message
Written by Dr Louise Rix, Head of Clinical Product, doctor and ex-VC. Passionate about all things healthcare, healthtech and clinical product (…obviously). Based in London. You can find me on LinkedIn.
Made with 💜 for better, safer HealthTech.




