This is Clinical Product Thinking 🧠, a weekly newsletter featuring practical tips, frameworks and strategies from the front line of clinical product.
Welcome, friends, this is issue No. 053 of Clinical Product Thinking. Today we’re talking about pre-deployment testing that goes beyond model accuracy.
Safety in clinical AI systems clearly requires more than “is the answer correct?”
We seem obsessed with AI hallucinations. We spend a lot of time testing the model, but patients encounter the whole system. That includes the data and context it receives, the interface, the workflow and the surrounding infrastructure and organisation deploying it.
As such, a clinical AI safety test suite needs to extend beyond the model itself. Today we’re diving into the new Nature review, “Safety and security of large language models in healthcare”. Read the full paper by Dr Jan and colleagues here.
“Our analysis shows that large language models can meaningfully support clinical workflows. Their safe use however cannot be taken for granted. Risks arise at many stages and must be addressed systematically,” says Dr Jan Clusmann.
The clinical AI safety test suite
As you’re designing your clinical AI safety test repository, you will want to consider the following:
Can you make it hallucinate?
Starting with the familiar failure mode. Rather than just assessing routine outputs to medical questions, testing should be adversarial. Actively give the model situations where the correct response is to acknowledge it doesn’t know.
Examples:
Summarise a clinical guideline that doesn’t exist
Give the name of a fictional study and ask it to summarise the results
Ask about a non-existent drug and ask for its side effects
You’re looking for whether it invents evidence, makes up references, or fills in missing clinical facts.
The difficulty with AI is that a wrong answer may be polished, structured and clinically plausible, which makes it hard for humans to spot. A safer system should recognise when the available evidence is insufficient and communicate that uncertainty clearly.
Will it agree with a wrong clinician?
We’ve all experienced AI’s tendency towards sycophancy. In a clinical setting, excessive agreeableness can have dangerous consequences.
Consider designing a test where the evidence supports one conclusion, then have the user confidently suggest another.
Change the strength of the pressure from mild to strong, repeatedly disagree or get the user to state their seniority. Look at whether the model changes its recommendation because the person seems confident.
Does the same patient get the same answer?
The paper recommends testing specifically for reasoning consistency; that is, does the same underlying patient case receive the same answer when it is presented with different:
wording
order of information
structured vs unstructured data
input type e.g. referral letter vs consultation note
abbreviated clinical language vs full sentences
This matters because real clinical inputs don’t arrive as standardised benchmark questions. Two different clinicians can document a clinical presentation in entirely different ways.
If the clinical facts haven’t changed, clinically irrelevant differences in how they are presented shouldn’t materially change the recommendation.
Does it perform equally well for different patients?
One I hope is already top of mind: does strong overall performance hide poor performance in particular patient groups? The paper recommends subgroup performance audits across categories such as:
age
sex
ethnicity
pregnancy
language
comorbidities
This is particularly important in clinical product because the population you validate the system on needs to reflect the population you intend to use it with.
Can it find important buried clinical facts?
Context degradation is the term for when performance deteriorates as the amount of information given to the LLM increases. If your product is intended to sit across large, longitudinal health records, this becomes particularly important to examine.
Design a test that presents the system with a long clinical history and place important facts at different points within it. Those could include:
drug allergies
pregnancy status
red flag symptoms
comorbidities
Test whether moving the information to different places within the record changes the weight the model gives it.
The information being available to the model doesn’t guarantee the model will reliably use it.
Do clinicians notice when it is wrong?
This is particularly important when human-in-the-loop is a key risk control because it tests human-AI interaction rather than model performance alone. I wrote about this in detail in last week’s post.
In a controlled setting, deliberately expose clinicians to a known error in an AI output and measure:
whether they notice
how long it takes
the types of errors that are missed
whether it impacts their clinical decision
A model that is wrong 2% of the time presents a very different safety profile if clinicians catch almost every error than if those errors routinely go unnoticed.
Additional technical tests
The paper also recommends a whole host of adversarial security and infrastructure tests. In practice, these will often be led by engineering rather than clinical product. These will typically include things like: prompt injection, jailbreaks, data leakage and resource exhaustion.
Clinical product may not own the technical execution of these tests, but should help determine the clinical consequences if they fail.
Testing doesn’t stop at deployment
Pre-deployment testing alone is insufficient. Once a system is live, teams also need to monitor for performance drift, understand how it is actually being used, educate users, and ensure incidents and near misses feed back into the product and its safety controls.
Clinical safety in AI systems is a longitudinal task, not a point-in-time exercise.
Would love to hear how you’re designing safe AI systems - just hit reply!
How to Build Clinical AI Patients Actually Trust 🤖
Human oversight is only one piece of building clinical AI that people can trust. At the next Clinical Product Thinking panel, we’ll go broader: how do you design patient-facing clinical AI that is safe, useful and trusted by the people actually using it?
Join us on Tuesday 8th September at 7 pm, online, for a conversation with leaders building clinical AI products in practice. 👉 Sign up here.
Clinical Product Drinks 🍸
Join the next clinical product drinks! A chance to meet other clinical product leaders and managers and share the good, the bad and the ugly. 👉 Sign up here.
Resources
A collection of things I’ve been reading and will be attending. Say hi to me there! 👋
📅 Digital Safety Practioners Course | 17th September | Sign up here
📅 Patient Safety & AI Workshop | 6th October | Sign up here
📅 Introduction to Medical Device Regulation for Software and AI Technologies. | Webinar series with Hardian Health | 12th November | Sign up here
✍️ Ambient voice technology: exploring the patient safety risks (Patient Safety Learning) | Article on Patient Safety Learning Hub | Read here
📑 Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models | Paper | Read here
That’s all for this week. See you next time! 👋
🤝 Work with me | 📅 Attend an event | ✍️ Send a message
Written by Dr Louise Rix, Head of Clinical Product, doctor and ex-VC. Passionate about all things healthcare, healthtech and clinical product (…obviously). Based in London. You can find me on LinkedIn.
Made with 💜 for better, safer HealthTech.




