This is Clinical Product Thinking đ§ , a weekly newsletter featuring practical tips, frameworks and strategies from the front line of clinical product.
Welcome, friends, this is issue No. 056 of Clinical Product Thinking. Today weâre talking about using AI for clinical assurance.
Last week I went to the Digital Safety in Practice event by NHSE, a day focused on recent developments in the world of clinical safety. Iâd highly recommend it to anyone working in the space!
One fascinating talk was from Prof Mark Sujan, Chair in Safety Science at the University of Yorkâs Centre for Assuring Autonomy. His research focuses on safety assurance for systems involving AI.
Mark presented work on using LLMs to review patient-safety learning responses. What interested me about it most was how LLMs differ from human reviewers and what that means when weâre applying AI to clinical assurance.
Here are the lessons Iâll be taking away in my own work:
Define the task narrowly
The lesson from the research was that âreview this safety caseâ was too vague to give meaningful results. Instead, they had multiple agents (8, I believe!), each with a different role or task, e.g:
identify claims without supporting evidence
find missing hazards
extract controls
find gaps in the argument
As Mark put it, youâre not really assuring the LLM. Youâre assuring it for use on a specific task.
Give it a rubric
During the research, they found that loose criteria gave poor results. This improved significantly when they explicitly defined what evidence to look for.
Rather than: âIs this hazard adequately controlled?â
Try: âIdentify each claimed control. Show the evidence that it exists, assumptions being made, gaps in implementation evidence and remaining risk.â
Making the AI show itâs working means you have the ability to critique.
Use AI to find evidence, but be cautious when it interprets it
This was an interesting finding from the research. LLMs are very good at finding evidence but were markedly less reliable at deciding whether that evidence genuinely demonstrated what was being assessed. The AI was also less reliable at systems-level reasoning.
Based on this, a useful division of labour might be:
AI: retrieve, collate, compare, categorise and surface evidence
Human: interpret, contextualise, and apply judgement
Watch for proxy evidence
AI can find something that it thinks looks like evidence without actually proving the underlying claim.
Mark gave an example of a family being treated compassionately at the time of an incident occurring. The model mistakenly attributed that the incident itself had been compassionately investigated. The two separate events had been flattened into one.
You can imagine how this can happen in a safety case:
A control is listed â evidence it exists
Testing is mentioned â therefore adequate testing has happened
The words âclinical reviewâ are present â meaningful clinical review must have occurred.
Make sure you ask: has the AI found evidence of the thing itself or something that only resembles it?
Donât just check whether you agree with the answer
The research showed a significant rate of false agreement. That is when a human and an LLM give the same overall assessment based on totally different reasoning.
Unfortunately, this increases the human review burden. Even when you agree with an AIâs answer. You still need to check the thread of reasoning it used to get there.
AI in the human loop
One person in the audience offered this phrase:
âAI in the human loopâ, rather than âhuman in the AI loopâ.
Thatâs where most of us are at with AI assurance today. Keep the clinical assurance process human-led, then use AI where it is particularly useful.
The strongest use cases here arenât âtell me whether this is safe.â
Theyâre things like:
Find every unsupported claim.
Show me where evidence is missing.
Compare this against a defined rubric.
Surface inconsistencies for a human to investigate.
Extract the evidence relevant to this specific assurance question.
For now, I like the idea of AI in the human loop: let AI do more of the searching, sorting, comparing and challenging, while keeping interpretation, judgement and accountability with the human reviewer.
If youâre using AI in clinical safety or assurance, Iâd love to hear what youâre doing with it. Just hit reply to this email.
Clinical Product Thinking Events đ
I would love to hear what kind of events youâd love to see next on the CPT calendar. Let me know in the poll below:
Clinical Product Drinks - Tuesday!đ¸
Join the next clinical product drinks! A chance to meet other clinical product leaders and managers and share the good, the bad and the ugly. đ Sign up here.
Thatâs all for this week. See you next time! đ
đ¤ Work with me | đ Attend an event | âď¸ Send a message
Written by Dr Louise Rix, Clinical Product, AI & Safety, doctor and ex-VC. Passionate about all things healthcare, healthtech and clinical product (âŚobviously). Based in London. You can find me on LinkedIn.
Made with đ for better, safer HealthTech.



