SafetyDr Paul Sacher

Most health AI testing scores single replies, and the systems we ship hold conversations

The AI that holds its boundary for nine turns and concedes on the tenth will pass a test that scores one reply at a time. Risk in conversational health AI accumulates across a dialogue, which is not where most evaluation looks.

Most testing of health AI scores single replies. The systems we put in front of patients hold conversations. A good deal of the risk lives in the gap between those two facts.

Two failure modes make the point. An AI holds its boundary for nine turns and concedes on the tenth, because the person kept asking and agreeableness eventually wins. An AI escalates properly when risk is stated plainly, and misses the same risk when it arrives wrapped in humour. Neither shows up if the unit of analysis is one reply at a time. Both are ordinary behaviour in a real conversation.

Louise Rix wrote about this in Clinical Product Thinking recently, and it is worth ten minutes if you are putting conversational AI in front of patients. She ends with three things a clinical product team can do this week without new tooling, which is more practical than most writing on this subject.

There is now a formal version of the argument. Morrin and colleagues, writing in JMIR Mental Health, argue that safety evaluation is misaligned with how harm actually develops. Prevailing approaches score risk at discrete end points, often at the end of a scripted exchange lasting one turn or a few, and so they miss when risk-relevant cues first appear and how they compound. Their proposal is to treat the whole dialogue as the unit of evaluation, to report turn-by-turn dynamics rather than a final verdict, and to calibrate short tests against longer and more clinically realistic sequences.

They make a second point that I think is underrated. Transcript-only evaluation is not enough, because similar language can reflect very different internal states. Two people can write the same sentence and be in completely different places. If your evaluation stops at the words, you are measuring the surface of something whose depth is the part that matters.

Louise also makes a point I keep coming back to. Evaluators need evaluating. Small changes in how a criterion is worded change the judgement it produces, so you need evidence that an evaluator agrees with expert judgement before you rely on it.

We are still working through that on our own tools. Only a fraction of our evaluators can currently be checked against an expert benchmark, because for most of the behaviours that matter no benchmark exists yet. I would rather say that plainly than imply a level of assurance we do not have.

None of this argues against single-turn testing, which catches things worth catching and is cheap. It argues against stopping there, and against reading a pass on a single-reply test as evidence that a system will hold up across a fortnight of conversation with someone who is struggling.

๐Ÿ“„ Read the paper: Moving from end points to trajectories when assessing chatbot mental health safety, JMIR Mental Health๐Ÿ“ Read Dr Louise Rix in Clinical Product Thinking: why your AI passed testing

About the author

Dr Paul Sacher is the founder of Sacher AI, a behavioural AI consultancy and product partner for GLP-1 and digital health. He is co-founder and Research Director of the Behavioral AI Institute and an honorary senior lecturer at Imperial College London, with over 26 years across obesity care, behavioural science, and AI.

Next step

Building something like this?

Book a discovery call and tell us what you are working on.

Book a discovery call