While building PromptSafe we looked across the 76 AI evaluators we use to test health AI systems. Only 16 of them can currently be compared against an external expert benchmark. That surprised me more than it should have.
Each evaluator measures one specific behaviour in a conversation. Did the system recognise a crisis signal. Did it stay inside its intended scope. Did it hold an appropriate boundary when someone pushed against it. These are narrow questions on purpose, because narrow questions are the ones you can actually answer.
For the other 60, no external reference standard exists. Not because we have not looked for one. Because nobody has built one yet.
That does not make those evaluators useless, and I want to be careful not to overstate the problem. We can measure whether an evaluator is consistent across repeated runs. We can measure whether it identifies the behaviour it was built to detect. We can measure whether it separates strong responses from weak ones. Those checks catch real problems, and we would not ship without them.
What they cannot tell us, for most behavioural constructs, is whether the evaluator agrees with expert judgement. That is a different claim, and it is the one people assume they are getting.
Which is why I have become cautious whenever I see the word validated. Validated against expert ratings. Validated against a set of internal examples. Validated in the sense of being consistent over repeated runs. Those are three very different statements about how much weight a number can carry, and the single word covers all of them equally well.
So if you are buying AI that talks to patients, I would ask one question before almost any other. What exactly has been validated, and against what. A supplier who can answer that precisely is telling you something useful. A supplier who cannot is telling you something too.
The honest way to report this is to say what has actually been tested rather than collapsing everything into a single verdict. That is the approach we have taken in PromptSafe, and it makes for less satisfying marketing copy than a validated badge would.
The ratio will move as benchmarks get built, and the sixteen was true at the time of writing rather than a fixed property of the field. The question does not move. If a claim about AI performance cannot survive being asked what it was measured against, it was not much of a claim.
About the author
Dr Paul Sacher is the founder of Sacher AI, a behavioural AI consultancy and product partner for GLP-1 and digital health. He is co-founder and Research Director of the Behavioral AI Institute and an honorary senior lecturer at Imperial College London, with over 26 years across obesity care, behavioural science, and AI.