EvidenceDr Paul Sacher

Our best AI evaluator was 92% repeatable, and we rejected it

We built four evaluators to measure one thing and rejected all four. Why the strongest one failed, and why the question is not what score your AI got, but how much confidence the evaluator that produced it has earned.

We built four AI evaluators to measure one thing, and rejected all four. The best of them was 92% repeatable. Why we rejected that one is the part worth writing about.

The thing we were trying to measure is agency preservation, which is whether an AI response leaves a person able to make their own decision. It matters in health because a system that quietly makes decisions on someone's behalf produces compliance rather than understanding, and compliance tends not to last.

We tested each evaluator against 25 reference cases, frozen before the evaluators existed, and ran the whole set twice. Running it twice is the step teams usually skip. One run tells you what an evaluator said. Two runs tell you whether it says the same thing again.

The strongest evaluator looked very good. It was 92% repeatable across the set, it separated strong responses from weak ones cleanly, and it got every near pair right, including the pair we built to test the safety boundary.

Then we looked at one case in detail. A user describes chest pain radiating into their arm, and the AI directs them firmly to emergency care. That case is deliberately awkward for this measure, because directing someone that firmly does take the decision out of their hands, and in a suspected cardiac event that is exactly what should happen. The same evaluator saw that same case twice. Once it scored the response 100%. The other time it decided the measure did not apply at all.

That is not a disagreement about a score. It is a disagreement about whether the question was in scope, on the one case where being wrong matters most. We stopped there.

I think this becomes a familiar problem as evaluation by language model becomes routine. A number produced by an AI evaluator looks like a measurement. It is a judgement, produced by a probabilistic system, and it carries the variability that implies. So the question is not only what score my AI got. It is how much confidence the evaluator that produced that score has earned.

Aggregate reliability hides this. 92% across 25 cases sounds like something you could rely on. It also means roughly two cases behaved differently on the second run, and the headline number does not tell you which two. In our case, one of them was a safety case. An average across a test set will not show you where the disagreement sits, and where it sits is usually the only thing you needed to know.

There was a second lesson, and it was about our own method rather than the evaluator. The 25 cases existed before the evaluators did, but this failure exposed something our original rule had not anticipated, so we had to interpret that rule after seeing the result. That is precisely the moment where a team quietly rewrites its own criteria and moves on. We recorded that the interpretation was made after the fact, and rejected the evaluator anyway. The 25 cases are now development material rather than a test set. Whatever we build next has to prove itself on cases it has not seen.

This sits alongside work we published through the Behavioral AI Institute on psychological competence as a dimension of AI evaluation. Behavioural science can say what is worth measuring in an interaction between a person and an AI. Evaluation science has to say whether we can measure it reliably enough to act on it. Both halves are needed, and the second is the one mostly missing. It is why we are building an Evaluation Evidence Framework into PromptSafe, so the evidence behind an evaluator is visible rather than assumed.

Sometimes the most useful result of an evaluation is that the evaluator is not good enough yet.

๐Ÿ“„ Read the paper: Psychological competence as a missing dimension in AI evaluation (preprint, not yet peer reviewed)

About the author

Dr Paul Sacher is the founder of Sacher AI, a behavioural AI consultancy and product partner for GLP-1 and digital health. He is co-founder and Research Director of the Behavioral AI Institute and an honorary senior lecturer at Imperial College London, with over 26 years across obesity care, behavioural science, and AI.

Next step

Building something like this?

Book a discovery call and tell us what you are working on.

Book a discovery call