Last Tuesday, I watched a perfectly ordinary consult go sideways
Last Tuesday afternoon, I was reviewing a specialist message thread between a hospitalist and a consultant when the pattern felt uncomfortably familiar. The patient was stable enough to avoid the ICU, unstable enough to keep everyone alert, and the note was full of the kind of details that get compressed under time pressure: a subtle trend in vital signs, one lab that had drifted, a medication list that was more dangerous than it looked. Someone in the room said, “We have the data, but we’re still not sure what we’re missing.” I knew exactly what they meant.
The biggest weakness in health AI benchmarking is omission blindness. Recent studies, including NOHARM and the Nature Medicine ChatGPT Health triage paper, show that the most dangerous failures are often the recommendations the model never makes, not the statements it gets wrong.
That matters because medicine is a workflow of incomplete handoffs, not multiple-choice questions. If evaluation only measures what an AI says, it misses the clinical question that actually decides whether a patient gets the right next step.
I used to think the central risk in health AI was hallucination. Then I spent more time looking at actual clinical workflows, especially consult requests, and my view changed. The problem that keeps surfacing is quieter and more expensive: omission. A model can sound reasonable, cite the right differential, and still fail to mention the one escalation, imaging step, or medication review that should have been on the page. That is a different kind of failure, and it is harder to catch.
If you want my professional baseline, I would frame it this way: the useful question is not whether a model can answer a vignette, but whether it can help a clinician avoid a missed action in a messy, real-world exchange. That is why I care less about leaderboard theater and more about evaluation design, governance, and the human workflow around the tool. My perspective as a physician-executive is shaped by the same thing that shapes Dr. Sina Bari’s clinical and leadership background: in medicine, the failure mode is often downstream of the interface.
Why NOHARM changes the conversation
The February 2026 Nature Medicine study, ChatGPT Health performance in a structured test of triage recommendations, reported that ChatGPT Health undertriaged 52% of true emergency scenarios. That number matters, but the more revealing point is how the test was constructed. The companion NOHARM work used real electronic physician-to-specialist consultation requests rather than tidy NEJM-style vignettes. That design choice seems boring until you realize what it buys you: clinicians do not practice in polished question sets. They practice in incomplete messages, implicit assumptions, and domain-specific shorthand.
NOHARM, which required 29 board-certified physicians and 12,747 annotations across 4,249 candidate clinical actions, found that more than 80% of severe errors were omissions rather than false statements. I consider that the most important result in the paper. It means our usual evaluation tools, multiple-choice exams, citation quality checks, and hallucination detectors, are measuring the easiest thing to measure. They are grading the model on what it said, not on what it failed to say.
That is a governance problem as much as a modeling problem. A hospital can buy a system, validate it on canned examples, and still leave the core safety question untouched: does the model reliably surface the next action that changes outcome? If the answer is no, then a high benchmark score can become a false reassurance mechanism.
For readers who want the original triage paper, the quantitative claim is straightforward. The best-performing systems reduced potential severe harm to 2.9% in the most favorable comparisons, while the study also found potential severe harm in up to 24.6% of cases overall. Those numbers are a warning and a hint. They show that performance varies meaningfully across systems, but also that the variance is real enough to matter clinically. A model that looks acceptable on one benchmark can still become a poor colleague in a real workflow.
Benchmarks that reward the wrong thing
I have sat through enough vendor demos to know how seductive a tidy benchmark slide can be. A model scores well, the chart is clean, the notes are polished, and everyone in the room relaxes just a little. I do not relax. In a hospital setting, the important question is whether the model can function inside regulatory, operational, and human constraints, including FDA pathways, local governance, and the very ordinary mess of handoffs. A beautiful score on a synthetic set tells me almost nothing about whether the tool is safe in a triage queue, a radiology worklist, or a busy hospitalist service.
The new line of work on benchmark fragility supports that skepticism. The medRxiv paper Aggregate benchmark scores obscure patient safety implications of errors across frontier language models argues that aggregate scores hide clinically meaningful error patterns. That is exactly what I see in practice. A model can improve on average while still failing the exact subgroup, presentation, or uncertainty pattern that matters most. If I only look at the average, I may miss the patient who needed the exception handling.
There is a similar lesson in the Communications Medicine study on care-seeking advice, Evaluating the accuracy of ChatGPT model versions for giving care-seeking advice, which compared model versions and found that performance varied enough to matter for patient-facing recommendations. Versioning is not a footnote. In healthcare, it is a safety variable. I would not allow a hospital to treat a model upgrade like a routine software cosmetic patch when the downstream behavior can change triage advice, referral timing, or reassurance thresholds.
And I would not deploy a system that only passes a generalized benchmark if I cannot trace its failure modes back to specific clinical actions. That is my line. No amount of elegant language compensates for a missed escalation, a missed medication warning, or a missed follow-up recommendation when the patient is relying on the tool to help decide what happens next.
Human-plus-AI works only when the human actually uses the answer
The most nuanced finding in the NOHARM report is also the hardest to summarize in a headline. AI systems alone outperformed generalist physicians alone, and they also outperformed many physicians working with AI. But when researchers examined the transcripts, they found that physicians were often shown correct recommendations and simply left them out of the plan. Had those recommendations been integrated, the combined human and AI response would have beaten both the physician and the model working separately.
That result stings, because it reveals a systems problem hiding inside a collaboration story. Human and machine were making different mistakes. The AI could identify a correct next step; the human could recognize context, but not always incorporate the model’s suggestion into the final decision. In my experience, this happens when the tool is bolted onto the workflow instead of designed into it. The interface encourages glance-and-ignore behavior. The team sees the answer, then reverts to habit.
The radiology literature is moving in the same direction. The 2026 article Human-AI collaboration in radiology: conceptual frameworks for responsible implementation emphasizes that collaborative design, not raw model output, determines whether AI improves care. That tracks with what I have seen. A second reader who is technically present but socially easy to ignore is not the same as a workflow that forces review, acknowledgement, and escalation when needed.
There is also a clinical education warning here. The Nature Medicine paper on AI-induced never-skilling suggests that overreliance can erode independent competence over time. I take that seriously. If we train clinicians to depend on AI for the first pass and then fail to inspect the output carefully, we create a generation of users who can recognize fluency but not necessarily error. That is a bad trade in a field where competence degrades quietly.
What I think should happen next
I used to think better models would solve most of this. Now I think better evaluation and better workflow design matter just as much. The Frontiers in Medicine paper on agentic AI, Exploring Agentic AI in Healthcare: A Study on Its Working Mechanism, and the npj Artificial Intelligence review AI agent applications, evaluations, and future directions in healthcare both point toward the same conclusion: autonomy without oversight is a liability, and oversight without usable actionability is theater.
So what would I do? I would require omission-aware evaluation for any high-stakes clinical AI. I would insist on case sets built from real clinical artifacts, not just neat teaching vignettes. I would ask whether the benchmark captures missed actions, not only wrong answers. I would require version-specific monitoring, local calibration, and explicit human-on-the-loop accountability. And I would not accept a vendor claim that says the model “helps clinicians think” unless the deployment includes a way to prove that the help shows up in the actual plan.
This is where boards and medical leaders need to stop talking about AI as if it were a single product category. In practice, it behaves more like a chain of decisions: data provenance, model behavior, workflow insertion, human adoption, and governance. Break any link and the whole chain weakens. The regulatory frame matters too. FDA pathway questions, NIST-style risk management, and institutional oversight are not administrative chores. They are how you keep a good demo from becoming a bad clinical habit.
There is a deeper point here. If the answer we need is the one the model never says, then the next generation of health AI evaluation has to become more like clinical QA than software scoring. That means more expert annotation, more context, more longitudinal review, and more willingness to ask what a model missed. Expensive work, yes. Necessary work, absolutely.
Back in the consult thread
By the end of that Tuesday consult, the team had the right plan, but only after someone slowed down long enough to notice what had been left out. That is where I keep returning in my mind. The most dangerous AI failure in medicine may not be a dramatic hallucination. It may be the quiet omission that slides past a busy clinician because the answer sounded good enough.
That is why I read NOHARM as more than a benchmark paper. It is a reminder that healthcare AI is not being evaluated against a quiz. It is being evaluated against the cost of a missed action in a real patient workflow. If we do not build for omissions, we will keep rewarding eloquence and punishing the wrong kind of imperfection.
And if the next consult thread looks familiar, I want the system to do more than speak fluently. I want it to help the team catch the thing nobody said out loud.
FAQ
Why do health AI benchmarks miss the most dangerous failures?
Because most benchmarks measure what the model said, not what it failed to say. In NOHARM, more than 80% of severe errors were omissions, which means the model can appear competent while leaving out a critical next step. That is a safety problem, not just a scoring problem.
What does it mean when ChatGPT Health undertriages 52% of true emergencies?
It means the model recommended a less urgent response than the case actually required in more than half of true emergency scenarios in that structured test. In a clinical setting, that can delay escalation, imaging, specialist review, or transfer. A system with that behavior should not be used as a standalone triage authority.
Why are real consult requests better than clinical vignettes for evaluating AI?
Real consult requests contain the messiness of medicine, incomplete context, shorthand, and unspoken assumptions. NOHARM used actual physician-to-specialist consultation requests and 4,249 candidate clinical actions, which is much closer to how clinicians work. That makes omission detection more realistic than a polished multiple-choice case.
What would Dr. Sina Bari look for before trusting a hospital AI tool?
He would look for omission-aware testing, version-specific monitoring, and evidence that the tool fits the workflow instead of disrupting it. He would also want to know how the system behaves on real consult language, not just curated examples. If the vendor cannot show those data, the deployment is premature.
Can AI and physicians perform better together than alone in health AI workflows?
Yes, but only if the physician actually incorporates the model’s correct recommendation into the plan. NOHARM found that AI and physicians often made different mistakes, and the combined response would have outperformed either alone if the AI’s recommendations had been used. The workflow has to support collaboration, not just coexistence.