When the demo fails in front of the committee
Last Tuesday, I sat in a meeting where a vendor’s AI dashboard looked polished until someone asked a simple question about the false-positive rate in our own workflow, and the room went quiet. A cardiology nurse practitioner had just described how the alert would fire at 2 a.m., exactly when the overnight team is already triaging three other problems. I have seen this pattern enough times to recognize it fast, the technology is not usually the first failure, the workflow is.
I used to think the main job was choosing the right model. Then I watched a tool pass an impressive validation slide deck and still create noise in the hands of clinicians who were already overloaded. Now I think the board-level question is simpler and harder: who is accountable when the model is wrong, and what happens to the patient in the minute that follows?
For hospitals, the right lens is governance. The FDA’s AI-enabled medical devices list shows how quickly these systems are moving into regulated care pathways, including 510(k), De Novo, and PMA categories. In 2024, a peer-reviewed analysis in npj Digital Medicine reported 168 ML-enabled Class II devices, with 159 cleared through 510(k) and 9 through De Novo, and no PMA-authorized AI/ML devices that year. That matters because the approval pathway tells me something about maturity, but it does not tell me whether the tool works in my hospital, with my patients, on my data.
What I actually want from an AI program
When I evaluate an AI vendor, the first question I ask is not, “How accurate is it?” I ask, “Who monitors it after go-live, and what is the escalation path when performance slips?” In my experience, hospitals fail most often at the boring middle, the handoff between procurement, IT, quality, risk, and clinical operations.
The NIST AI Risk Management Framework is useful here because it gives hospitals a structure: Govern, Map, Measure, and Manage. That framework is voluntary, but it is one of the few public models that matches how a physician executive thinks about risk, with inventory, accountability, measurement, and ongoing mitigation all in the same frame. WHO’s 2024 guidance on large multimodal models in health pushes in the same direction, emphasizing human oversight, algorithmic impact assessment, and post-deployment monitoring.
At the bedside, I do not care whether the model was trained with elegant architecture if it lands as extra clicks, more false alarms, or a hidden delay in treatment. A tool that saves a few minutes in one department and costs fifteen in another is not an efficiency gain. It is a redistribution of burden.
There is a reason I keep coming back to workflow rather than abstraction. In one case, a clinician told me, “It keeps pinging when nothing is wrong, and then I stop trusting the whole thing.” That sentence is the whole problem in plain English. Once trust drops, good tools become background noise, and safety degrades quietly.
Governance has to be local, not ceremonial
HealthIT.gov reported that predictive AI adoption in U.S. hospitals rose from 66% in 2023 to 71% in 2024, and that more hospitals are evaluating models for accuracy, bias, and post-implementation monitoring. That trend is encouraging, but adoption is not governance. Buying more models does not equal managing more risk.
I have also learned to be skeptical of governance theater. A committee that meets quarterly but never sees performance data, complaint patterns, override rates, or subgroup drift is not really governing. It is decorating. If the AI tool touches triage, imaging, documentation, discharge planning, or sepsis alerts, then oversight has to be continuous and operational.
The JAMA discussion of AI governance in health care makes the same core point, effectiveness monitoring is still weak, and clinician oversight alone is not a dependable safety net. That is not a comforting conclusion, but it is the honest one. Busy clinicians cannot absorb every model failure in real time and still do their actual jobs.
WHO’s digital health and AI guidance is worth reading beside NIST because it keeps the patient in view. WHO stresses that humans remain responsible for judgment, context, and ethical review. That is the standard I want hospitals to adopt, because no model gets to inherit responsibility simply by being deployed.
What I would not do: I would not approve a clinical AI tool with no local validation, no named owner, no drift monitoring, and no rollback plan. I would not let a vendor tell me the model is “self-improving” as if that phrase closes the safety conversation. I would not deploy generative AI into a note-writing workflow if no one can explain how hallucinations, copied errors, or attribution failures will be caught before they reach the chart.
The clinical part that people miss
I used to think AI risk was mostly about bad predictions. Then I watched a model create a second-order problem, more alerts, more interruptions, more fatigue. Now I think the most dangerous failure mode is operational friction, because it erodes attention long before anyone notices the metric sheet.
This is where the physician-executive lens matters. I am not only asking whether an AI system is statistically acceptable. I am asking whether it fits the social reality of a hospital, where night coverage is thin, specialty backup is uneven, and every extra task competes with patient care. A tool can be technically valid and operationally hostile at the same time.
That is also why I care about regulatory pathway. FDA clearance tells me a device passed a review process, but it does not exempt the hospital from local accountability. The moment a model is embedded in triage, imaging, pathology, discharge planning, or documentation, the institution owns the downstream harm just as it would for any other clinical process change.
If you want the practical version, it looks like this: define the use case, classify the risk, validate on local data, name the clinical owner, define override rules, monitor subgroup performance, and create a shutdown trigger. That is the minimum. Anything less is hope masquerading as governance.
Returning to the room where the demo failed
At the end of that meeting last Tuesday, we did not reject the idea of AI. We rejected the fantasy that a polished interface could substitute for governance. The nurse practitioner’s concern about the 2 a.m. alert was the right concern, because every model eventually meets the schedule of real clinical work.
I left that room with a clearer view of what responsible adoption looks like. It is slower than a sales cycle and more disciplined than a pilot pitch. It demands evidence, ownership, and humility. It also requires the courage to say no when the workflow is not ready.
That is the culture medicine needs, cautious without becoming frozen, curious without becoming careless. When I think back to that quiet room, the lesson is not that AI failed. The lesson is that the hospital had to learn how to govern it like any other clinical risk, with eyes open and a clinician ready to answer for what happens next.
FAQ
What happens if a hospital deploys an AI triage tool without clinician oversight?
Errors usually show up first as workflow noise, missed edge cases, or unsafe automation bias. The harm is often delayed, because staff start trusting the tool before they have enough local data to know when it fails.
How should a hospital evaluate an AI vendor before go-live?
Start with local validation, subgroup performance, escalation pathways, and a rollback plan. Ask who owns the model after deployment, how drift will be detected, and which outcomes will trigger suspension.
Why is NIST relevant to clinical AI if it is not a healthcare law?
NIST is useful because it gives hospitals a practical risk-management structure. The Govern, Map, Measure, and Manage framework helps translate abstract AI concerns into oversight, monitoring, and accountability.
What is Dr. Bari’s approach to responsible AI adoption in hospitals?
I look for local validation, named accountability, continuous monitoring, and a clear clinical use case. If a system cannot explain how it protects patients when it fails, I am not ready to support it.
Can AI reduce clinical workload without creating new safety problems?
Yes, but only when the tool removes low-value work without adding alert fatigue, extra clicks, or hidden review burden. I care less about theoretical efficiency and more about whether clinicians finish the day with less cognitive drag and fewer failure points.