Analysis / 001

The AI Question Hospitals Keep Asking Too Late

Last Tuesday, a clinician showed me an AI demo that looked polished until we asked a simple question: what happens when the model is wrong on a crowded ward at 2 a.m.? Hospitals do not need more AI hype, they need governance that treats deployment like a clinical intervention.

Author

Dr. Sina Bari, MD

Physician-Technologist | Healthcare AI Executive | Stanford Medicine

Published

September 1, 2026

Reviewed

September 1, 2026

Last Tuesday, I sat in a conference room with a hospital team while a vendor walked us through an AI tool that promised faster notes, cleaner workflows, and fewer late nights. The demo was smooth. The room was quiet. Then one of the nurses asked the question that matters most in medicine: “Who catches it when this thing is wrong at 2 a.m.?”

Hospitals should treat AI as a clinical governance problem, not a software purchase. The right question is not whether the model looks impressive in a demo, but whether it has a regulatory path, a measurable safety case, human oversight, and a plan for drift after deployment.

In practice, that means using FDA pathways, NIST risk management, and WHO ethics as the baseline, then demanding local validation before AI touches patient care.

I have seen this pattern enough times to be skeptical on purpose. A polished interface can hide weak evidence, unclear liability, and workflow damage that only appears once the tool meets real patients, real fatigue, and real interruptions.

Hospitals keep buying the promise before buying the proof

The most common mistake I see is simple. Leaders evaluate AI like procurement, then discover too late that deployment behaves more like introducing a new medication or a new device into the hospital ecosystem. The FDA makes this distinction explicit in its guidance on artificial intelligence in software as a medical device, where AI/ML tools are handled through regulated medical-device pathways such as 510(k), De Novo, and PMA depending on the product and risk profile.

I used to think the main problem was overenthusiastic vendors. Then I spent more time inside operational AI reviews and realized the bigger problem was internal: hospitals often do not know what evidence threshold they want before the contract is signed. Now I think the first safeguard is not technical at all. It is governance.

My clinical test is blunt

When I evaluate an AI tool, I ask a few questions in plain language. What is the failure mode? Who reviews the output? What happens when the model sees a patient population different from the one used to train it? How fast will performance degrade if the documentation pattern changes, the scanner changes, or the staffing pattern changes?

That last question matters more than people admit. In clinical settings, distribution shift is not an academic term. It is Tuesday afternoon when the emergency department is packed, the resident is cross-covering, and the algorithm starts seeing messier inputs than it saw in the pilot.

I do not approve AI deployments that cannot answer those questions with local data, named owners, and an escalation plan. If a team cannot explain who halts the system when safety signals appear, the system is not ready for patient care.

What the best frameworks actually say

The useful part of the current policy conversation is not the hype. It is the convergence. The NIST AI Risk Management Framework, released in 2023, gives organizations a practical way to map, measure, manage, and govern AI risk. WHO’s Ethics and governance of artificial intelligence for health lays out six ethical principles, including autonomy, accountability, transparency, equity, and sustainability. Those are not decorative concepts. They are the minimum scaffolding for clinical AI.

In my experience, hospitals often say they want “responsible AI,” but they mean they want a tool that works most of the time without creating a committee. That is too thin. Responsible AI in healthcare means the opposite: clear accountability, explicit monitoring, and a willingness to stop use when the evidence fails.

Clinical vulnerability is part of the job

I have been wrong before. Early in the ambient documentation wave, I assumed the biggest risk would be accuracy in transcription. Then I watched how a note-writing assistant could distort tone, flatten nuance, and shift what clinicians believed the patient had actually said. The problem was not only words on a screen. It was how those words changed handoffs, coding, and downstream assumptions.

That experience changed my thinking. Now I care as much about workflow integrity as I do about model performance. If a tool saves time but quietly changes the clinical record in a way the team does not notice, the cost can surface days later in the chart, in communication, or in the next clinician’s judgment.

AI in medicine is often an operations issue wearing a clinical costume

Some of the most meaningful AI use cases in hospitals are not dramatic. They are boring in the best sense. Schedulers, ambient scribes, triage support, radiology prioritization, pathology workflow triage, and no-show prediction can each save time if they are implemented carefully. But “carefully” is doing a lot of work.

For example, the Straight Arrow News feature Move fast and heal things? AI tests regulation and medicine’s cautious culture captured the tension well, including my own point that physicians should not have to choose between charting during the visit, running late all day, or taking work home at night. That tension is real, and AI can help. It can also add a new layer of cognitive overhead if every note needs cleanup or every suggestion demands a second review.

The question is not whether the model is smart. The question is whether the entire workflow is smarter after you add it.

Evidence matters, not vibes

The published record is mixed enough to keep everyone humble. Gommers et al. in The Lancet (2026) reported in the MASAI trial that AI-supported mammography screening had higher sensitivity, 80.5% versus 73.8%, with specificity of 98.5% in both groups. That is meaningful, but it is not a blank check for every breast imaging workflow in every health system.

Similarly, the FDA’s regulatory posture exists because performance in controlled settings does not guarantee safe performance after deployment. A model can look excellent in validation and still struggle when the patient mix, scanner hardware, documentation style, or staffing pattern changes. Hospitals need post-market monitoring, not just pre-purchase enthusiasm.

WHO’s 2021 guidance is useful here because it insists that health AI must serve human well-being, accountability, and inclusiveness. NIST adds the operational discipline. Together they push leaders toward a question I wish more boards asked: what does ongoing trustworthiness look like after the go-live applause fades?

There is also a labor angle that healthcare executives cannot ignore. Documentation relief is valuable, but only if the gains are real. A 2025 ambient-scribe study in the Journal of the American Medical Informatics Association reported significant reductions in perceived documentation time and after-hours work, while objective documentation minutes did not always change in the same way over the study window. That mismatch is exactly why I prefer reading both subjective and operational metrics before declaring victory.

What I would not do

I would not deploy a clinical AI tool because the demo looked elegant and the sales team sounded thoughtful. I would not accept a pilot that measures only user satisfaction and ignores false positives, override rates, time-to-action, or downstream chart contamination. I would not let a hospital call something “safe” before it has a named owner, a rollback plan, and a monitoring dashboard that someone actually reviews.

I would also not use AI as a substitute for clinician judgment in high-stakes decisions. A model can rank, flag, summarize, and accelerate. It should not quietly become the de facto decision-maker while everyone pretends humans are still in charge.

The board-level question is simpler than the pitch deck

When I brief hospital leaders, I frame AI in three layers. First, regulatory classification: does this fall under FDA oversight, and if so, what pathway? Second, governance: who is accountable, what is the monitoring plan, and how do we handle drift? Third, clinical utility: does it improve a workflow that matters to patients, clinicians, or both?

That sequence is deliberate. If the answer to the first two questions is fuzzy, the third one almost does not matter. Good intentions do not reduce liability. Good dashboards do not prevent harm by themselves.

There is a broader policy lesson here too. AI in medicine is advancing faster than most institutions can absorb, but speed is not the same thing as readiness. The clinicians I trust most are not the ones who love or hate AI reflexively. They are the ones who ask, with a straight face, “What problem are we solving, what could go wrong, and who notices first?”

Back to that conference room

When the nurse asked who catches the error at 2 a.m., nobody laughed. That was the right silence. The best answer was not a promise of magic. It was a governance structure, a monitoring plan, and a willingness to say no until the tool earned its place.

That is where I am now. I used to think the central issue was whether AI could keep up with medicine. Now I think the more important question is whether medicine can keep its standards intact while AI rushes in. The answer should still begin with patient safety, and it should still end with a human being accountable at the bedside.

For more on my background and clinical perspective, see Dr. Sina Bari’s physician bio and Stanford training. For broader writing on the intersection of medicine and technology, visit sinabarimd.com.

FAQ

What happens if a hospital deploys an AI triage tool without clinician oversight?

The tool can become a silent decision-maker, especially when staff are busy and trust the output too quickly. In practice, that can mean delayed escalation, missed edge cases, and overreliance on a score that was never validated for the local patient mix.

How should a health system evaluate an AI vendor before signing a contract?

Start with regulatory status, then demand evidence on local workflow impact, false positive and false negative rates, and post-deployment monitoring. I would also want named clinical and technical owners, a rollback plan, and a clear explanation of how the model handles drift.

Why do ambient scribes help some clinicians but frustrate others?

They help when they reduce documentation burden without distorting the note or creating cleanup work. They frustrate clinicians when the draft changes tone, misses nuance, or adds enough correction time that the promised time savings disappear.

What is Dr. Sina Bari’s approach to hospital AI governance?

I start with the patient, the workflow, and the failure mode, then work backward to the evidence and oversight structure. If a tool cannot show me a realistic safety case and a monitoring plan, I do not treat it as ready for routine care.

How do FDA and NIST fit together for clinical AI?

FDA addresses whether the tool belongs in a regulated medical-device pathway, while NIST gives institutions a framework for managing AI risk after they buy or build it. Hospitals need both because approval and governance solve different problems.