Analysis / 001

The hospital AI tools I trust are the ones built for boring accountability

A physician-executive view of why hospital AI should be judged less by flashy demos and more by regulatory pathway, workflow fit, and post-deployment monitoring. The safest tools are usually the least theatrical.

Author

Dr. Sina Bari, MD

Physician-Technologist | Healthcare AI Executive | Stanford Medicine

Published

August 25, 2026

Reviewed

August 25, 2026

Last Tuesday, I was in a clinic room with a patient whose chart had already been touched by three different systems, and none of them had actually helped me answer the question in front of us. The intake note was cleaner than usual, the radiology order had a little more structure than I liked, and the vendor demo I had seen earlier that morning promised a smoother future. I still ended up doing what I always do first, I slowed down and looked at the patient.

Hospitals should treat AI like a regulated clinical utility, not a performance demo. The tools worth buying are the ones that can survive FDA scrutiny, local governance, and post-deployment monitoring without making clinicians babysit them all day.

I used to think the key question was whether an AI tool was accurate enough on paper. Then I watched a system look excellent in a demo and fail in the workflow where it mattered, at the handoff between ordering, triage, and documentation. Now I think the real question is whether the tool can be governed after it leaves the slide deck.

I write this from the physician-executive side of the table, where AI is never just software. It is a purchasing decision, a safety decision, and a labor decision. When I evaluate a system, I ask three questions immediately: what regulatory pathway did it come through, what clinical workflow does it actually touch, and who notices when it drifts?

Why the regulatory pathway matters more than the pitch

For hospital leaders, the FDA pathway is not a footnote. It is the first clue about risk, intended use, and the amount of discipline a vendor had to survive before selling the product. The FDA’s De Novo classification pathway exists for novel devices without a predicate, while 510(k) clearance depends on substantial equivalence, and PMA is reserved for the highest-risk class of devices.

In a 2025 analysis in npj Digital Medicine, the FDA had listed 1,016 AI and machine learning enabled medical device authorizations by December 20, 2024. Of those, 96.4% came through 510(k), 3.2% through De Novo, and 0.4% through PMA. That distribution tells me something blunt: the market is dominated by incremental tools, not miracle products, and the burden is on us to separate useful increment from decorative automation.

The other number I keep in mind comes from a JAMA-linked analysis of FDA-cleared AI devices, where only 505 devices, or 55.9%, reported clinical performance studies at approval. That leaves too much room for selective reporting and too much work for hospital buyers who assume an FDA label equals broad clinical readiness. It does not.

When I hear a vendor say the model is “validated,” I want to know validated against what, in whom, and under what distribution shift. The hospital is full of shifts: different scanners, different note styles, different patient populations, different staffing patterns, and yes, different levels of clinician impatience.

The kind of AI I would actually buy

I would buy systems that reduce a concrete burden and expose their failure modes. A radiology prioritization tool that flags an intracranial bleed is not impressive because it uses the word AI. It is useful if it improves turnaround without flooding the worklist with false alarms, and if the alert burden stays tolerable on a night when the ED is already packed.

A 2024 review of radiology AI reported up to 95% nodule-detection sensitivity, up to 94% segmentation accuracy, 30% to 75% scan-time reductions, and 30% to 50% faster reporting in representative workflows. Those are meaningful numbers, but I read them with caution. The same body of literature also shows how fragile these gains can be once the tool leaves a controlled study and meets real operational mess.

For a physician leader, that is the point. I am not buying a performance chart. I am buying a process that can survive implementation.

The strongest public standard for thinking about this is still the NIST AI Risk Management Framework, which pushes organizations toward mapping, measuring, and managing risk across the full lifecycle. That framework fits healthcare better than the usual product launch language, because hospitals already live inside lifecycle thinking. We do not stop caring after procurement. We care after go-live, after the model is retrained, and after the first quiet failure nobody wanted to report.

What I have seen go wrong in practice

In my experience, the worst failures are rarely dramatic. They are boring, and that makes them dangerous. A model overcalls urgency for a few days, the team adapts, then the signal gets normalized. A documentation assistant drops a nuance from the assessment, then a resident copies the phrasing, then the error gets immortalized in the chart.

I have also seen the opposite problem, a tool that looks clinically cautious but quietly creates delay. The patient does not see the delay as a model artifact. They experience it as waiting, rescheduling, or another callback.

One colleague told me during a rollout meeting, “If this thing makes me click three more times, I am done.” That was not resistance to innovation. That was a clinician naming the real cost of low-value automation. I listened, because she was right.

What I would not do is deploy a black-box triage model into an emergency department and call it quality improvement because the dashboard is pretty. I would not let an AI assistant write patient-facing language without clinician review. I would not buy a tool that depends on one champion and collapses when that person is on vacation. And I would not confuse vendor confidence with clinical evidence.

Hospital governance is the product, not the paperwork

There is a reason the governance conversation keeps getting louder. A policy-oriented JAMA discussion framed AI oversight as a lifecycle problem, not a one-time clearance event. That lines up with what hospitals actually need: premarket review, local performance monitoring, escalation pathways, and a mechanism for retiring tools that stop working in the real world.

Clinically, I think the best governance question is embarrassingly simple: if this tool starts behaving badly on Monday morning, who knows by lunch? If the answer is nobody, then the institution has bought risk and called it efficiency.

The physician side of this matters too. Hospitals often think governance is a committee. It is not. It is a muscle. If the committee meets only when there is a complaint, it is already behind.

That is why I prefer AI programs that are narrow, measurable, and easy to audit. A tool that predicts no-show risk, suggests coding support, or surfaces imaging abnormalities can be monitored more cleanly than a vague “clinical copilot” that wants to sit in every workflow at once. Broad claims create broad harm when they fail.

The labor question nobody wants to put on the slide

AI changes labor before it changes medicine. That is uncomfortable for vendors, but obvious to anyone who works in a hospital. It changes who reviews work, who corrects it, who is held responsible, and who gets blamed when the model saves no time at all.

I used to believe the main issue was whether AI could replace some administrative work. Then I saw how often the tool simply moved work from one person to another. Now I think the better question is whether the model reduces cognitive friction for the clinician who still has to make the final call.

There is a supply-chain side to this as well. The labor behind model development includes labeling, annotation, data cleaning, and clinical validation. The people doing that work are often invisible to the final buyer. So are the sites that provided the data. When I brief hospital leadership, I tell them that ethical procurement includes asking where the training data came from and whether the vendor can explain its dataset governance without hiding behind trade secrets.

That is also where sustainability enters the conversation. Training and serving large models costs energy, money, and staff time. A hospital should ask whether a high-compute model delivers enough clinical value to justify its footprint. Some do. Many do not.

What I would tell a hospital board before approving an AI purchase

First, buy for a specific clinical problem, not for brand aura. Second, require evidence that includes subgroup performance and local validation. Third, insist on monitoring after launch, not just during procurement. Fourth, define who can pause the tool when it misbehaves. Fifth, make sure the savings are real, not shifted onto clinicians as hidden work.

The AMA, WHO, and NIST all point in the same direction here, even if they use different language. AI governance should be transparent, measurable, and tied to patient safety. That is boring. Good. Boring is what hospitals are supposed to be when the stakes are high.

If you want a concise external reference point for the regulatory side, the FDA’s public materials on AI-enabled medical devices are useful starting places, especially the agency’s pages on AI-enabled medical devices and De Novo classification. They will not tell you whether a tool fits your hospital, but they will remind you that a label is not a guarantee of safety in use.

Back in the clinic room

By the time I returned to that patient last Tuesday, the cleaner note and the fancier workflow still had not answered the clinical question. I did. That was the point. AI can help me get to the answer faster, but only if it stays accountable to the same standards I would apply to any other clinical tool.

That is where I have landed. I want AI that is humble, auditable, and local enough to be corrected. I want it to make the work lighter without making the judgment sloppier. And when I look up from the screen, I want the patient to matter more than the dashboard.

Dr. Sina Bari, Stanford-trained clinician and physician-leader, writes about AI from inside the operational reality of medicine, where the question is never whether a tool sounds smart, but whether it helps a real person in front of you.

FAQ

What happens if a hospital deploys an AI triage tool without clinician oversight?

The tool can quietly amplify bad routing decisions, increase alert fatigue, or delay care when the model is wrong. In practice, the harm often shows up as extra work for nurses and physicians before it shows up in any formal incident report.

How do FDA 510(k), De Novo, and PMA pathways differ for AI medical devices?

510(k) relies on substantial equivalence to a predicate device, De Novo is for novel lower-risk devices without a predicate, and PMA is the highest-risk pathway. For hospital buyers, the practical lesson is that the pathway signals risk, but it does not replace local validation.

Why do radiology AI tools fail even when published accuracy looks high?

They often fail because the study setting is cleaner than the hospital environment, with different scanners, workflows, and patient mix. A model that looks strong on sensitivity can still create noise, delay, or work redistribution once it meets real operations.

What is Dr. Sina Bari's approach to choosing hospital AI tools?

I look for narrow use cases, clear evidence, local monitoring, and a shutdown plan if the system starts drifting. I avoid tools that depend on hype, vague validation claims, or hidden work for clinicians.

How should a medical group think about AI governance after go-live?

Governance should continue after deployment with ongoing performance checks, escalation pathways, and periodic review of whether the tool still helps. If no one is watching drift, the group is not governing the system, it is just hoping.