Last Tuesday, in a crowded clinic room, I watched a resident hover over a new triage dashboard while a patient waited for an answer about whether her chest pain needed the emergency department or another day of observation. The vendor rep had left twenty minutes earlier, the screen was still open, and the resident said the part I hear more often than I like: “It looks smart, but I do not know what it is actually seeing.”
I used to think the main problem with clinical AI was accuracy. Then I spent enough time inside real workflows to see that accuracy is only one layer, and often not the first one that breaks. Now I think the harder problem is governance, because a model can be technically strong and still create harm if it adds noise, hides its reasoning, or pushes clinicians into rubber-stamp behavior.
What I look for before a hospital buys an AI tool
When I evaluate an AI vendor, I start with the same blunt question I would use for a new medication: what is the actual risk, and who is responsible if the answer is wrong? The FDA already separates medical software into pathways such as 510(k), De Novo, and PMA, and it also makes a crucial distinction for clinical decision support, whether a clinician can independently review the basis for the recommendation. That is the first gate I want a hospital to respect, because if the rationale cannot be inspected, the tool is already too opaque for routine bedside use.
The FDA’s Artificial Intelligence in Software as a Medical Device page is useful because it shows how the agency is thinking about lifecycle oversight, not just initial clearance. I also want boards to read the NIST AI Risk Management Framework and the WHO ethics and governance guidance for AI in health, because both put structure around the questions hospitals usually postpone until after launch.
In my experience, the first failure mode is not catastrophic misdiagnosis. It is smaller, duller, and more common. A tool adds one more alert to a clinician’s already overloaded queue, or it flags the wrong patient segment, or it becomes another tab that no one trusts enough to open during a busy shift. That is how a promising algorithm becomes shelfware.
The clinical workflow is where the truth shows up
I have seen well-meaning teams fall in love with model metrics and forget the bedside. AUC looks impressive in a slide deck, but a clinician lives in the next seven minutes of work, the handoff note, the lab callback, the family question, the interruption in the hallway, the pager that will not stop. If an AI tool does not fit that flow, it creates friction even when it is “right.”
The literature has been pointing in this direction for years. Melnick et al. in Mayo Clinic Proceedings (2020) found that EHR usability scored 45.9 on the System Usability Scale, a reminder that clinicians already work inside systems that are harder to use than they should be. A JAMA randomized vignette study in 2023 showed that AI support can improve diagnostic performance, but only when the support is usable and clinically integrated, not just technically impressive. And in a 2024 NIST profile on generative AI risk management, the message is similar: manage risks across the lifecycle, not just at the moment of purchase.
I should admit something else here. I have been wrong before. There were cases where I assumed the problem was model quality, when the actual problem was documentation burden, poor alert logic, or a weak escalation pathway. That kind of error matters, because it changes where you spend the money and the time.
What I would not do
I would not deploy a hospital AI tool simply because it cleared a pilot with a friendly department champion and a clean PowerPoint. I would not let a model influence triage, radiology prioritization, or inpatient deterioration monitoring without documented validation in the local population, explicit fallback rules, and a named owner who can be called when the system drifts. And I would not accept “the vendor monitors it” as a governance answer. It is too thin for clinical care.
That is especially true in areas like imaging, sepsis detection, bed management, and discharge prediction, where the consequence of a false negative is not a tidy spreadsheet error. It is a delayed scan, a missed escalation, a confused nurse handoff, a family sitting in uncertainty, a patient who waits too long. In hospital AI, downstream workflow is part of patient safety.
One reason I keep returning to standards bodies is that they force discipline. The WHO guidance emphasizes ethics and human rights. NIST gives institutions a way to talk about validity, reliability, safety, accountability, transparency, explainability, privacy, and resilience. The FDA keeps reminding developers that device software lives in a regulated ecosystem, not a sandbox. Those are not bureaucratic decorations. They are the minimum architecture for responsible use.
The contrarian view: hospitals do not need more AI enthusiasm, they need fewer unowned decisions
The most useful AI program I have seen, or at least the most credible, is not the one with the slickest demo. It is the one that knows exactly which decisions it can influence, which ones it cannot, and who gets paged when the output is off. That sounds plain. It is. Clinical medicine runs on plain things done consistently.
I used to think governance was what happened after the technology worked. Now I think governance is part of the technology. If the model cannot be audited, if the workflow cannot absorb it, if the clinicians cannot challenge it, and if the hospital cannot explain its use to a patient, then the system is incomplete. Not futuristic. Incomplete.
For a physician-executive, the right question is not whether AI can help. It can. The question is whether the institution can absorb the risk without exporting it to the bedside team. That is where boards, medical staff leadership, compliance, informatics, and frontline clinicians need to sit in the same room before deployment, not after an incident review.
If you want the credentialing lens behind that position, my background is outlined on Dr. Sina Bari’s physician profile and clinical leadership background. I mention it because this perspective comes from watching tools meet real patients, not from a conference stage.
Back in the room
By the time I left that clinic visit last Tuesday, the resident had closed the dashboard. We went back to the basics, symptoms, vitals, exam, the parts of medicine that still carry the weight when software is uncertain. The patient needed judgment, not theater. That is the standard I keep coming back to. AI can support the decision, but the hospital still has to own it, and the clinician still has to live with it.
FAQ
What happens if a hospital deploys an AI triage tool without clinician oversight?
It usually creates two risks at once, false reassurance and hidden work. Clinicians may trust the output too much, or they may ignore it after a few bad experiences, which defeats the point of the tool. The safer model is clinician-in-the-loop review with clear escalation rules and local validation.
How does FDA oversight apply to AI clinical decision support in a hospital?
FDA oversight depends on whether the software meets the statutory exclusion criteria for clinical decision support and, if not, whether it is regulated as a device through 510(k), De Novo, or PMA. In practical terms, hospitals should ask whether a clinician can independently review the basis of the recommendation and whether the product’s intended use matches the setting where it will be deployed.
What should a board ask before approving a generative AI pilot in healthcare?
The board should ask who owns the risk, what data the model was trained and tested on, how performance will be monitored after go-live, and what happens when the tool is wrong. I would also ask whether the pilot changes documentation, triage, or ordering behavior, because those are the places where patient safety and workflow collide.
What is Dr. Sina Bari’s approach to hospital AI governance?
I start with clinical utility, then I test for explainability, workflow fit, and accountability. I am comfortable with AI that makes clinicians faster or more consistent, but I am not comfortable with AI that is unreviewable, unowned, or deployed as a substitute for judgment.
Why do AI tools that look good in demos fail in real clinical settings?
They often fail because the demo hides the messy parts of care, interruptions, uncertain data, handoffs, and competing priorities. A good demo shows prediction quality, but a good deployment shows whether the tool fits the actual work of medicine. If it creates more clicking, more alerts, or more ambiguity, clinicians will stop using it.