The AI governance test every hospital misses on the first try
Last Tuesday, I sat in a meeting where a vendor was demonstrating a triage model that promised to sort emergency department patients in seconds. The screen looked polished. The demo sounded smooth. Then one of my colleagues asked a question I have learned to ask early, before anyone gets seduced by a clean dashboard: “Who watches the model when it is wrong, and what happens to the patient in room 12?”
I used to think the main problem was model performance. If the AUROC looked strong enough, if the validation cohort was respectable, if the vendor had a glossy regulatory slide, I thought the rest would mostly work itself out. Then I started seeing how often the failure was operational, not statistical. A model can be technically “right” and still arrive too late, sit in the wrong part of the EHR, or create alert fatigue so dense that nurses stop trusting the queue entirely.
That is why I now think hospital AI governance begins with a clinical question, not a procurement packet. The right question is not whether a tool is impressive. The right question is whether it changes a decision in a way that improves care without adding a new mode of harm. The WHO guidance on ethics and governance of artificial intelligence for health makes the same basic point, with more formal language: AI in health needs ethics, accountability, and human rights built into design and deployment.
In my experience, the first serious error is letting novelty outrun classification. Hospitals talk about AI as if every tool belongs to the same bucket. It does not. A readmission prediction model, a radiology prioritization engine, and a documentation assistant carry different risks, different oversight burdens, and different regulatory pathways. In the United States, those distinctions matter because FDA clearance routes are not interchangeable. A substantial share of AI-enabled medical devices has moved through 510(k), while De Novo and PMA are used far less often. In one large peer-reviewed analysis, 96.7 percent of FDA-classified AI/ML devices were cleared via 510(k), 2.9 percent via De Novo, and less than 1 percent via PMA, which tells me something important: low-friction clearance does not equal low-risk deployment.
I would not buy an AI system because it is “FDA cleared” and then stop asking questions. That is the wrong finish line. Clearance is one gate, not the whole building. A board-level governance process still has to ask about external validation, drift, data provenance, failure modes, escalation pathways, and the exact clinical owner of the model after launch. The NIST AI Risk Management Framework is useful here because it gives hospitals a vocabulary for mapping risks, measuring trustworthiness, and tracking the system after deployment instead of pretending the model is static.
There is a practical reason I keep returning to post-deployment oversight. Clinical AI is not a one-time event. Patient mix changes. Documentation habits change. Imaging protocols change. A model trained on last quarter’s data can become a different creature after workflow redesign, a new note template, or a change in coding incentives. In one hospital meeting, an administrator told me, “We thought it was set once we signed the contract.” That sentence is how drift becomes a patient safety issue.
For radiology and pathology, the stakes are especially concrete. AI often enters these specialties as a labor-savings story, but the hidden clinical question is whether the tool changes the miss rate in a way that matters. I care less about the marketing phrase “workflow optimization” than about the downstream consequence: does the system reduce unnecessary review without letting subtle disease disappear into the background? In a 2024 radiology study discussed in the literature, AI was evaluated as a way to exclude normal chest radiographs and reduce radiologist workload. That is exactly the kind of use case where governance has to be strict, because a tool that safely filters normal studies can also create catastrophic trust if the false-negative profile is not understood well enough.
I have seen a version of this problem in real life. A triage model placed a patient lower in the queue because the inputs looked reassuring, but the chart told a messier story. The patient had a normal-looking number in one field and a dangerous trend in another. The model did what it was told. The system around it failed to notice that the story was incomplete. That is the vulnerability: clinical AI often inherits the incompleteness of the record, then amplifies it with confidence.
There is also a supply-chain issue that gets too little attention. Every AI system depends on labor and data that are usually invisible to the clinician using it. Someone labeled the images. Someone curated the dataset. Someone built the integration. Someone paid for the compute. Someone absorbed the local workflow pain when a model added clicks. If I am evaluating a vendor, I ask who owns the training data rights, how bias was audited, and what happened to the model when it met a population unlike the one in the slide deck. That is not academic fussiness. It is clinical risk management.
One of the few honest sentences in hospital AI is this: I do not trust a model that cannot be monitored after launch. Monitoring has to include performance drift, subgroup behavior, override frequency, and downstream clinical outcomes. If an alert fires 400 times a day and no one acts on it, the model is not supporting care. It is becoming ambient noise. If a model silently shifts its recommendations after a new EHR integration, that is not an innovation story. That is an incident.
What I would not do is let a hospital deploy a high-impact AI tool without a named physician owner, a defined escalation path, and a rollback plan. I would not accept a system that has no local calibration or no mechanism for frontline clinicians to report bad behavior. I would not permit a vendor to hide behind proprietary claims when the model is influencing triage, diagnosis, or resource allocation. If a tool shapes care, the health system has to be able to explain how it works well enough to defend its use in front of patients, staff, regulators, and, if needed, a plaintiff’s attorney.
The broader policy picture is moving in the same direction. The World Health Organization wants ethics and governance at the center. NIST wants risk management to be operational. In medicine, the FDA asks for evidence, not vibes. That combination should make hospital leaders less dazzled by the idea of AI and more serious about the conditions under which it earns clinical trust. I am not anti-AI. I am anti-unstructured adoption.
The best physician-executive posture is probably the least glamorous one. Start with the harm you are trying to prevent. Define the patient population. Name the owner. Measure the baseline. Set the threshold for intervention. Decide who can pause the model. Then test it again after deployment, because the first month is where real-world behavior starts to separate from the pitch deck.
Three weeks after that vendor demo, I was back in clinic, looking at a patient who had already waited too long for the right answer. The conversation was ordinary, which is to say it mattered. I thought about the model in the meeting and about how easy it would be for a system to mistake speed for safety. If we build AI well, it can help clinicians notice more, sooner, and with less waste. If we build it carelessly, it will simply make our blind spots faster. The patient in room 12 deserves the first version, not the second.
FAQ
What happens if a hospital deploys an AI triage tool without clinician oversight?
The tool can start shaping patient flow before anyone notices its weak points. That raises the risk of missed deterioration, inappropriate queue placement, and alert fatigue, especially if the model is calibrated on a population that does not match the local hospital. A clinician owner should be accountable for overrides, escalation, and ongoing review.
How does FDA clearance differ for hospital AI systems that influence diagnosis or triage?
FDA clearance is only one part of the governance picture. Many AI-enabled medical devices have entered through 510(k), while fewer use De Novo and very few reach PMA, so the pathway tells you something about regulatory category, not everything about clinical safety. Hospitals still need local validation, monitoring, and a rollback plan.
What is Dr. Sina Bari’s approach to evaluating a hospital AI vendor?
I start with the workflow, not the marketing. I want to know who owns the model, how it was validated outside the training site, what failure modes were seen, how drift will be tracked, and how frontline staff can report problems. You can read more about Dr. Sina Bari’s physician-executive perspective and clinical background.
Why do AI models that work in one health system fail in another?
They often fail because the data environment changed. Different note templates, patient mix, coding behavior, scan protocols, or staffing patterns can change how the model performs. A system that was accurate in one hospital can become misleading when the input stream and workflow are different.
How should a board think about AI ethics in healthcare?
A board should treat AI ethics as patient safety and operational accountability, not as a branding exercise. The key questions are whether the system is transparent enough to govern, monitored enough to trust, and reversible enough to stop if it causes harm. WHO’s ethics guidance and NIST’s risk framework are useful reference points for that discussion.