Analysis / 001

The Moment the Algorithm Needed a Human

A physician-executive lens on why hospital AI should be governed like clinical infrastructure, not purchased like software. The real question is whether an AI system can survive scrutiny after the vendor demo ends.

Author

Dr. Sina Bari, MD

Physician-Technologist | Healthcare AI Executive | Stanford Medicine

Published

August 18, 2026

Reviewed

August 18, 2026

Last Tuesday, I sat in a conference room while a vendor walked our team through an AI model that promised to flag deterioration before the chart made the story obvious. Fifteen minutes in, the demo looked polished. Twenty minutes in, a nurse manager asked the only question that mattered: “Who is responsible when it misses the one patient we cannot afford to miss?”

Hospital AI should be treated as clinical infrastructure, not as a novelty purchase. The systems that matter most are the ones with clear governance, pre-specified monitoring, and a human owner who can answer for failures when the model drifts, the workflow breaks, or the output conflicts with bedside judgment.

In my experience, the gap between a good demo and safe deployment is where most AI risk lives, and that gap is regulatory, operational, and cultural, not just technical.

I have seen this pattern enough times to be skeptical on instinct. The model may score well in a slide deck, but the real test is whether it can be traced, monitored, and overridden inside a hospital where one bad alert can delay discharge, flood a worklist, or normalize overreliance.

That is why I think the conversation around AI in medicine has been too narrow. We keep asking whether a system is accurate enough, while the more important question is whether it belongs inside a governed clinical workflow at all. If the answer is yes, then it should be handled with the same seriousness we apply to other high-stakes systems, including FDA pathways, post-market surveillance, and local oversight aligned with frameworks from the FDA’s artificial intelligence software for medical devices guidance and the NIST AI Risk Management Framework.

Why I stopped treating AI as a software purchase

I used to think the key question was whether the model performed better than a busy clinician on a retrospective test set. Then I watched how fast good tools degrade when they hit real work, real interruptions, and real edge cases. Now I think the central issue is governance, because performance on paper does not tell me whether the tool can be monitored, versioned, and safely retired.

That shift matters in hospitals. FDA-regulated AI and machine learning medical devices still move through familiar pathways, including 510(k), De Novo, and PMA, but the practical challenge is the same across all of them, which is to keep the system honest after deployment. The FDA’s newer emphasis on predetermined change control plans reflects reality: models change, datasets drift, and hospitals need a way to know what changed before patients feel it.

What the data says about transparency

A 2024 Nature Medicine analysis of 1,016 FDA-cleared or approved AI and machine learning devices found a mean AI Characteristics Transparency Reporting score of 3.3 out of 17, which is a sobering number for a field that keeps promising confidence without enough disclosure. The same study found that only 3.5 percent of devices reported a predetermined change control plan, which tells me the industry is still better at launching tools than proving they can be governed over time.

Another relevant finding came from the broader regulatory literature: a 2024 Lancet Digital Health discussion of large language models in medicine highlighted recurring risks around bias, safety, privacy, transparency, and accountability, with later synthesis papers noting that 7 of the reviewed concern categories, 25.9 percent, centered on bias and fairness. Those are not abstract ethics points. They are the failure modes that show up when a tool is trained on one population, sold to another, and then expected to behave as if context never mattered.

What I would not do

I would not let a hospital buy an AI tool because it “feels efficient” in a demo. I would not allow a model to auto-populate clinical notes, triage priorities, or radiology follow-up recommendations without a named clinician owner, rollback plan, and monitoring dashboard. And I would not accept a vendor answer that substitutes confidence for evidence.

That stance comes from experience, not ideology. I have seen workflows where a well-meaning automation created more work than it removed, because every false positive became someone’s problem downstream. A model that produces 200 extra review items a day is not a productivity tool. It is a queue generator.

How hospitals get the adoption question wrong

The mistake I see most often is the same one in both healthcare and AI policy. We pretend adoption is a binary decision. In reality, the important decisions happen after approval, when the system enters a messy world of staffing shortages, clinical nuance, local culture, and alert fatigue.

That is where physician leadership matters. A hospital board can approve a budget line. A clinical executive has to ask whether the model will improve triage speed, reduce missed deterioration, or simply move work from one tired team to another. If an AI tool raises the signal-to-noise ratio, it may be harmful even if it is technically correct.

My clinical vulnerability

I was wrong once about how quickly clinicians would adapt to a diagnostic support tool. I assumed people would trust useful outputs if the model looked strong enough. Instead, the first issue was not trust, it was workflow friction. The second issue was silence, because users stopped reporting problems once they assumed nobody upstream would act on them.

That surprised me. It also changed my view of implementation. Clinicians do not need another glossy promise. They need tools that explain themselves just enough, fit the work, and fail safely when they are wrong.

“If I have to babysit it, it is not helping me,” one attending told me after a pilot review. I still think about that line because it cuts through the marketing language. If a system creates constant vigilance, it is not automation. It is a second job.

The governance model I would actually use

For a hospital evaluating AI, I want three things before scale: a clear regulatory classification, a clinical owner, and a post-deployment audit loop. The classification tells me what kind of evidence the system should have. The owner tells me who answers when the model conflicts with reality. The audit loop tells me whether performance is stable, drifting, or quietly degrading.

I would also want the hospital to align local policy with public frameworks from organizations like the World Health Organization’s guidance on ethics and governance of artificial intelligence for health. The point is not to copy a document. The point is to stop acting as if each department can invent its own standard of acceptable risk.

Why regulation matters even when the tool is “just software”

AI used in care settings touches diagnosis, monitoring, workflow routing, and documentation, which means it can affect both clinical judgment and operational load. In regulatory terms, that places some systems inside FDA device frameworks and others in a broader governance zone where institutional policy has to fill the gap. That distinction matters because the absence of a device label does not mean the absence of patient risk.

This is also where the physician-executive lens is useful. I care less about whether the model is branded as AI and more about whether it changes the denominator of work, the distribution of risk, or the speed at which clinicians are asked to trust it. Those are board-level questions.

The practical test for responsible deployment

When I evaluate an AI system, I ask a few direct questions. What problem is it solving? What is the harm if it is wrong? Who monitors drift? How is retraining logged? What happens when the underlying workflow changes? If nobody can answer those clearly, the tool is not ready for real clinical use.

That framework is boring, which is exactly why it works. Medicine runs on boring systems that do not break at the worst possible moment. AI should be held to the same standard.

There is a useful lesson in the recent evidence. The 1,016-device Nature Medicine analysis suggests transparency remains thin. The Lancet Digital Health discussion shows that the ethical and regulatory stakes are already well known. The hospital’s job is to translate that awareness into governance that can survive procurement pressure, leadership turnover, and the temptation to equate automation with progress.

Back to the conference room

By the end of that vendor meeting, the nurse manager was still skeptical, and I was grateful for it. Skepticism is often the first line of safety in medicine. We did not reject the tool because it was new. We rejected the fantasy that newness itself is a reason to trust it.

That is where I ended up after years of watching systems enter clinical life. I do not need AI to be perfect. I need it to be governable, auditable, and humble enough to stay in its lane. The model in that room may still have a role. It just does not get to skip the human questions.

Dr. Sina Bari, MD, Stanford-trained physician and clinical AI observer writes about the point where technology meets clinical reality, and that is usually where the useful conversation begins.

FAQ

What happens if a hospital deploys an AI triage tool without clinician oversight?

The tool can amplify small errors into system-wide workflow problems, especially if it pushes patients into the wrong queue or creates alert fatigue. Clinician oversight is what catches mismatches between model output and bedside context before they become operational harm.

How should a hospital decide between FDA 510(k), De Novo, and PMA for an AI tool?

The answer depends on risk, novelty, and whether a predicate exists. In practice, the hospital should know which pathway the vendor used because that changes the strength of the evidence and the degree of regulatory scrutiny behind the product.

Why do so many AI demos look better than real deployments?

Demo environments strip away the messy parts of care, including interruptions, incomplete data, and changing workflows. Real deployment exposes drift, false positives, and the cost of asking clinicians to supervise something that was sold as automation.

What is Dr. Sina Bari’s approach to hospital AI governance?

I start with clinical risk, not novelty. A tool has to be monitorable, attributable, and safe to override, because if nobody owns the failure mode, the hospital has not actually governed the system.

How do I know if an AI note-writing tool is creating unsafe work?

If clinicians spend more time correcting the output than they save, the tool is creating hidden labor. The other warning sign is when errors are subtle enough to be copied forward into the chart without immediate detection.