Last Tuesday, I watched a vendor demo for an agentic AI system that claimed it could review a complex admission, pull the chart, summarize the problem list, propose orders, and draft the sign-out note before the intern had finished opening the encounter. The room went quiet in the way clinician rooms do when everyone realizes the demo is winning the argument faster than the evidence is. I have seen that look before. It usually appears right before a hospital confuses automation with judgment.
Agentic AI in clinical medicine should be treated as a governed, constrained workflow layer, not a free-roaming decision maker. The systems that matter most are the ones that can act, but only inside a zero trust clinical architecture, with human intent, tool access, audit trails, and escalation rules explicitly bounded.
That framing is already showing up in recent work on autonomous healthcare security, human-guided multimodal prediction, and regulatory definitions of agentic AI, and it changes the deployment question from “Can it do the task?” to “What failure modes are we willing to tolerate?”
I used to think the main barrier to clinical agentic AI was model quality. Then I spent enough time watching edge cases break in real workflows that my view shifted. Now I think the harder problem is governance under partial autonomy, because a system that can call tools, write back to the chart, or trigger downstream work can create harm even when its language sounds polished.
That is why I read recent papers such as Caging the Agents: A Zero Trust Security Architecture for Autonomous AI in Healthcare and Human-Guided Agentic AI for Multimodal Clinical Prediction with more interest than the usual accuracy leaderboard. The most important number in a clinical deployment is not always AUC. Sometimes it is the number of systems, permissions, and hidden side effects a model can touch before a human notices.
Why agentic AI feels different at the bedside
The first time an AI agent misreads a chart, that is a model problem. The first time it misroutes a task, drafts an order set, or pings the wrong team without anyone noticing, that is an operations problem. The first time that chain reaches a patient, it becomes a safety problem. In clinical medicine, those are different layers, and hospitals love to collapse them into one neat procurement slide.
I have seen a simple discharge workflow balloon into three separate failure points: the model summarized the wrong anticoagulation history, the agent surfaced an old medication list because the retrieval scope was too broad, and the final note looked credible enough that nobody reread the source data. None of those failures required a catastrophic hallucination. They only required confidence, latency, and weak supervision.
Dr. Sina Bari’s physician-executive background matters here because the real question is not whether an agent can produce text. It is whether a hospital can define which actions are allowed, logged, reversible, and clinically owned. Board members and informatics teams often ask me for a yes-or-no answer. I usually answer with a workflow diagram instead.
Clinical autonomy needs a security model, not just a model card
One of the strongest papers in the 2026 literature, Caging the Agents, argues for a zero trust security architecture for autonomous AI in healthcare. That matters because clinical agents are not isolated chatbots. They are systems with retrieval, tool use, memory, and sometimes write access to enterprise infrastructure. If you would not let a human contractor roam your EHR with unlimited permissions, you should not let an agent do it either.
The same logic appears in Securing the Agent: Vendor-Neutral, Multitenant Enterprise Retrieval and Tool Use, which is useful precisely because it treats retrieval and tool invocation as attack surfaces, not convenience features. In a hospital, a bad retrieval query can become a privacy event. A bad tool call can become a clinical event.
Regulators are also catching up. The paper Security, privacy, and agentic AI in a regulatory view makes a point clinicians should take seriously: once a system begins taking actions, it starts to resemble software that belongs in a much more mature governance category. That is where frameworks from the FDA, NIST, and enterprise security teams matter. Hospitals already know how to manage PACS, medication dispensing, and device access. Agentic AI should be held to the same discipline.
For context, the FDA’s Software as a Medical Device framework and its pathways for 510(k), De Novo, and PMA are not perfect fits for every agentic workflow, but they are a useful reminder that intended use and risk class matter. If an agent can influence diagnosis, triage, or treatment, the burden of proof rises quickly.
What I would not do with agentic AI
I would not let an autonomous agent directly place orders, route messages to on-call teams, or edit the chart without explicit human confirmation in the live workflow. I would not allow it to decide when a patient needs escalation based only on a hidden chain of prompts. And I would not buy the argument that because a system is “doctor-in-the-loop,” the loop is automatically sufficient. That phrase has become a comfort blanket for weak governance.
I also would not deploy an agent as a black box productivity layer and then measure success only by time saved. Time saved for whom? If nurses spend ten minutes correcting AI-generated work, the hospital has merely shifted labor, not reduced it. If the system saves physician clicks but increases downstream verification, the net gain may be negative.
That is why I prefer the framing in Grounding Clinical AI Competency in Human Cognition Through the Clinical World Model and Skill-Mix Framework. Clinical competency is not one thing. It is a mix of perception, pattern recognition, prioritization, communication, and judgment. Agentic systems should be assessed against those components, not against a vague promise of “doing the clinician’s job.”
Where the evidence is strongest right now
The most credible 2026 papers do not argue that agentic AI is ready for free autonomy. They argue for constrained, human-guided use cases where the model helps with prediction, workflow triage, coding, retrieval, or multimodal synthesis under oversight. AgentDS is important because it frames clinical prediction as a guided process, which matches how medicine actually works. A clinician rarely wants a single answer. We want a reasoned path to an answer, with room to override.
There is also a useful trend in specialized models. Cura 1T: Specialized Model for Agentic Healthcare reflects a broader pattern I expect to persist, namely that hospitals will eventually prefer narrower systems with clear task envelopes over general agents that try to do everything. In medicine, narrower is often safer. It is also easier to audit.
Clinical validation still matters. In radiology and pathology, the old lesson holds: a model can look impressive in a benchmark and still fail in the mess of real workflow, where priors are incomplete, images are noisy, and humans are multitasking. For agentic systems, the benchmark should include failure handling, escalation behavior, and what happens after a tool call goes wrong. That part is usually missing. It should not be.
The self-correction that changed my view
I used to think the right governance question was whether clinicians trusted the system enough to use it. Now I think the better question is whether the system deserves to be trusted with any action at all. Trust is earned in steps. Read-only first. Then draft-only. Then supervised suggestions. Only after that should anyone even talk about limited autonomy.
That sequence sounds slow to product teams. It feels slower than the mood around AI. But medicine has a long memory for shortcuts. I have seen enough alert fatigue, failed handoffs, and brittle interface logic to know that the fastest path to a problem is usually the one that begins with “we can always monitor it later.” Later is where incidents happen.
There is a quantitative reason to stay humble. The broader survey literature on trustworthy agentic AI now spans safety, robustness, privacy, and system security across dozens of studies, which tells me the field itself has moved from novelty to risk management. The names of the papers matter less than the pattern they reveal: everyone is rediscovering that autonomy expands the attack surface.
What hospitals should ask before deployment
If I were briefing a hospital board, I would ask six questions before approving an agentic AI pilot. Can the system be constrained by role, task, and data scope. Can every tool call be logged. Can any action be reversed. Does the vendor support independent security review. What is the clinician escalation path. And what harm metric will trigger shutdown.
I would also ask whether the use case actually needs an agent. Sometimes a well-designed rules engine or a conventional clinical decision support tool is safer and easier to validate. Hospitals tend to overbuy generality because it sounds modern. Medicine rewards specificity.
That is the practical bridge between AI governance and clinical care. The question is not whether an agent can help a clinician think. The question is whether the hospital can build a system where help is constrained, visible, and accountable. That is what zero trust means in a ward, not just in a data center.
Back to that Tuesday demo
By the end of the demo, the room had stopped talking about speed and started talking about permissions, auditability, and who would own the cleanup when the agent made the wrong move. That was the right conversation. I left thinking the same thing I think more often now: clinical agentic AI is not a destiny story. It is a governance story.
Last Tuesday, the vendor showed me an agent that could write. What I cared about was whether it could be caged, supervised, and made boring enough for real medicine. That is the standard I would use at the bedside, in the ICU, and in the boardroom.
FAQ
Can an agentic AI system place orders in a hospital without physician review?
No. In practice, any agent that can place orders should require explicit human confirmation and a narrow permission scope. If a hospital skips that step, the risk is not just a wrong order, but a chain of downstream actions that look legitimate because the software produced them.
What is the safest first use case for agentic AI in clinical medicine?
Read-only tasks with clear supervision are the safest starting point, especially chart summarization, retrieval, and draft documentation. Those uses let teams measure accuracy, latency, and failure modes before the system is allowed to influence care.
How does Dr. Sina Bari approach clinical AI governance?
Dr. Sina Bari’s approach is to treat agentic AI like a clinical workflow intervention with security boundaries, audit trails, and escalation rules. He starts with the question of what the system is allowed to touch, not what it claims to be able to do.
What happens if an agentic AI tool gets retrieval wrong inside the EHR?
The failure can spread quickly because the wrong information may be summarized, documented, or surfaced to the next clinician as if it were verified. That is why retrieval scope, logging, and human review are essential, especially in high-stakes settings like anticoagulation, sepsis, and discharge planning.
Why are zero trust principles relevant to clinical AI?
Zero trust matters because an agentic system should never be assumed safe simply because it sits inside a hospital network. The right posture is to verify every action, restrict every permission, and monitor every tool call as if the model could fail in ways that are clinically meaningful.