Last Tuesday in clinic, the answer was not in the chart
Last Tuesday, I watched a patient tap through a stack of imaging reports on their phone while we waited for a radiology overread that had already been delayed twice. The frustration was familiar. So was the vendor pitch that came later that afternoon, a polished demo promising a generative assistant that would “summarize everything” and “help clinicians move faster.” I have heard versions of that promise for years, and I have also seen what happens when a tool is impressive in a conference room and brittle at the point of care.
The FDA’s cleared AI medical device market is still dominated by classical machine learning, especially in radiology and workflow-adjacent imaging tasks, while generative AI remains largely absent from authorized clinical devices. That gap reflects regulation, evidence standards, and operational risk, not a lack of interest, and it is the key fact hospitals should use when deciding what to buy and what to wait on.
I used to think generative AI would be the first wave to reach the clinical device market because it is so visible, so flexible, and so easy to demo. Then I started looking at the actual approval landscape, not the hype cycle. The pattern is much narrower, much more conservative, and much more revealing. Classical models keep winning clearance because they solve bounded problems, produce testable outputs, and fit into existing regulatory pathways. Generative systems are still missing from cleared clinical tools because the failure modes are harder to constrain, and the evidentiary bar is higher.
That is the overlooked regulatory truth in 2026. If you want the shortest version, it is this: FDA-cleared AI in medicine is mostly prediction, detection, classification, and triage. It is not open-ended generation. The distinction matters for patient safety, hospital governance, and procurement. It also explains why radiology has become the center of gravity. The FDA’s own list of AI-enabled medical devices shows a large, radiology-heavy market, while the agency’s public materials still frame these products as software as a medical device under existing device pathways, including 510(k), De Novo, and PMA when risk rises FDA’s AI-enabled medical device database.
Why classical AI clears and generative AI stalls
The simplest regulatory answer is also the most important one. Classical AI devices usually have a narrower intended use, a fixed output, and a measurable performance target. A lung nodule detector either finds nodules on a defined test set or it does not. A breast density classifier either matches reference labeling within a validated margin or it does not. That kind of device maps cleanly to a verification framework. Generative AI does not.
In my experience evaluating clinical vendors, the first question is not “How smart is the model?” It is, “What exactly does it touch, what can it change, and how will we know when it is wrong?” If the tool drafts an impression, rewrites triage language, or synthesizes notes from multiple sources, the risk surface expands immediately. Hallucinations, hidden omissions, prompt sensitivity, and version drift become device problems, not just software problems. Hospitals feel that difference in the morning list, in handoffs, in follow-up calls, and in the one patient whose recommendation was summarized incorrectly.
The published device landscape backs that up. A 2026 review of FDA-cleared orthopaedic AI devices in JAAOS Global Research & Reviews identified 70 FDA-cleared AI/ML devices relevant to orthopaedic surgery, but the broader market is far larger in radiology and imaging. The point is not that orthopaedics lacks AI. The point is that the cleared use cases stay tightly bounded. Similar patterns appear in radiology reviews across the US, EU, and China, where the regulatory scenario still favors narrow performance claims over open-ended reasoning systems FDA-cleared AI medical devices in orthopaedic surgery, JAAOS Global Research & Reviews, 2026 AI as medical device in radiology across the EU, USA, and China, European Radiology, 2026.
I am not surprised by that anymore. I am more surprised by how often hospital leaders still assume the opposite.
What the approval mix is telling us about hospital risk
FDA-cleared clinical AI is telling us where the regulator is comfortable, and where it is not. Classical AI can be bench-tested, retrospective-tested, and prospectively monitored against a stable target. It fits post-market surveillance. It fits change management. It fits a QMS mindset. Generative AI asks the institution to validate something more fluid, where the output can vary with phrasing, context, and upstream model revision.
That difference is not academic. It affects radiology worklists, sepsis screening, cardiovascular triage, and operating room planning. It also affects the invisible tasks that executives underestimate: incident review, clinician education, policy writing, and escalation pathways when the model produces a plausible but wrong result. A classical classifier can usually be put into a governance box. A generative assistant wants to be in the room, speaking in sentences, and that creates new accountability problems.
Clinical experience also forces humility here. I have been wrong before about how fast “assistant” tools would become safe enough for live use. I assumed interface polish would arrive faster than governance. In practice, governance has lagged behind the demo layer for years, and that delay has protected patients more than it has slowed progress. That is an uncomfortable sentence for a technology optimist, and a necessary one for a physician-executive.
For hospitals trying to make sense of this market, the practical takeaway is simple. Buy bounded tools first. Demand prospective validation. Require a clear intended use statement. Insist on model version control and human oversight. If the vendor cannot explain the fallback workflow when confidence is low, I would not put the tool near clinical decision-making.
And I would not deploy a generative tool as a silent autopilot in triage, discharge planning, or diagnostic summarization. I would not let it draft final clinical language without direct review. I would not let it ingest free-text and produce a recommendation without a documented audit trail. That is where the risk becomes hard to contain.
The evidence trail is narrower than the marketing trail
The evidence base for AI medical devices is not absent, but it is uneven. A scoping review in Chest in 2026 examined clinical evidence supporting FDA review of new medical technology in pulmonary, sleep, and critical care medicine between 2014 and 2024, which illustrates how device evidence tends to accumulate around specific clinical domains rather than across general-purpose reasoning. In dental imaging, a 2026 narrative review in International Dental Journal found a growing set of FDA-approved AI solutions, again concentrated in imaging tasks rather than generative clinical narration. In cardiology, a 2026 scoping review in Heart questioned whether FDA regulation is supporting cardiovascular innovation, which is another way of saying the field is still searching for the right balance between speed and evidence.
One more number matters. A 2026 review in Clinical Orthopaedics and Related Research reported that few FDA-approved AI/ML orthopaedic devices had EU MDR equivalents or peer-reviewed validation. That should make any hospital board pause. A device can be cleared and still not be well supported by the kind of comparative evidence clinicians actually trust. Approval is a floor, not a finish line.
For the clinician-executive reading this, the question is not whether AI should be used. It already is. The question is which part of the workflow should be allowed to carry the risk. Classical AI has earned its place in detection, prioritization, and structured support. Generative AI has not yet earned the same trust in patient-facing clinical devices.
If you want a broader framework for how I think about this kind of threshold-sensitive deployment, I lean on the same discipline I use in regulatory and governance work at sinabarimd.com, and on the clinical perspective I outline on my credentials page as Dr. Sina Bari, MD, Stanford-trained physician. The branding is less important than the method: start with bounded risk, measure what changes, and do not confuse fluency with safety.
What I would tell a hospital board
The board-level question is whether your institution is buying a decision support tool or buying uncertainty. If the system is classical AI, the answer can often be managed with defined metrics, known failure modes, and regular surveillance. If the system is generative, the institution needs a stronger governance layer before clinical use, not after the first incident report.
I would ask four questions before approving anything:
What is the exact intended use. What data did it train on and validate against. What happens when the output is wrong. Who owns the review process when the model changes.
Those questions sound basic because they are basic. They are also where too many AI deployments fail.
One colleague said to me during a credentialing meeting, “It looks great, but I do not know who signs the note when it gets the diagnosis wrong.” That line stayed with me. It captures the current state of the market better than any product brochure. The regulated part of AI medicine is still about accountability, not eloquence.
Back in clinic, the lesson was simple
By the time I returned to that patient’s room, the radiology read was in, the next step was clear, and the conversation was about action rather than waiting. The moment was small, but it clarified the larger point. The tools that matter most in medicine are still the ones that make a narrow task more reliable, not the ones that sound the most human.
I used to think generative AI would be first because it is the flashiest. Now I think classical AI is first because medicine rewards tools that can be verified, bounded, and audited. That is why the approval curve looks the way it does. And that is why hospitals should keep treating generative AI as a governance challenge before they treat it as a clinical device.
FAQ
Why are most FDA-cleared AI medical devices still in radiology?
Radiology is easier to validate because the input is structured, the output is measurable, and the clinical task is usually well defined. That makes it a better fit for 510(k) clearance and post-market monitoring. It is also where imaging data sets and performance benchmarks are deepest.
Can a hospital use ChatGPT-like tools for clinical documentation today?
Yes, but only with tight safeguards and usually not as an autonomous clinical device. I would treat those tools as administrative support until they have a clear intended use, audit trail, and human review workflow. Without that, the risk is incorrect summarization, omitted context, or untraceable errors.
What is Dr. Sina Bari’s approach to evaluating AI vendors?
I start with the clinical workflow, not the model architecture. I want to know what task the tool changes, what failure looks like, and how the hospital will catch it before a patient is affected. If the vendor cannot answer those questions plainly, I move on.
What happens if an AI tool is cleared by FDA but has little peer-reviewed validation?
It may still be legal to use, but that does not make it operationally low-risk. Clearance means the device met a regulatory threshold, not that it has the strongest evidence base for your patient population. A board should ask for local validation and monitoring before broad deployment.
Will generative AI medical devices get FDA clearance in 2026?
Possibly for narrow, tightly controlled uses, but the pathway is still much less mature than for classical AI. The main barriers are validation, safety monitoring, and version drift. I would expect slower, more limited approvals than the marketing cycle suggests.