micheledpierri.com

AI & Digital Health: Quarterly Review 2026Q3

Article authored by Michele D. Pierri, MD
Last updated: October 2026
Reading time: 10 minutes


Three months ago, the question was whether artificial intelligence was really entering clinical practice.

This quarter, that question feels almost outdated.

AI is already inside clinical workflows: reading images, screening patients, drafting notes, answering clinical questions, supporting diagnosis, monitoring patients at home and, increasingly, being designed to act rather than merely advise.

But Q3 2026 also produced a much more important lesson.

A better model does not necessarily produce better medicine.

One of the most informative studies of the quarter was not a spectacular new benchmark. It was a negative randomized trial. In ESTOP-AKI, machine learning successfully identified hospitalized patients at high risk of acute kidney injury, but triggering an early nephrology consultation did not improve kidney outcomes. Only 41% of selected nondiet and nonelectrolyte recommendations in the intervention arm were fully followed.

The algorithm had done its job. The clinical pathway had not.

That distinction runs through almost everything that mattered between July and September: AI-guided ultrasound, clinical agents, oncology foundation models, diagnostic copilots, ambient documentation and even regulation.

The model is becoming only one component of a larger system.

The workflow is now the model.


Prediction Is Not Prevention

Clinical AI has spent years optimizing discrimination, calibration and benchmark performance. Q3 provided a useful reminder that none of these metrics is a patient outcome.

The ESTOP-AKI randomized trial enrolled 180 high-risk hospitalized patients. An ML system identified patients at increased risk of acute kidney injury and triggered a structured nephrology intervention.

The primary outcome did not improve.

Seven-day peak creatinine change was 0.04 versus −0.03 mg/dL. Stage 1 or greater AKI occurred in 42% versus 36% of patients, while stage 2 or greater AKI occurred in 19% versus 13%. None of these differences was statistically significant.

Perhaps more revealingly, adherence to selected clinical recommendations was incomplete.

This exposes a weakness in the way clinical AI is often evaluated. We tend to test:

patient data → prediction

when the clinically meaningful chain is:

patient data → prediction → notification → clinical action → adherence → outcome.

Every link can fail.

A model with an excellent AUROC can therefore produce no measurable benefit if the intervention triggered by its prediction is ineffective, late, inconvenient or ignored.

That may sound obvious. It has nevertheless taken remarkably long for clinical AI research to move systematically from model evaluation to pathway evaluation.


AI Can Also Improve the Workflow Without Improving Diagnosis

A second randomized study produced an equally useful result.

The ALLIANCE multicenter trial evaluated Prof. Valmed, a certified LLM-based clinical decision-support system, in 82 physicians assessing rheumatology cases.

Diagnostic accuracy improved after physicians received assistance — but it improved similarly whether that assistance came from the LLM system or conventional diagnostic resources.

Top-1 accuracy increased from 22.2% to 33.3% with LLM assistance and from 23.3% to 35.0% with conventional resources. There was no significant difference between the groups.

Yet something else changed dramatically.

Case-processing time was approximately 94 seconds with the LLM versus 206 seconds with conventional resources.

The AI did not make the physicians more accurate than conventional decision support. It made information substantially faster to obtain.

That distinction matters.

The value proposition for clinical AI may not always be “better diagnosis.” It may be faster access to evidence, less cognitive friction, broader differential diagnoses or reduced administrative workload.

In other words, the endpoint worth optimizing may sometimes be the workflow rather than the diagnostic score.

There was also a warning: confidence increased more readily than accuracy, and exploratory analyses raised the familiar concern of AI over-reliance.


“Human in the Loop” Is Not a Safety System

The standard answer to concerns about AI hallucinations is reassuringly simple: keep a clinician in the loop.

A multicenter study published in September tested that assumption.

Junior clinicians reviewed GPT-4o-generated responses containing medical hallucinations across simulated clinical scenarios.

They identified only 15.8% of the hallucinations.

Even more strikingly, 13.1% of clinicians failed to detect a single hallucination across the scenarios they evaluated.

Detection did not meaningfully improve when the clinical risk of the hallucination increased.

This changes the interpretation of one of the most common phrases in medical AI.

A human being present somewhere between model output and patient is not automatically an effective safety mechanism.

Human oversight needs an architecture.

That may mean independent evidence retrieval, explicit uncertainty flags, structured verification steps, escalation thresholds, audit logs or forcing functions for high-risk decisions.

The relevant question is no longer:

Is there a physician in the loop?

It is:

What exactly is the physician expected to verify, when, using what information, and what happens if the physician and the AI disagree?

That is a much harder question — and a much more clinically useful one.


AI-Guided Ultrasound Shows What a Designed Workflow Looks Like

One of the best examples this quarter came from cardiovascular imaging.

In a prospective study of aortic-stenosis screening, nine novice operators received only four hours of training before scanning 1,302 participants using handheld ultrasound with AI acquisition guidance.

The AI-only system achieved an AUC of 0.979.

But the more interesting result came from the complete pathway.

Experts reviewed only around 10% of examinations, and just 4.8% of participants required comprehensive echocardiography.

The final pathway achieved 85.4% sensitivity, 99.7% specificity, a positive predictive value of 91.1% and a negative predictive value of 99.4%.

The important innovation here is not simply that AI interpreted an ultrasound.

AI helped a novice acquire the image, another component interpreted it, uncertainty triggered expert review, and only selected patients progressed to full echocardiography.

Expertise was redistributed rather than replaced.

A second prospective study showed the same principle in obstetrics. Novice operators using blind-sweep ultrasound generated data from which AI estimated gestational age with a mean absolute error of 4.2 days, noninferior to the 4.5-day clinical standard, across settings in Chicago and Nairobi.

This may become one of the most scalable patterns in medical AI:

AI acquisition → automated interpretation → selective specialist escalation.

The health-economic implications could ultimately matter more than another incremental improvement in image-classification accuracy.


Foundation Models Become Clinical Infrastructure

At the research frontier, model scale continued to grow.

NeuroVFM was trained on 5.24 million MRI and CT volumes generated by routine clinical care. Rather than depending primarily on curated public datasets, the project illustrates what its authors call health-system learning: using the heterogeneous data produced inside a health system to build generalist clinical models.

Pathology moved in the same direction.

PRISM2 used approximately 2.3 million whole-slide images, 700,000 pathology reports and 14 million derived question-answer pairs to connect histopathology with clinical dialogue.

COMPASS approached precision oncology from another direction, learning biologically grounded immune concepts from transcriptomic data across more than 10,000 tumors and evaluating them across multiple cancer types and immunotherapies.

NEVA combined pathology and language in neuroblastoma across multiple institutions.

These are impressive systems.

But scale should not be confused with maturity.

None of these projects has yet demonstrated that deploying the model prospectively improves treatment selection, morbidity, mortality or diagnostic outcomes.

The bottleneck has moved.

The central research problem is increasingly not whether we can train a sufficiently powerful medical foundation model. It is whether that model remains reliable when moved across institutions, scanners, populations and workflows — and whether using it actually changes patient outcomes.


Clinical Agents Move From Answers Toward Actions

September also made “agentic AI” considerably less theoretical.

A Nature Medicine study introduced MoChiAgent, an LLM-based system that orchestrates multiple tools to analyze longitudinal maternal and infant electronic health records.

Its predictive engine was developed and internally evaluated using more than 4.4 million longitudinal clinical visits and externally validated in independent maternal and infant cohorts.

For maternal complications, reported AUROCs included 0.89 for placental abruption, 0.89 for premature rupture of membranes and 0.91 for preterm labor.

The architecture matters as much as the numbers. Rather than asking a language model to generate a medical answer directly, MoChiAgent combines longitudinal prediction with a knowledge-search component that retrieves intervention and treatment information from curated medical literature and guidelines.

Earlier in the quarter, the Stanford SCM Navigator provided another glimpse of this direction. The EHR-integrated system screened 6,193 surgical cases for potential medical co-management during real clinical operations, achieving 94% sensitivity against in-workflow physician feedback.

These systems remain bounded.

But the direction is clear: medical AI is moving from isolated inference toward systems that retrieve information, call specialized tools, maintain context and participate in multistep workflows.

And in September, that trajectory acquired a major institutional experiment.


Cardiology Becomes a Test Bed for Autonomous Clinical AI

On September 9, ARPA-H announced the teams selected for its ADVOCATE program — Agentic AI-EnableD CardioVascular CAre TransfOrmation.

The four-year program is worth up to $62.7 million, with up to $33.7 million committed in its first year.

Its goal is unusually ambitious: develop a patient-facing clinical AI system for cardiovascular care capable of performing some actions autonomously while escalating others to human clinicians — and ultimately pursue FDA authorization.

The architecture is particularly interesting.

Separate teams are building patient-facing clinical agents, supervisory safety systems, multisite validation programs and deployment infrastructure. The program is designed to test not just whether an agent can produce a plausible answer, but whether an ecosystem of clinical, supervisory and evaluation layers can operate safely at scale.

Nothing about ADVOCATE should yet be interpreted as evidence that autonomous cardiovascular care works. It is a development and evaluation program, not an approved clinical service.

But it represents something important.

The regulatory question is beginning to shift from:

How do we approve an AI model?

to:

How do we supervise a system that can observe, reason, act and change patient management over time?

For cardiovascular medicine in particular, that is a major transition.


A Small Randomized Trial Shows Where LLMs May Help First

Not every useful clinical application requires autonomous decision-making.

A prospective randomized phase II trial published on September 26 provides a much more modest — and perhaps more immediately deployable — example.

Researchers randomized 268 patients with newly diagnosed prostate cancer undergoing radical prostatectomy to standard preoperative communication or an AI-assisted pathway.

Patients in the AI group first received personalized responses generated by a locally deployed LLM and then underwent routine physician communication.

Post-communication anxiety was lower in the AI-assisted group: mean GAD-7 score 3.2 versus 5.7.

Physician workload also fell substantially. NASA-TLX scores decreased from 56.8 to 39.9, while mean communication time fell from 19.9 to 11.3 minutes.

This is not autonomous medicine.

The LLM did not diagnose the cancer or choose the operation. It handled a bounded, information-intensive task before the physician-patient conversation.

That may be precisely why the result is interesting.

Some of the first clinically meaningful benefits from generative AI may come not from replacing difficult medical decisions but from restructuring the information work surrounding them.


Regulation Finally Starts Chasing Generative and Agentic AI

Regulation also changed significantly during Q3.

The FDA published a discussion paper specifically addressing generative-AI-enabled medical devices, asking for input on risk assessment, premarket evaluation and postmarket monitoring.

By September, the agency reported that it had authorized more than 1,600 AI-enabled medical devices for marketing in the United States.

That number needs careful interpretation. The FDA itself emphasizes that its AI-device list is not comprehensive, and authorization says nothing by itself about comparative effectiveness or clinical outcome benefit.

But the direction is unmistakable: AI-enabled devices are no longer an edge case in medical-device regulation.

In Europe, transparency obligations under the AI Act became operational during the quarter, while the broader framework continues to evolve around high-risk systems and medical-device regulation.

Across jurisdictions, the common regulatory problem is becoming clearer.

Static approval was designed for relatively static products.

Modern medical AI may be updated, connected to external knowledge, embedded in workflows, monitored continuously and — in the agentic case — allowed to initiate actions.

Regulation therefore has to follow the lifecycle, not merely the launch date.


The Most Important Q3 Result May Be a Change in What We Measure

Looking across the quarter, the most important development is not a single model.

It is a change in the unit of evaluation.

The most informative studies increasingly measure systems rather than algorithms.

The aortic-stenosis study measured how many examinations required expert review and how many patients needed comprehensive echocardiography.

ESTOP-AKI measured whether a prediction actually changed kidney outcomes.

ALLIANCE measured diagnostic accuracy and time.

The prostate-cancer trial measured anxiety, physician workload and communication time.

The hallucination study tested whether human oversight actually works rather than merely assuming it does.

And ADVOCATE is being designed around separate clinical, supervisory, evaluation and implementation layers.

This is progress.

Clinical AI should ultimately be judged the same way we judge other medical interventions: not by how impressive the underlying technology appears, but by what happens to patients, clinicians and health systems when we use it.


Key Takeaways

  • Prediction alone is not clinical impact. ESTOP-AKI showed that successfully identifying high-risk patients does not guarantee better outcomes when the downstream intervention fails to translate prediction into effective care.
  • LLMs may improve efficiency before they improve diagnosis. In ALLIANCE, certified LLM support did not significantly improve diagnostic accuracy over conventional resources, but roughly halved assisted case-processing time.
  • “Human in the loop” is insufficient as a safety claim. Junior clinicians detected only 15.8% of experimentally evaluated LLM hallucinations, showing that oversight itself needs to be engineered and tested.
  • Hybrid workflows look increasingly credible. AI-guided ultrasound demonstrates how novice acquisition, automated interpretation, selective expert review and escalation can be designed as one clinical pathway.
  • Foundation models are becoming health-system infrastructure. NeuroVFM, PRISM2, COMPASS and NEVA demonstrate extraordinary scale, but prospective clinical utility remains largely unproven.
  • Agentic AI is moving toward formal clinical evaluation. MoChiAgent illustrates tool-orchestrating clinical architectures, while ARPA-H’s ADVOCATE program will attempt to build and evaluate FDA-authorizable patient-facing cardiovascular agents with dedicated safety and validation layers.
  • Bounded generative-AI applications may reach useful evidence sooner. In a randomized prostate-cancer trial, LLM-assisted preoperative communication reduced patient anxiety, physician workload and consultation time.
  • Regulation is shifting from products toward lifecycles and systems. The FDA is explicitly considering generative AI medical devices, and its authorized AI-device inventory has passed 1,600.

Looking Ahead

The next quarter should be less about asking whether medical AI can perform clinical tasks and more about whether the emerging systems survive contact with real healthcare.

Five questions are worth watching.

  • Will prospective trials of LLMs and clinical agents begin measuring patient outcomes rather than vignette accuracy?
  • Will “human oversight” evolve into explicit, testable safety architectures?
  • Will health systems publish comparative data on AI-generated versus clinician-generated documentation errors?
  • Will regulators define practical pathways for continuously changing and agentic medical AI?
  • Will any of the enormous clinical foundation models now appearing in the literature cross the boundary from impressive research infrastructure to independently validated clinical deployment?

Q3 2026 leaves us in a more interesting place than Q2 did.

The question is no longer whether AI can enter medicine.

It has.

The question is whether we can build the clinical systems around it well enough that entering medicine actually improves care.

The workflow is now the model.


References

  • Early Nephrology Consultation and Acute Kidney Injury in Hospitalized Patients: A Randomized Clinical Trial. JAMA Network Open, July 10, 2026. DOI.
  • Generalization of AI-Based Gestational Age Assessment Using Blind Sweep Ultrasonography. JAMA Network Open, July 9, 2026. DOI.
  • Health system learning enables generalist neuroimaging models. Nature Medicine, July 10, 2026. DOI.
  • Generalizable AI predicts immunotherapy outcomes across cancers and treatments. Nature Medicine, July 3, 2026. DOI.
  • End-to-end multimodal pathology foundation model with clinical dialogue. Nature Medicine, July 31, 2026. DOI.
  • A unified vision-language model for precision oncology and biomarker prediction in neuroblastoma. Nature Communications, July 9, 2026. DOI.
  • An Electronic Health Record–Integrated, Large Language Model–Powered Tool to Triage Surgical Patients. JAMA Network Open, August 20, 2026. DOI.
  • Certified large language model-based diagnostic decision support in rheumatology: the ALLIANCE multicentre randomised controlled trial. medRxiv, September 2026. DOI.
  • Prediction of maternal and infant outcomes from longitudinal electronic health records with a Mother-Child AI agent. Nature Medicine, September 4, 2026. DOI.
  • A multicenter assessment of human oversight of generative AI outputs in simulated clinical decision making. npj Digital Medicine, September 2026. Article.
  • Large language model–assisted preoperative communication reduces patient anxiety and physician workload in prostate cancer: a prospective randomized phase II trial. npj Digital Medicine, September 26, 2026. DOI.
  • ADVOCATE — Agentic AI-EnableD CardioVascular CAre TransfOrmation. ARPA-H, September 2026. Program page.
  • Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. U.S. FDA, 2026. FDA.
  • Artificial Intelligence-Enabled Medical Devices. U.S. FDA, updated September 2026. FDA.
  • AI & Digital Health: Quarterly Review 2026Q2. Michele D. Pierri, July 1, 2026. Previous quarterly review.