micheledpierri.com: statistics, data analysis and coding

Nexus of Statistics, Data analysis, Coding, Art and Medicine

Menu
  • Home
  • Courses
    • Python Foundation
    • Statistics
    • Data Analysis
    • Machine Learning
  • Blog
    • All Pages
    • Health Informatics
    • Programming
    • Art
  • Illustrations
  • About
  • Contact
Menu
Home / Archives for Michele Danilo Pierri / Page 2

Author: Michele Danilo Pierri

Michele D. Pierri is a cardiac surgeon and cardiovascular physiopathology researcher with a strong interest in artificial intelligence, medical data science, clinical decision support, and digital health. His work focuses on the intersection between medicine, technology, and computational methods, with the aim of translating complex biomedical concepts into clear, practical, and clinically meaningful insights.
an apple on top of a stylized wooden stand shaped like a head (training dummy), the apple surrounded by thin concentric rings like a target, a single arrow embedded near the apple but within an ‘acceptable ring

When “Not Worse” Is Enough: Understanding Non-Inferiority Trials in Medicine

Posted on June 12, 2026August 11, 2026 by Michele Danilo Pierri

Non-inferiority Trials in Medicine:

Last updated: June 2026

Author: Michele D. Pierri

Reading time: 15–20 minutes


In clinical research, we are used to thinking that a new drug, device, or procedure should be tested by asking a simple question:

Is it better than what we already have?

This is the logic of a superiority trial. A new treatment is compared with an existing treatment, placebo, or standard care, and researchers try to demonstrate that the new option produces better outcomes.

But many modern clinical trials — especially those involving new drugs, medical devices, anticoagulants, antibiotics, interventional cardiology devices, surgical technologies, or minimally invasive procedures — are not designed to prove superiority. Instead, they are designed to prove non-inferiority.

At first glance, this may sound suspicious. Why would we accept a new treatment that is not clearly better? Is this just a convenient way for companies to get a product approved without proving that it improves patient outcomes?

The answer, though, is less clear-cut. Non-inferiority trials are scientifically legitimate, but they are also easy to misunderstand, and in some cases deliberately misused.

Superiority, non-inferiority, and equivalence: three different questions

The first mistake is to confuse three different concepts.

A superiority trial asks:

Is the new treatment better than the control?

A non-inferiority trial asks:

Is the new treatment not unacceptably worse than the control?

An equivalence trial asks:

Are the two treatments similar enough in both directions?

These are not interchangeable.

A non-inferiority trial does not prove that two treatments are identical. It does not prove that the new treatment is equivalent to the old one. It only tries to exclude that the new treatment is worse than the standard treatment by more than a pre-specified clinically acceptable amount. This amount is called the non-inferiority margin (Δ). Regulatory and reporting guidance documents (e.g., FDA guidance; the CONSORT extension for non-inferiority/equivalence) emphasize that this margin must be defined and justified before results are known.

Why use a non-inferiority design at all?

Non-inferiority trials are often used when an effective standard treatment already exists.

Imagine a disease for which a proven treatment reduces mortality. In that situation, it may be unethical to compare a new drug with placebo, because patients in the placebo arm would be denied effective therapy. This is one of the classic reasons for using an active-control non-inferiority trial, where the new treatment is compared against the accepted standard.

But there is another important reason.

A new treatment may not be more effective than the old one, but it may offer other advantages:

  • fewer adverse effects;
  • simpler administration;
  • no need for laboratory monitoring;
  • shorter hospital stay;
  • lower invasiveness;
  • better patient adherence;
  • lower cost;
  • easier logistics;
  • improved quality of life.

For example, a new oral anticoagulant may not reduce thromboembolic events more than warfarin, but it may avoid repeated INR monitoring. A new antibiotic may not cure more infections than the standard drug, but it may require fewer daily doses. A new device may not reduce mortality more than an established procedure, but it may be less invasive or easier to implant.

Cardiology offers the clearest example of this logic at scale: the low-risk TAVI trials were designed against SAVR as non-inferiority studies, and the 2025 ESC/EACTS guidelines translated that evidence into a lowered age threshold for the transcatheter approach.

In these cases, the clinical question is not necessarily “Is the new treatment better?”. It may be:

Can we accept a small possible loss of efficacy in exchange for other meaningful benefits?

That, at least, is the setting where these trials make genuine clinical sense.

A simple numerical example

Suppose we are testing a new device against a standard device. The endpoint is a negative outcome, such as major complications within 30 days.

The standard device has a complication rate of 5%.

The new device has a complication rate of 6%.

The absolute difference is:

6% − 5% = +1%

So the new device appears slightly worse.

But before the trial started, investigators defined a non-inferiority margin of 2%. This means that the new device would be considered clinically acceptable if it did not increase complications by more than 2 percentage points.

Now imagine the 95% confidence interval for the difference is:

+0.2% to +1.8%

The entire confidence interval is below the non-inferiority margin of +2%.

So the study may conclude:

The new device is non-inferior to the standard device.

But notice something important: the confidence interval is entirely above zero. That means the new device is also statistically worse than the standard device, although still within the pre-defined acceptable margin.

This is not a contradiction. It reflects the fact that non-inferiority is not the same as superiority. A treatment can be statistically worse and still be considered non-inferior if the difference is smaller than the accepted clinical margin.

How to read the confidence interval in a non-inferiority trial (quick guide)

Most non-inferiority conclusions are CI-based. The key is: which side of the CI must stay within the margin depends on how the effect is defined.

  • If the endpoint is an undesirable event (lower is better) and you analyze a risk difference (new − control), non-inferiority is typically shown when the upper bound of the CI is below Δ.
  • If the endpoint is a desirable event (higher is better) and you analyze a difference (new − control), non-inferiority is typically shown when the lower bound of the CI is above −Δ.
  • If you use ratios (risk ratio, hazard ratio), the same logic applies but with a ratio margin (e.g., HR < 1.25). Always check the paper’s estimand and margin definition.

If this sounds pedantic, it is not: many misinterpretations come from not being explicit about direction and scale.

The non-inferiority margin: the most important number in the trial

The whole trial depends on one number: the non-inferiority margin (Δ).

This is the largest loss of efficacy (or increase in harm) that investigators are willing to accept.

If the margin is too strict, the trial may require a very large sample size and may fail to show non-inferiority. If the margin is too generous, almost any treatment can look acceptable.

That is where things get difficult.

Suppose the standard treatment has a 5% event rate. A new treatment has a 7% event rate. That is a relative increase of 40%:

from 5% to 7%

If the non-inferiority margin is set at +2%, the new treatment may still be declared non-inferior.

But would clinicians and patients really accept a 40% relative increase in complications?

Maybe yes, if the new treatment has major advantages. Maybe no, if those advantages are marginal. The answer is rarely obvious in advance.

This is why the margin cannot be only a statistical decision. It must be a clinical decision.

Regulators often discuss margin choice in terms of two related concepts:

  • M1: the estimated effect of the active control compared with placebo, based on historical evidence;
  • M2: the largest clinically acceptable loss of that effect.

Put simply: the non-inferiority margin should not be so wide that it allows the new intervention to lose most (or all) of the benefit that made the standard treatment worth using in the first place.

Why non-inferiority trials can be attractive to sponsors

Non-inferiority trials can be scientifically appropriate. But they can also be commercially attractive.

If a company knows that its new drug or device is unlikely to be better than the existing standard, a superiority trial may be risky. A non-inferiority trial offers another path: the product does not need to prove that it is better on the primary endpoint. It only needs to prove that it is not worse beyond the accepted margin.

That logic holds, but only when the new product actually delivers on those advantages.

The risk is that the marketing message becomes:

“As effective as the standard treatment.”

When the more accurate interpretation may be:

“Not shown to be unacceptably worse than the standard treatment, according to a margin chosen before the trial.”

These two statements sound similar, but they are not the same.

Why non-inferiority is not automatically easier

It is tempting to say that non-inferiority trials are “easier” because they do not require superiority.

But statistically and methodologically, a good non-inferiority trial is not necessarily easier. In some ways, it is more fragile.

In a superiority trial, poor adherence, imprecise measurements, protocol deviations, or crossover between groups usually make it harder to detect a true difference. They push the trial toward a neutral result.

In a non-inferiority trial, the same problems can be dangerous. If everything blurs the difference between treatments, the new treatment may appear “not much worse” simply because the trial was not good enough to detect meaningful differences.

This is why the concept of assay sensitivity is so important.

Assay sensitivity means that the trial is capable of distinguishing an effective treatment from an ineffective or less effective one. In a placebo-controlled trial, this can often be assessed directly. In an active-control non-inferiority trial, it is harder to verify, because there is usually no placebo group. The trial must rely on the assumption that the active control would have performed as expected if placebo had been included.

The “constancy assumption”

Another hidden assumption is the so-called constancy assumption.

This means that the benefit of the standard treatment observed in older trials is assumed to still apply in the current trial.

But medicine changes.

Patients change. Diagnostic criteria change. Background therapies improve. Surgical and interventional techniques evolve. Event rates decline. Follow-up becomes different. Endpoints are redefined.

A standard treatment that showed a large benefit decades ago may have a smaller absolute benefit today because baseline risk is lower. If the historical treatment effect is no longer valid, the non-inferiority margin may become questionable.

This is particularly relevant in areas such as cardiology, oncology, infectious diseases, intensive care, and surgery, where standards of care evolve rapidly. It is not always clear, when designing a new trial, whether the historical reference still holds.

A second example: when the margin changes everything

Imagine a trial comparing a new antibiotic with a standard antibiotic.

Clinical cure rate:

  • standard antibiotic: 90%
  • new antibiotic: 86%

Difference:

−4 percentage points

The new antibiotic is less effective.

Now consider two possible non-inferiority margins.

Scenario A: margin = 5%

The new antibiotic loses 4 percentage points. This is within the allowed margin.

Conclusion:

Non-inferior.

Scenario B: margin = 3%

The new antibiotic loses 4 percentage points. This exceeds the allowed margin.

Conclusion:

Not non-inferior.

The data are identical. The conclusion changes only because the margin changes.

This is why the non-inferiority margin is not a technical detail. It is the foundation of the entire study.

Common pitfalls (what can make a non-inferiority trial misleading)

  • Comparator issues: the “standard” treatment may be suboptimal (dose, timing, operator expertise), making non-inferiority easier to show.
  • Adherence and crossover: non-adherence tends to make groups look artificially similar (dangerous in non-inferiority).
  • Endpoint choices: composite endpoints can hide clinically important trade-offs.
  • Over-interpretation: “non-inferior” is sometimes communicated as “equivalent” or “just as good,” which is not what the design proves.

What readers should look for

When reading a non-inferiority trial, especially one involving a new commercial product, the most important questions are:

1) Was a non-inferiority design justified?

Was there already an effective standard treatment? Would placebo have been unethical? Does the new treatment offer plausible advantages other than efficacy?

2) Was the margin clinically reasonable?

Do not just ask whether the margin was pre-specified. Ask whether it was clinically acceptable. A margin can be statistically convenient and clinically unacceptable.

3) Was the comparator appropriate?

The control treatment must be the real standard of care, used at the correct dose, with correct timing, and under appropriate conditions. A weak comparator makes the new treatment look better.

4) Were both intention-to-treat and per-protocol analyses reported?

In superiority trials, intention-to-treat analysis is usually conservative. In non-inferiority trials, it may not be. Non-adherence and crossover may make treatments look artificially similar. For this reason, both intention-to-treat and per-protocol analyses are often important, and consistency between them strengthens the conclusion.

5) Was superiority also tested (and was it pre-specified)?

Sometimes a trial first shows non-inferiority and then tests superiority. This can be valid if properly pre-specified and statistically controlled. But post-hoc claims of superiority should be treated cautiously.

6) Are the claimed advantages real?

If the new treatment is slightly less effective but safer, cheaper, simpler, or better tolerated, it may be clinically valuable. But those advantages must be demonstrated, not merely suggested.

The practical interpretation

A non-inferiority result should never be read as:

“The two treatments are the same.”

A better interpretation is:

“The trial did not show that the new treatment is worse than the standard by more than the pre-specified acceptable margin.”

That sentence is longer, less attractive, and less marketable.

But it is more accurate.

Conclusion

Non-inferiority trials are not inherently problematic. They are often necessary and ethically appropriate when an effective treatment already exists and placebo would be unacceptable. They can help identify treatments that are slightly less effective but safer, simpler, less invasive, or more acceptable to patients.

They do, however, require careful reading.

The question worth asking is not only whether non-inferiority was demonstrated. The more important one is:

Was the amount of possible inferiority clinically acceptable?

In other words, the entire interpretation depends on the margin.

A well-designed non-inferiority trial can support good clinical decision-making. A poorly designed one can turn “not clearly worse” into a misleading impression of “just as good.”

For clinicians, researchers, and patients trying to make sense of trial reports, the practical takeaway is straightforward:

Whenever a study says “non-inferior,” always ask: non-inferior by how much, compared with what, and in exchange for which real benefit?

FAQ

Does “non-inferior” mean “equivalent”?

No. “Non-inferior” means the trial excluded that the new treatment is worse than the control by more than the pre-specified margin. “Equivalent” requires showing differences are small in both directions.

Why not just run a superiority trial?

Because in many settings placebo is unethical, and the goal is to trade a small potential loss of efficacy for meaningful advantages (safety, simplicity, cost, quality of life).

What is the non-inferiority margin (Δ)?

It is the maximum difference considered clinically acceptable. The entire conclusion depends on how Δ is chosen and justified.

References

  • Food and Drug Administration. (2016). Non-inferiority clinical trials to establish effectiveness: Guidance for industry. U.S. Department of Health and Human Services. https://www.fda.gov/media/78504/download
  • International Council for Harmonisation. (2000). ICH E10: Choice of control group and related issues in clinical trials.
  • Kaul, S., & Diamond, G. A. (2006). Good enough: A primer on the analysis and interpretation of noninferiority trials. Annals of Internal Medicine, 145(1), 62–69.
  • Piaggio, G., Elbourne, D. R., Pocock, S. J., Evans, S. J. W., Altman, D. G., & CONSORT Group. (2012). Reporting of noninferiority and equivalence randomized trials: Extension of the CONSORT 2010 statement. JAMA, 308(24), 2594–2604.
  • Schumi, J., & Wittes, J. T. (2011). Through the looking glass: Understanding non-inferiority. Trials, 12, Article 106.
  • Snapinn, S. M. (2000). Noninferiority trials. Current Controlled Trials in Cardiovascular Medicine, 1(1), 19–21.
a surgical team, including a surgeon, an assistant, an anesthetist, and a scrub nurse, is about to perform an operation but instead of a human being, the patient is a computer

A Day Off From the OR. Or So I Thought

Posted on June 10, 2026August 16, 2026 by Michele Danilo Pierri

Upgrading a local AI workstation from an RTX 3060 12GB to an RTX 3090 24GB with real-world LLM, Flux and Whisper benchmarks.

Article authored by Michele D. Pierri, MD

Cardiac Surgeon & Medical Technology Researcher

Last checked: June 10, 2026

Reading time: 5 minutes


There are days when you don’t go to the operating room. Sundays, public holidays, the occasional free afternoon when the on-call schedule aligns in your favor. On those days, surgeons do what everyone else does: read, rest, spend time with family.

Or they operate on their PC.

This is a clinical account of what happened on June 8, 2026, in a home study in Ancona. The patient — a desktop computer, age approximately four years — was admitted for elective hardware replacement. Diagnosis: GPU insufficiency with secondary power supply compromise. The procedure: combined transplantation of graphics processing unit and power supply unit. Anesthesia: none required. The patient does not feel pain. Or if it does, it hasn’t complained yet.


June 8, 2026 — Operative Report, Case #001

The PC-patient
The patient

Preoperative Assessment

Every surgical intervention begins with a proper workup. In cardiac surgery, we assess ventricular function, coronary anatomy, risk scores. Here, the workup was somewhat different, but the logic was identical: identify the failing organ, understand why it’s failing, and select the appropriate replacement.

The patient presented with a well-documented history of progressive performance degradation. The primary GPU (an NVIDIA RTX 3060 with 12 GB VRAM) had been adequate for its original function. It was not adequate anymore. The ecosystem I’ve been building over the past months (local LLMs, image generation with Flux, audio transcription, the occasional multi-agent pipeline) had pushed the hardware past its operational envelope. Mistral-Nemo inference at 30 tokens per second. Flux image generation at 92 seconds per frame. Acceptable numbers two years ago. Not acceptable when you’re trying to build a production workflow.

The secondary finding was the power supply: a Corsair 600W non-modular unit. Non-modular means the cables are permanently attached: you get all of them whether you need them or not. In a patient already showing signs of internal disorganization (cable routing that could generously be described as “organic”), a non-modular PSU is a contraindication to clean surgical access. It had to go.

The selected implants: an NVIDIA RTX 3090 24GB Gainward and a Corsair RM850x 850W fully modular. The VRAM doubles. The power budget nearly doubles. The cable chaos — theoretically — resolves.


The Donor Organ

Corsair RM850x — connectors panel
Corsair RM850x — connectors panel

Before any transplant, you prepare the donor organ. In thoracic surgery, the harvested heart is inspected, flushed, assessed for structural integrity. Here, the Corsair RM850x was unboxed on the same floral tablecloth that would serve — entirely unintentionally — as the operative field for the entire procedure.

Look at that connector panel. MOTHERBOARD. CPU/PCIe. 12V-2×6. SATA/PATA. Each port labeled, each cable optional, each connection deliberate. There is a certain surgical elegance to a fully modular power supply: you attach only what the anatomy requires. No redundant structures, no extraneous tissue. The organ is clean. The implantation will be clean. In theory.

NVIDIA RTX 3090 24GB Gainward
NVIDIA RTX 3090 24GB Gainward

The RTX 3090 Gainward, for its part, arrived in packaging that suggested its designers were also not entirely unaware of the theatrical potential of high-end graphics cards. It is a large device. Substantially larger than the RTX 3060 it would replace. Sometimes the anatomy demands upsizing.


Opening the Field

The operative field
The operative field

The case was placed on its side on the operative table — the floral tablecloth, which by this point had committed fully to its role as surgical draping. The side panel was removed. The field was opened.

What emerged was exactly what an intraoperative view sometimes reveals: more complexity than the preoperative imaging suggested. The existing cable routing — I will be generous here — reflected four years of incremental decisions made by someone who was always planning to fix it properly later. The mediastinal anatomy was present. The structures were identifiable. The access was not ideal.

The red-handled scissors visible in the background were, I want to clarify, for cable ties. Not for anything else. Though I understand why the photograph raises questions.

With the field open, the sequence became clear: explant the PSU first (it sits at the bottom of the case, access is direct), then address the GPU (which occupies the PCIe slot mid-chassis and requires disconnection of the PCIe power cables before extraction), then close in reverse order with the new components.

Textbook operative planning. The execution, as always, is where things get interesting.


Explantation: The Old Heart Comes Out

non-modular Corsair 600W
non-modular Corsair 600W

Here it is. The explanted organ, photographed post-removal on the operative field.

The non-modular Corsair 600W, with its full complement of permanently attached cables — every wire that was ever wired into it, regardless of whether it was ever used, spread out on the floral tablecloth like a vascular pedicle after harvest. The braided sleeving in black, yellow, and red is not decorative: those are the actual voltage rails. Twelve volt. Five volt. Three point three volt. The colors of the cardiovascular system, if the cardiovascular system ran on direct current.


Anastomosis: Connecting the New Organ

RM850x outside the case
RM850x outside the case

This photograph documents what in surgical terms would be called the anastomotic phase: the moment when the new organ, still partially external to the operative field, is connected to the host’s vascular supply before final positioning.

The RM850x sits on the table beside the open case. The modular cables — all black, all selected specifically for this configuration, no excess — run from its connector panel into the chassis interior. This is cable management done properly, which is to say: done once, deliberately, before everything is enclosed and inaccessible for the next four years. In surgery, you want your anastomoses under direct vision. You don’t want to be working around other structures, correcting an angle you can’t quite reach. Same principle applies here. This is the moment for precision.

The 12V-2×6 connector for the GPU — a relatively recent standard, replacing the older 8-pin PCIe configuration — carries up to 600W to the graphics card. To put that in physiological terms: the RTX 3090 at full load draws more power than some cardiac surgical bypass circuits. This seemed like relevant context while routing the cables.


The New Anatomy in Situ

rear view with RM850x in position
rear view with RM850x in position
full internal view pre-closure

By early afternoon, the new anatomy was in place.

The RTX 3090 Gainward occupies the primary PCIe slot. “GEFORCE R—” is visible on the shroud — the rest obscured by the case edge, which seemed appropriately modest for a device about to receive considerably more attention than it had any right to expect. The PCIe power cables converge on it from the lower right like coronary vessels approaching a hypertrophied ventricle. The Corsair RM850x sits at the base of the chassis, label facing outward, entirely settled into its new home.

The frontal view shows the full operative result before closure: MSI cooling solution on the CPU (the pulmonary apparatus, in this extended metaphor — keeping everything cool, keeping everything running), the XPG motherboard visible behind the GPU, the cable routing through the chassis floor grating, the Thermaltake case providing structural containment for everything inside it.

It looked correct. It looked clean. The question — as it always is, at this point in any procedure — was whether it would work.

Post-Operative Course and Outcome Measures

The panel was replaced. The case was closed. The patient was connected to power and display and powered on.

POST-OP DAY 0: BIOS detected both components correctly. No POST errors. Windows loaded normally. nvidia-smi confirmed RTX 3090 24576 MB VRAM. Temperature at idle: 39°C core. Hemodynamically stable.

POST-OP DAY 0 (LATER): Stress testing with HWInfo64. GPU core peak 60.3°C. Hotspot 71.5°C. Voltages within normal range throughout. No instability events. The patient tolerated the workload without complications — which is, in both medicine and hardware, the outcome you were hoping for but never quite take for granted until you see it.

The benchmark results — the objective outcome measures of this intervention — were collected the following day using the same protocol applied to the pre-operative system:

MetricPre-op (RTX 3060)Post-op (RTX 3090)Delta
Mistral-Nemo gen speed30.1 tok/s51.5 tok/s+71%
Phi-4 gen speed27.4 tok/s48.7 tok/s+78%
LLM gen speed (long prompt)33–40 tok/s76–86 tok/s+114–126%
Flux image generation91.8 s/image44.1 s/image−52%
Whisper transcription10.1× realtime11.6× realtime+15%

The LLM inference improvement — particularly the generation speed on long prompts, which more than doubled — reflects the additional VRAM: with 24 GB, models that previously required partial CPU offloading now run entirely in GPU memory. The Flux result is the most immediately visible change in daily workflow: images that took a minute and a half now take forty-four seconds. Over a production session of ten images, that’s nearly eight minutes recovered per session.

Whisper improved modestly. Whisper large-v3 was already comfortable in 12 GB; the 3090 doesn’t resolve a bottleneck that didn’t exist.


Discharge Summary

The patient was discharged to home the same day, in good general condition, with improved functional capacity across all measured domains.

The RTX 3060 and Corsair 600W are awaiting evaluation for secondary use or resale — both functional, both simply superseded by the requirements of the case. In cardiac surgery, explanted tissues occasionally end up in tissue banks. In hardware, they end up on eBay.

The operative field — the floral tablecloth — survived the procedure without incident and has since returned to its primary function.

The surgeon went to bed at a reasonable hour, which is more than can be said for most OR days.

small “town band” made of poorly dressed boys plays using improvised instruments

Multi‑Agent AI in Healthcare

Posted on June 2, 2026August 16, 2026 by Michele Danilo Pierri

From passive Q&A to autonomous agents that debate, challenge, and assist clinicians.

Article authored by Michele D. Pierri, MD

Cardiac Surgeon & Medical Technology Researcher

Last updated: June 2026

Reading time: 15 minutes

Key takeaways

  • Agentic AI is not just a better chatbot. It is goal-driven, tool-using, stateful, and designed to work with clinicians.
  • Multi-agent / multi-persona deliberation can mitigate diagnostic cognitive biases (anchoring, premature closure, availability) by forcing structured counterarguments.
  • Real clinical value comes from grounding (EHR/FHIR + knowledge tools), traceability, and evaluation, not from eloquent text alone.
  • In healthcare, human-in-the-loop and safety-by-design are non-negotiable: agents should prepare, summarize, and request confirmation; they should never execute irreversible actions autonomously.
  • The same blueprint scales beyond diagnosis (med rec, trial matching, safety monitoring, patient journey).

1. Redefining AI in Medicine: What Is an Agentic System?

The first wave of AI in healthcare gave us chatbots. Helpful for fetching information, yes, but essentially reactive: you ask a question, you get an answer. Today, a different paradigm is taking shape, one that I find genuinely more interesting from a clinical standpoint: agentic systems, AI architectures that set goals, use tools, retain memory, and orchestrate multiple specialized components to complete complex, multi-step tasks with minimal human hand-holding.

An agentic system is more than a language model that answers prompts. It’s an architecture in which one or more autonomous AI agents perceive their environment, make decisions, call external tools (APIs, databases, calculators), and adapt their strategy to achieve a predefined clinical or operational goal. The distinction matters more than it might seem.

To avoid confusion, a brief vocabulary:

  • Agentic: goal + planning + tool use + state/memory.
  • Multi-agent: multiple specialized agents coordinated by an orchestrator/controller.
  • Multi-persona: a multi-agent pattern where each agent embodies a distinct cognitive/clinical style (e.g., pragmatist, scholar, devil’s advocate).

Contrast this with familiar alternatives:

  • Passive chatbot: responds only when addressed; no initiative, no state.
  • Deterministic automation: follows rigid rules (e.g., “IF creatinine > 1.2 THEN alert”) without contextual reasoning.
  • Single LLM pipeline: one model, one prompt, one answer. No memory, no interplay of perspectives.

Key characteristics of an agentic healthcare system:

  • Goal-oriented: Not “give me information about chest pain” but “determine the most likely causes of this patient’s chest pain, considering their history, and propose the next three diagnostic steps ranked by appropriateness.”
  • Tool use: The agent doesn’t just rely on its internal knowledge; it actively queries electronic health records (via FHIR APIs), drug databases, clinical guidelines, calculators (e.g., Wells score), and even PubMed.
  • Memory and state: It maintains a longitudinal view of a patient journey that may span months, remembering previous interactions and decisions.
  • Multi-agent orchestration: Specialized agents, each with a distinct role, persona, and tool set, collaborate, debate, and converge on a better outcome than any single model could achieve alone.
  • Human-in-the-loop by design: In clinical contexts, agents never execute irreversible decisions (prescriptions, procedures) autonomously; they prepare, summarize, and ask for confirmation.

This is not a distant vision. The building blocks are already here: large language models with function calling, vector databases for memory, and orchestration frameworks. The real question is how to assemble them for clinical value. And that, frankly, is where most implementations still fall short.


2. Under the Hood: The Building Blocks of a Clinical Agent

Before diving into the case study, let’s demystify the tech stack that makes a clinical agent work:

  • LLM as reasoning core: Models like GPT-4 or open-source alternatives (Meditron, Llama 3 fine-tuned on biomedical text) generate plans, synthesize information, and simulate different clinical perspectives.
  • Tool calling / function calling: The agent can execute API requests, retrieving lab results from a FHIR server, checking drug interactions with a medication knowledge base, or searching medical literature via PubMed.
  • Memory systems: Short-term memory (conversation history) and long-term memory (vector stores of patient summaries, clinical guidelines) allow the agent to accumulate context.
  • Orchestration logic: A “controller” agent manages the flow, parsing the user’s goal, dispatching sub-tasks to specialized agents, and synthesizing their outputs.
  • Safety guardrails: Output filters, clinical confidence thresholds, mandatory human validation steps, and “no-autonomy” policies prevent hazardous recommendations.

Minimum viable stack (practical blueprint)

  • An LLM with tool/function calling
  • A controller/orchestrator prompt (or a simple state machine)
  • Retrieval (RAG) over local guidelines + a curated knowledge base
  • Structured logging (inputs, tool calls, outputs, decisions)
  • Evaluation harness (see section 5) + red-team cases
  • UX/Workflow integration (EHR context, confirmation step, traceability)

For healthcare specifically, interoperability standards like HL7 FHIR are the plumbing that lets agents talk to real hospital systems. Even a prototype can use public FHIR test servers (e.g., the SMART Health IT sandbox) to simulate real-world data flows. In practice, getting this integration right is considerably harder than the demos suggest.


3. Case Study: A Multi-Persona Agent for Differential Diagnosis

To test these ideas, I built a rudimentary agentic system for differential diagnosis that simulates a team of experts with distinct cognitive styles. The premise is simple: a clinician enters a chief complaint and a few anamnestic clues; the system engages an internal panel of AI personas that reason, critique each other, and converge on a weighted list of hypotheses. A bit like a real case discussion, though with its own obvious limitations.action, healthcare data governance

Agent Learning Hero Section

Why this matters

Diagnostic errors often stem from cognitive biases: premature closure, anchoring, availability. A single AI can easily fall into the same traps. By introducing multiple deliberative agents, we can expose blind spots, consider counterfactuals, and broaden the differential. Whether this actually reduces diagnostic error rates at scale is still an open question, one that proper clinical validation would need to answer.

The Agent Trio

  1. The Clinical Pragmatist: the frontline expert. Thinks in terms of probabilities, common presentations, and “what you can’t miss.” References real-world prevalence, prioritizes immediate actionable steps, and avoids over-investigation.
  2. The Academic Scholar: the literature-driven mind. Retrieves the latest evidence, rare syndromes, recent cohort studies, and genomic associations. Prioritizes pathophysiological coherence and flags emerging disease phenotypes.
  3. The Provocateur (Devil’s Advocate): the critical foil. Generates counterfactual reasoning: “What if the main symptom is a red herring?” “What if two diseases are co-occurring?” “What if a lab result is falsely normal?” Challenges assumptions and forces alternative explanations.

How it works: a narrative walkthrough

Let’s trace a real interaction. The user input is a sparse clinical snapshot:

“72-year-old man, 3-month history of fatigue, unintentional weight loss of 8 kg, mild diffuse abdominal pain. No fever, no blood in stool. Former smoker.”

Step 1 – Data parsing and goal setting:

The orchestration layer structures the input into a problem representation and assigns the goal: Generate a ranked differential diagnosis, highlighting red flags and evidence gaps.

Step 2 – Independent reasoning:

Each agent receives the structured case and is instructed to produce:

  • A ranked list of 3-5 possible diagnoses
  • The clinical rationale
  • One key “what else?” counterpoint
Agent Learning Control Panel

Step 3 – Deliberation and synthesis:

The controller feeds these outputs into a moderator step, asking the agents to comment on each other’s contributions and adjust their confidence. The Provocateur forces the group to address the “colon cancer without bleeding” scenario; the Academic adds that the absence of jaundice doesn’t exclude a pancreatic body/tail tumor. The Pragmatist absorbs the suggestions and updates the list, now including TSH and a medication review as immediate low-cost steps.

Step 4 – Final output artifacts (what the clinician actually gets):

Agent Learning: top 3 diagnoses

Tools used in this prototype

The system leverages an LLM with function calling to:

  • Query PubMed for recent case reports (Academic agent)
  • Compute pre-test probabilities using simple epidemiological priors
  • Check drug databases for medications that cause weight loss/fatigue

Still rudimentary, it demonstrates a fundamental shift: the AI stops being a black-box oracle and becomes a transparent, multi-voiced assistant that explicitly questions its own reasoning. Whether that shift translates to measurable clinical benefit remains to be tested properly.


4. Beyond Diagnosis: Other Agentic Patterns in Healthcare

The multi-agent paradigm is not limited to differential diagnosis. The same architectural patterns apply across the entire patient journey. A few high-impact scenarios worth considering:

  • Patient Journey Concierge: a proactive agent that monitors appointment schedules, medication adherence, and patient-reported symptoms. It reaches out via WhatsApp, adjusts reminders based on mood and engagement, and escalates to a human when it detects worsening.
  • Safety Watcher for Inpatient Wards: an agent that streams vitals and labs in real time, correlates subtle trends (e.g., rising heart rate + falling urine output), and sends structured, actionable alerts to the responsible nurse, moving beyond noisy threshold alarms.
  • Trial Matching Navigator: an agent that nightly scans clinical documentation and matches patients to active trials, flagging eligible candidates and requesting missing genetic tests, thus accelerating research enrollment.
  • Medication Reconciliation Mediator: at admission and discharge, an agent reconciles home medications with hospital orders, flags dangerous interactions and omissions, and generates patient-friendly medication plans with follow-up reminders to prevent readmissions.

Each follows the same agentic recipe: clear goals, tool integration, memory, evaluation, and, in many cases, the interplay of multiple specialized agents. The execution details vary considerably, and so does the risk profile.


5. The Critical Layer: Safety, Governance, Evaluation, and the Human in the Loop

If you’re a developer or clinician reading this, the safety question is probably already forming: Who is liable when a multi-agent discussion misses a key diagnosis?

This is the most important design consideration. In any agentic healthcare system, four principles are non-negotiable:

  1. Human-in-the-loop for clinical decisions: The system provides decision support, not autonomous orders. The final differential, test ordering, and treatment remain firmly in the clinician’s hands.
  2. Explainability / traceability: Every recommendation must be traceable to the agent reasoning and to the tools used (inputs, tool calls, retrieved snippets, timestamps).
  3. Legal and UX framing: The output is “clinical decision support” or “evidence synthesis,” not a medical act. This distinction must be explicit in the interface and documentation.
  4. Data privacy and consent: Agents that access patient data must run inside secure, compliant environments. In prototyping, synthetic or de-identified data should be the default.

Evaluation (often missing, always required)

To move from a demo to something clinically credible, you need an evaluation layer. This is the part most teams skip, and the omission shows:

  • Case-based benchmarking: a curated set of vignettes (common + edge cases) with expected differentials and failure modes.
  • Error taxonomy: track where the agent fails (missed “can’t miss,” over-testing, hallucinated evidence, wrong prioritization).
  • Calibration: compare confidence vs correctness; force abstention or escalation below thresholds.
  • Human review workflow: structured clinician feedback loops (what changed, why, and whether the tool helped or harmed).

Our differential diagnosis agent, for example, wraps every final output in a disclaimer and always includes a confidence statement. If internal consensus drops below a threshold, the system explicitly suggests escalation to a human consultation rather than pushing a low-confidence guess. It’s a small design choice, but an important one.


6. From Insight to Action: A Starter Kit for Agentic Healthcare

To move from theory to practice, the community needs open, safe sandboxes. I propose a Medication Reconciliation Agent Starter Kit, a minimal but extensible agent that:

  • Connects to a public FHIR test server (e.g., HAPI FHIR) populated with synthetic patient data.
  • Ingests a mock admission note and home medication list.
  • Deploys two agents: a Pharma Checker that flags interactions/omissions and a Patient Communicator that generates a plain-language discharge medication plan.
  • Produces concrete artifacts: interaction flags with sources, a reconciled med list for clinician confirmation, and a patient-facing summary.
  • Operates with guardrails by default: no dosing changes suggested, mandatory clinician confirmation, and full trace logging.

This starter kit would give developers a concrete blueprint for building agentic healthcare tools that are safe by design, interoperable, and grounded in real-world workflows. The same principles can then be extended to more ambitious multi-persona diagnostic agents. Whether hospitals will actually adopt something like this depends on factors well beyond the technical, including procurement, liability, and clinician trust.


Conclusion: A Debate Club Inside Your Clinical Workstation

Agentic AI in healthcare is not about replacing clinical judgement. It’s about augmenting it with a tireless, transparent, and multi-perspective reasoning partner. Even a simple orchestration of three distinct agent personas can surface diagnoses that might otherwise be missed and, critically, explain why.

The next step is to harden these prototypes with real data integration, rigorous evaluation, and thoughtful clinician-in-the-loop interfaces. The open-source community and health-tech builders have an opportunity to create the building blocks of a new generation of clinical AI, one that debates, double-checks, and ultimately serves the patient.

If you’re building in this space: what agentic pattern would you add to this list?


See also on this site: OpenClaw and Cowork in Healthcare


a solemn council / meeting of physicians, gathered around a large wooden table, each doctor representing a different era of medical “seeing”.

From Hippocrates to Artificial Intelligence

Posted on May 27, 2026August 9, 2026 by Michele Danilo Pierri

I am a cardiac surgeon. My work is rooted in the operating room, in clinical decisions, and in the care of individual patients. At the same time, I have become deeply interested in medical data, statistics, coding, and artificial intelligence. My fascination with the history of medicine, however, has remained just as strong. At first glance, these interests may appear distant: ancient medical texts on one side, machine learning pipelines on the other. Yet the more I have studied the history of medicine, the more that separation has dissolved.

Medicine has always been technological.

Not only because it uses instruments, but because every era has developed new ways of seeing. The Hippocratic physician used observation as a technology of attention. Vesalius used dissection and illustration to reopen the human body to direct evidence. Harvey used experiment and quantitative reasoning to transform the heart from a symbolic organ into a pump. Morgagni connected symptoms to anatomical lesions. Virchow moved disease into the cell. Laennec made the body audible through the stethoscope. Lister and Koch made infection visible through the logic of germs. Modern clinicians use laboratory data, imaging, risk scores, electronic records, medical coding, and now artificial intelligence.

Each era had its own way of making the invisible visible.

This is why the great books of medicine are not simply historical monuments. They are records of changing perception. They show how physicians learned to look differently: at the patient, at the body, at disease, at evidence, and eventually at data.

Today, when we apply code, statistics, and machine learning to medicine, we are not abandoning the medical tradition. We are entering a new phase of it.

Thesis: the history of medicine can be read as a history of technologies of perception—instruments and methods that repeatedly change what counts as evidence, and therefore what can be seen.

In what follows, I sketch a compressed itinerary from Hippocratic bedside observation to digital medicine, ending with artificial intelligence as another instrument of visibility—powerful, but not self-justifying.

Medicine as a History of Seeing

Historical eraDominant instrument or methodWhat became visible
Hippocratic medicineBedside observationThe clinical course of disease
Classical and medieval medicineCompilation, commentary, classificationMedical knowledge as an organized tradition
Renaissance anatomyDissection and anatomical illustrationThe structure of the human body
Early modern physiologyExperiment and quantitative reasoningThe dynamic function of organs
Pathological anatomyAutopsy and clinicopathological correlationThe anatomical seat of disease
Clinical medicineBedside examination, auscultation, teaching hospitalThe patient as a clinical pattern
Cellular pathologyMicroscope and histologyDisease at the cellular level
Microbiology and antisepsisCulture, staining, germ theory, surgical hygieneInvisible infectious agents
Modern internal medicineLaboratory data, imaging, textbooks, evidenceThe measurable patient
Digital medicineElectronic health records, coding, statistics, AIPatterns, predictions, trajectories, and hidden phenotypes

To read the rest of this essay, keep three recurring moves in mind:

  • A new interface (a tool, a method, a representational technology).
  • A new object of knowledge (what becomes visible: lesion, cell, microbe, trajectory).
  • A new risk of reduction (what is lost when the new object becomes the whole story).

The Hippocratic Tradition: Seeing the Patient

The Hippocratic Oath is probably the most recognizable medical text in Western history. Its authorship and precise date remain uncertain; attributing it to Hippocrates as a single identifiable author is, at best, a useful convention. Still, its symbolic weight is considerable: it presents medicine not merely as a technical practice, but as a moral profession governed by obligations toward patients, teachers, colleagues, and the wider community of physicians. The Oath is commonly dated to the classical Greek period, though its exact origin continues to be debated. (einsteinmed.edu)

The Oath of Hippocrates

Yet the Oath alone does not capture Hippocratic medicine. The broader Hippocratic Corpus includes writings on prognosis, epidemics, environment, diet, and clinical observation. Illness becomes, in these texts, something that can be observed, described, compared, and followed over time. Medicine here begins to separate itself from magical or purely religious explanations of disease.

What preceded that separation is worth looking at directly: the Homeric poems already describe a working medical culture — wound care, drug knowledge, practitioners with recognised standing — with no theory of disease behind any of it.

This is the first great transformation of medical vision: the patient becomes a temporal phenomenon. Disease is not only a state; it is a course. The physician must watch, remember, compare, and anticipate.

In that sense, the Hippocratic physician already thinks in trajectories.

Celsus, Dioscorides, Galen, and Avicenna: Seeing Medicine as Organized Knowledge

Aulus Cornelius Celsus’ De Medicina, written in the first century CE, is one of the most important surviving Latin medical texts from antiquity. It covers general medicine, pharmacology, surgery, and bone disease, preserving one of the great organized accounts of ancient medical knowledge. Celsus is also associated with the classical signs of inflammation: redness, swelling, heat, and pain. (historyofinformation.com)

In De Materia Medica, Dioscorides organized medicinal substances derived from plants, minerals, and animals. For centuries, this kind of writing shaped the way physicians and healers thought about therapy: treatment was inseparable from careful observation of the natural world.

Galen then created one of the most powerful medical systems in history. His works integrated anatomy, physiology, humoral theory, therapeutics, and philosophy into a structure coherent enough to dominate European and Islamic medicine for centuries. Even when later physicians corrected him, they typically did so by first arguing with him. His importance lies not only in what he got right, but in the intellectual architecture he provided.

Avicenna’s Canon of Medicine, completed in the early eleventh century, became one of the most influential medical textbooks ever written. Through Latin translation it entered European medical education and remained authoritative for centuries. (biodiversitylibrary.org)

This era teaches a lesson that is easy to overlook. Before medicine can become experimental, it must become transmissible. Knowledge has to be collected, ordered, taught, and criticized.

In modern terms, this is the age of medical databases before databases existed.

Vesalius: Seeing the Body

The publication of Andreas Vesalius’ De humani corporis fabrica libri septem in 1543 marks one of the decisive moments in the history of medicine. The work was not simply an anatomical atlas. It changed the authority structure of medical knowledge. The National Library of Medicine presents Vesalius’ Fabrica as a landmark of historical anatomy, notable for its detailed anatomical woodcuts and its insistence on direct engagement with the dissected body. (nlm.nih.gov)

Andreas Vesalius' De humani corporis fabrica libri septem
http://www.metmuseum.org/Collections/search-the-collections/358129

For centuries, anatomy had been mediated through ancient authorities, Galen above all. Vesalius did not merely add new details. He placed the human body itself back at the center of medical truth.

The intellectual gesture was radical: when the text and the body disagree, the body must be observed again.

This is one of the deep roots of modern medicine. The physician is no longer only a reader of inherited knowledge. He becomes an observer, a dissector, a verifier. Vesalius transformed anatomy into a visual and empirical science.

In the history of medical seeing, this is the moment when the body becomes an object of direct evidence.

Harvey: Seeing Function

In 1628, William Harvey published Exercitatio Anatomica de Motu Cordis et Sanguinis in Animalibus, commonly known as De Motu Cordis. The work demonstrated the circulation of blood and the pumping function of the heart, and it is widely regarded as one of the foundational texts of modern physiology. (PMC)

Harvey’s importance is not only cardiovascular. It is methodological.

He did not simply describe the heart. He reasoned about it quantitatively. He asked how much blood the heart could eject per beat, whether the older Galenic model was physically plausible, and how venous valves behaved under pressure. He used observation, experiment, calculation, and mechanical reasoning together, in a way that still feels unmistakably modern.

For a cardiac surgeon, Harvey is not a remote historical figure. He belongs to the conceptual ancestry of hemodynamics, cardiac output, venous return, and circulatory physiology. I find it striking, when reviewing post-bypass hemodynamic data, that the fundamental framework we use is still recognizably Harveyan, nearly four centuries later.

Vesalius taught medicine to see structure. Harvey taught medicine to see function.

Boerhaave and the Teaching Hospital: Seeing the Clinical Pattern

Herman Boerhaave is often associated with the rise of clinical teaching in early eighteenth-century Leiden. The text in the original list, Observationes Medicae, is less representative of his historical role than Institutiones medicae and Aphorismi de cognoscendis et curandis morbis, published in 1708 and 1709 respectively. Britannica lists these among his principal works and emphasizes their wide use during and after his lifetime. (britannica.com)

Boerhaave’s significance lies in the transformation of medicine into a teachable clinical discipline. The patient was no longer only an individual case, nor only an illustration of a theoretical system. The patient became part of a reproducible educational method.

The hospital became a classroom. The bedside became a site of disciplined observation. The clinical case became a unit of knowledge.

This shift still shapes how medicine is practiced and taught today, even when we no longer trace it back to Leiden. Modern ward rounds, case presentations, morbidity and mortality conferences, and clinical reasoning exercises all preserve something of this tradition, often without anyone in the room being aware of it.

Morgagni: Seeing the Lesion

Giovanni Battista Morgagni’s De sedibus et causis morborum per anatomen indagatis, published in 1761, is one of the foundational works of pathological anatomy. Its central move was to correlate clinical histories with post-mortem anatomical findings. Disease was no longer only a general disturbance of the body. It had a seat. It could be localized. (sciencedirect.com)

This changed clinical reasoning profoundly.

Symptoms became clues pointing toward internal lesions. Autopsy became a method for verifying diagnosis. The clinic and the dissecting room became connected.

Modern imaging still operates inside this Morgagnian framework. CT, MRI, echocardiography, angiography, PET, and ultrasound all continue to ask a very old question in technologically new ways:

Where is the lesion?

Morgagni taught medicine to connect the story of the patient with the geography of the body.

Jenner: Seeing Prevention

Edward Jenner’s An Inquiry into the Causes and Effects of the Variolae Vaccinae, published in 1798, belongs to another great transformation: prevention becomes one of medicine’s most powerful instruments. Jenner’s work on cowpox and smallpox vaccination did not rest on modern immunology, which did not yet exist as a discipline, but it opened the path toward vaccination, public health, and the idea that disease could be prevented before it appeared. (resource.nlm.nih.gov)

This is a different kind of medical vision. The physician is no longer looking only at the sick body. He is looking at future disease.

Vaccination changes the temporal structure of medicine. The target is not only the present lesion, but the avoided event. A prevented death, a prevented epidemic, a prevented complication: these are invisible successes. And precisely because they are invisible, they tend to be undervalued, both clinically and politically.

Modern risk prediction, screening, population health, and artificial intelligence all inherit this preventive logic, whether or not we acknowledge it explicitly.

Laennec: Hearing the Invisible

René Laennec’s De l’auscultation médiate, published in 1819, introduced the stethoscope and transformed thoracic examination. The body became audible in a new way. Sounds from the chest could now be correlated with internal pathology, particularly diseases of the heart and lungs. (ajconline.org)

The stethoscope did not replace clinical judgment. It extended it. It created a new interface between the physician and the hidden body.

Auscultation also changed the ritual of the clinical encounter. The physician listened not only to the patient’s words, but to the patient’s organs. There is something worth preserving in that gesture, I think, even as we accumulate ever more sophisticated imaging data and remote monitoring systems.

In modern technological medicine, we tend to assume that instruments distance us from the patient. Laennec reminds us that instruments can also create new forms of intimacy.

Virchow: Seeing the Cell

Rudolf Virchow’s correct landmark text is Die Cellularpathologie in ihrer Begründung auf physiologische und pathologische Gewebelehre, published in 1858. Virchow’s concept of cellular pathology made the cell the fundamental unit of disease. (christies.com)

After Morgagni, disease had an organ.

After Virchow, disease had a cellular substrate.

This transition is more consequential than it might initially appear. The lesion was no longer only visible to the naked eye at autopsy. It could be microscopic, requiring instruments, staining, histological preparation, and trained interpretation. Medicine moved down in scale, and the tools had to follow.

This shift still underlies pathology, oncology, hematology, inflammatory disease, transplantation medicine, and much of contemporary biomedical reasoning. What Virchow established in the mid-nineteenth century, we are still elaborating and extending.

Virchow taught medicine that the most important lesion may be invisible until technology changes the scale of vision.

Semmelweis, Lister, and Koch: Seeing Infection

The nineteenth century transformed medicine’s understanding of infection. The process was neither smooth nor linear.

Ignaz Semmelweis’ work on puerperal fever showed that hand hygiene could dramatically reduce maternal mortality, though his ideas were resisted during his lifetime with a stubbornness that remains, in retrospect, difficult to fully account for. Joseph Lister’s On the Antiseptic Principle in the Practice of Surgery, published in 1867, applied antiseptic principles to surgical practice and helped transform surgery from a frequently lethal intervention into a safer therapeutic discipline. (PMC)

Ignaz Semmelweis' work on puerperal fever

Robert Koch’s work on tuberculosis, presented in 1882, made the tubercle bacillus visible and helped establish a new model of infectious causation. The consequences for microbiology, public health, and the etiological understanding of disease were far-reaching. (germanhistory-intersections.org)

Before germ theory, infection was often explained through miasma, constitutional weakness, bad air, or poorly defined contamination. After microbiology, disease could be linked to specific organisms, specific routes of transmission, and specific preventive measures.

For any surgeon, this is not simply history. It is the foundation of the operating room. Every sterile field, every preoperative antibiotic, every isolation protocol traces back, however indirectly, to this era.

Claude Bernard: Seeing Experimentally

Claude Bernard’s Introduction à l’étude de la médecine expérimentale, published in 1865, is one of the great methodological texts of modern medicine. Bernard helped define medicine as an experimental science, not merely an accumulation of clinical impressions. (sciencedirect.com)

His contribution is epistemological.

Observation is necessary, but not sufficient. The physician-scientist must formulate hypotheses, design experiments, control conditions, interpret results, and remain alert to the difference between association and causation. These are not trivial demands; they become harder, not easier, when the datasets grow large.

This remains directly relevant to modern medical data science. A machine learning model can identify patterns with high statistical confidence. But medicine still needs to ask whether those patterns are meaningful, causal, generalizable, and clinically useful. These are not algorithmic questions. They require human judgment.

Bernard reminds us that better data do not automatically produce better reasoning.

Osler, Harrison, and the Modern Clinical Textbook

William Osler’s The Principles and Practice of Medicine, first published in 1892, and Harrison’s Principles of Internal Medicine, first published in 1950, belong to a different category from Vesalius, Harvey, Morgagni, or Virchow. They are not books of a single discovery. They are architectures of clinical knowledge. (archive.org)

Osler represents the humanistic and bedside tradition of modern clinical medicine. Harrison represents the increasingly pathophysiological, laboratory-based, and systematic organization of internal medicine. Later editions of Harrison explicitly reflect the transformation of medicine through molecular genetics, imaging, robotics, bioinformatics, and information technology. (accessmedicine.mhmedical.com)

William Osler's The Principles and Practice of Medicine

These books show that modern medicine is not only a collection of discoveries. It is also a teaching system. A good textbook does not merely contain knowledge. It trains a way of thinking.

Mukherjee: Seeing Disease as Biography

Siddhartha Mukherjee’s The Emperor of All Maladies: A Biography of Cancer, published in 2010, is not a foundational scientific treatise in the same sense as Harvey’s or Virchow’s work. It is something different: a modern narrative history of cancer as a biological, clinical, scientific, social, and human phenomenon. It won the Pulitzer Prize for General Nonfiction. (Wikipedia)

Its place in a list of medical classics is defensible only if we understand the list broadly, and I think we should.

Modern medicine does not only need discoveries. It also needs memory, and narratives capable of connecting laboratory science, clinical practice, patient suffering, public policy, and cultural imagination. Cancer is not only a cellular disease. It is also a historical experience, a therapeutic battlefield, a social fear, and, for many patients I have known, an entirely personal catastrophe.

Mukherjee’s book reminds us that medicine must see not only mechanisms, but lives.

From Medical Texts to Medical Data

The history of medicine can be read as a history of changing visibility.

The Hippocratic physician saw the course of disease. Vesalius saw the anatomical body. Harvey saw circulation. Morgagni saw the lesion. Virchow saw the cell. Koch saw the microbe. Osler saw the clinical patient. Harrison saw the patient through pathophysiology, laboratory medicine, and organized internal medicine.

Today, digital medicine asks us to see something else: patterns distributed across data.

Electronic health records, ICD codes, laboratory time series, imaging datasets, operative notes, discharge summaries, wearable sensors, and genomic information are producing a new kind of medical object. Not simply the patient at the bedside, not simply the organ, not simply the cell, but the patient as a trajectory through complex systems of data.

This is where coding and artificial intelligence enter the story. Medical coding is not just administrative work. It is one of the ways medicine translates clinical reality into structured information. Machine learning, at its best, is a method for detecting patterns that are too complex, too distributed, or too subtle for ordinary clinical perception.

In practice, this “trajectory view” shows up in concrete clinical tasks. We use models (formal or informal) to recognize syndromes and phenotypes that are not single lesions but composite patterns—think of heterogeneous entities such as sepsis, ARDS, or HFpEF. We try to anticipate events before they declare themselves (AKI, decompensation, readmission risk). And we increasingly use NLP to extract latent structure from free text—operative notes, discharge summaries, and the narrative fragments that still carry clinical meaning after coding has done its work.

But the lesson of history is clear: every new way of seeing also creates new risks. Texts can become dogma. Anatomy can reduce the patient to a body. Pathology can reduce disease to a lesion. Laboratory medicine can reduce illness to numbers. Artificial intelligence can reduce clinical reality to patterns without meaning.

The task is not to reject technology. The task is to keep technology inside medicine.

Conclusion: Artificial Intelligence as Another Chapter in an Old Story

The future of medicine will not replace the history of medicine. It will extend it.

Artificial intelligence is not the opposite of clinical tradition. It is one more instrument in the long human effort to see disease more clearly, act earlier, and understand the patient more completely. Whether it will fulfill that potential remains, as of now, genuinely uncertain.

The great books of medicine matter because they remind us that medicine has never been static. It has always changed when physicians found new ways to observe, represent, measure, classify, and interpret disease.

From the Oath to the anatomical atlas, from the autopsy table to the microscope, from the stethoscope to the laboratory, from the textbook to the electronic health record, medicine has always been a dialogue between human judgment and technical mediation.

The challenge today is the same as it was in every previous era: not simply to see more, but to see better. And above all, to remember that behind every new instrument of vision there remains the same object of medicine: the patient.


Key takeaways

  • Medicine’s history can be read as a sequence of technologies of perception: new instruments and methods that make new aspects of disease visible.
  • These shifts repeatedly reorganize medical authority: from inherited texts to direct observation, from anatomy to experiment, from lesions to cells, from germs to statistics.
  • Digital medicine extends this trajectory by treating the patient as a data-rich path through time (EHRs, codes, labs, imaging, text), not only as a bedside encounter.
  • AI can be understood as a new perceptual instrument for medicine: powerful at detecting distributed patterns, but always in need of clinical interpretation and ethical constraint.
A doctor reads a discharge letter to a patient.

Natural Language Processing for Medical Report Analysis – Part 3

Posted on May 25, 2026August 16, 2026 by Michele Danilo Pierri

Machine Learning and Large Language Models for Clinical Text Analysis


Article authored by Michele D. Pierri, MD

Cardiac Surgeon & Medical Technology Researcher

Last updated: May 2025

Reading time: 30 minutes


Abstract

The first two articles in this series traced a progression from rule-based pattern matching (regular expressions) to linguistic analysis (spaCy/scispaCy) for extracting structured information from cardiac surgery discharge summaries. Both approaches, despite their methodological differences, share one fundamental characteristic: they are designed by humans who encode explicit knowledge about language structure into algorithms. This third and final article introduces a qualitatively different paradigm, built around systems that learn representations of medical language from data. We cover classical machine learning classifiers for complication detection using TF-IDF feature engineering, then advance to Large Language Models (LLMs) and their application to clinical text through prompt engineering and the Anthropic Claude API. The article closes with a complete hybrid NLP pipeline integrating all three approaches, followed by an essential discussion of validation methodology, GDPR/HIPAA compliance, and de-identification techniques required for any production medical NLP system.

Article Series:

  1. Introduction to NLP and Regex for Medical Reports
  2. Advanced Linguistic Analysis with spaCy and scispaCy
  3. Machine Learning and Large Language Models for Clinical Text Analysis (this article)

1. The NLP Paradigm Progression: A Unified View

Before introducing the techniques covered here, it is worth consolidating the conceptual framework underlying the entire series. A question naturally arises: are regex, spaCy, and LLMs all really “NLP”? The answer is yes. All three are approaches to Natural Language Processing, the field concerned with enabling computers to understand and generate human language. What distinguishes them is the level at which language is modeled and the source of the knowledge they encode.

Regex (Article 1) encodes linguistic knowledge as explicit character patterns written by a human programmer. The system has no model of language structure; it matches sequences of characters. Knowledge source: human expert rules.

spaCy/scispaCy (Article 2) encodes linguistic knowledge as statistical models trained on annotated corpora. The system has explicit representations of tokens, parts of speech, syntactic structure, and named entities. Knowledge source: supervised learning from human-annotated examples.

Classical ML (this article, Section 2) encodes linguistic knowledge as learned associations between numerical text representations (TF-IDF vectors) and output labels. The system has no explicit model of language structure, but learns statistical correlations between word co-occurrences and outcomes. Knowledge source: supervised learning from labeled examples, with human-designed features.

Large Language Models (this article, Section 3) encode linguistic knowledge as billions of learned parameters representing the statistical structure of language at every level: from character sequences to semantic relationships to pragmatic context. Knowledge source: self-supervised pre-training on massive text corpora, capturing language in an emergent, distributed representation.

The progression is not merely technical. It reflects increasing depth of language understanding, at the cost of increasing computational requirements, reduced interpretability, and (critically for medical applications) greater difficulty in validation and regulatory compliance. Each level adds capabilities the previous cannot provide, and each retains appropriate use cases where simpler approaches are preferable.


2. Classical Machine Learning for Clinical Text Classification

2.1 Why Classical ML Before LLMs?

Large Language Models achieve impressive results on many clinical NLP tasks. Yet classical machine learning (logistic regression, random forests, support vector machines) remains clinically relevant for several reasons:

  • Interpretability: Feature importance scores from a logistic regression are auditable by clinicians; a transformer’s attention weights are not
  • Data efficiency: A well-engineered classifier can be trained on hundreds of examples; LLMs require prompt engineering but cannot be fine-tuned without thousands
  • Regulatory transparency: Many healthcare jurisdictions require explainable AI for clinical decision support; “because the LLM said so” is not an acceptable audit trail
  • Computational cost: Classical classifiers run on a laptop CPU; LLMs require GPU inference or API calls
  • Speed: TF-IDF + logistic regression classifies a document in microseconds

For complication detection from discharge summaries (a well-defined binary classification task), a classical ML approach is not only adequate but often preferable.

2.2 Text Representation: TF-IDF Vectorization

The fundamental challenge in applying machine learning to text is representation: algorithms require numerical inputs, not strings. Term Frequency–Inverse Document Frequency (TF-IDF) is the classical solution. It represents each document as a vector where each dimension corresponds to a vocabulary term, and the value encodes both how frequently the term appears in this document (TF) and how discriminative it is across the corpus (IDF).

TF (Term Frequency): How often does term t appear in document d?

\text{TF}(t, d) = \frac{\text{count of } t \text{ in } d}{\text{total terms in } d}

IDF (Inverse Document Frequency): How rare is term t across all documents?

\text{IDF}(t) = \log\frac{N}{1 + |\{d : t \in d\}|}

TF-IDF: The product emphasizes terms that are frequent in a specific document but rare overall, precisely the discriminative terms useful for classification.

In clinical terms: “atrial fibrillation” appearing in a discharge summary has high TF-IDF if it appears frequently in that document but rarely across the whole corpus, making it a strong signal for arrhythmia-related classification tasks.

Expected output:

Training Complete:
  Documents: 14
  Positive cases: 8
  Cross-validation AUC: 0.875 ± 0.112

Classification Result:
  Prediction: COMPLICATED
  Probability (complication): 73.2%
  Confidence: Medium

  Top contributing terms:
    'atrial fibrillation': +0.4821 → complicated
    'cardioverted': +0.3914 → complicated
    'complications': -0.2143 → uncomplicated
    'without complications': -0.1987 → uncomplicated
    'amiodarone': +0.1654 → complicated

  Terms most associated with complications:
    'infection': +0.8234
    'atrial fibrillation': +0.4821
    'reintubation': +0.4510
    'dehiscence': +0.4201
    'stroke': +0.3987

The model correctly identifies “atrial fibrillation” and “cardioverted” as complication indicators. The “without complications” bigram moderates the prediction toward Medium confidence. This is actually a clinically reasonable behavior, reflecting the nuanced nature of the case; a binary label would not capture such gradient.


3. Large Language Models for Clinical Text Analysis

3.1 What LLMs Can Do That Classical NLP Cannot

Large Language Models represent a paradigm shift in NLP. Rather than learning task-specific mappings from labeled examples, LLMs are pre-trained on vast text corpora to predict the next token in a sequence (a self-supervised objective that forces the model to internalize grammar, semantics, factual knowledge, and reasoning patterns simultaneously). The result is a general-purpose linguistic intelligence that can be directed toward specific tasks through natural language instructions, rather than task-specific training.

For clinical text analysis, LLMs offer capabilities that are genuinely novel:

  • Zero-shot extraction: Extract arbitrarily complex structured data from text without labeled examples, simply by describing what you want in the prompt
  • Semantic equivalence recognition: Understand that “triple vessel disease,” “3VD,” and “severe multivessel CAD” refer to the same condition without explicit synonym lists
  • Clinical reasoning: Assess whether documented management aligns with published guidelines (ACC/AHA, ESC), something requiring medical knowledge, not just pattern recognition
  • Contextual inference: Determine that a patient with LVEF 35% on admission and 40% on discharge had cardiac function improvement, even if the word “improvement” never appears
  • Uncertainty quantification in narrative form: Recognize “cannot exclude” and “suspicious for” as expressions of diagnostic uncertainty, rather than confirmed findings

3.2 The Anthropic Claude API for Medical NLP

The Claude API provides programmatic access to Anthropic’s Claude models. For medical NLP applications, Claude offers several relevant strengths: strong performance on clinical reasoning benchmarks, support for long-context documents (up to 200K tokens in Claude 3 models), and structured output generation.

3.3 Prompt Engineering for Medical Text

Effective LLM performance on clinical tasks depends critically on prompt design. Several principles, validated empirically in medical NLP research, apply directly to our use case.

Role assignment: Giving the model a specific clinical identity (“You are a cardiac surgery quality improvement specialist”) activates domain-specific knowledge patterns and reduces generic responses. In published evaluations, appropriate role prompting can yield meaningful gains on some medical QA tasks, though the magnitude depends strongly on the dataset and evaluation setup.

Schema specification: Providing the exact JSON schema in the prompt dramatically improves structured output reliability. Rather than asking “extract the medications,” defining the exact field names, types, and constraints guides the model toward consistent output.

Constraint encoding: Clinical validation rules embedded in the prompt (“LVEF values should be between 15 and 80”) help the model self-correct obvious errors and flag unusual values.

Uncertainty handling: Explicitly instructing the model to use null for absent data rather than inferring or hallucinating prevents a major failure mode in medical LLM applications. The instruction is simple: “Never infer or hallucinate values not explicitly stated.” In practice, even this directive is not always followed perfectly, which underscores the need for post-hoc validators alongside any clinical prompt.

Disclaimer integration: For guideline comparison and clinical assessment tasks, embedding disclaimers directly in the prompt ensures they appear in every output, supporting appropriate use of the system.


4. Privacy, De-identification, and GDPR/HIPAA Compliance

4.1 The Regulatory Landscape for Medical NLP

Any NLP system processing clinical documents operates within a strict regulatory framework. In Europe, the General Data Protection Regulation (GDPR) classifies health data as a “special category” requiring explicit consent and heightened protection standards. In the United States, HIPAA’s Privacy and Security Rules govern Protected Health Information (PHI) handling.

For medical NLP systems, the critical requirement is de-identification: removing or transforming all information that could identify an individual patient before any processing that extends beyond direct clinical care. HIPAA defines 18 categories of PHI that must be addressed:

  • Direct identifiers: names, geographic identifiers smaller than state, dates (except year), telephone numbers, MRN, SSN, email addresses
  • Quasi-identifiers: age (if over 89), rare medical conditions in combination with other data

4.2 De-identification Pipeline

import re
import hashlib
from typing import Dict, List, Tuple
from dataclasses import dataclass

@dataclass
class DeidentificationResult:
    deidentified_text: str
    phi_found: List[Dict]
    replacement_map: Dict[str, str]

class MedicalTextDeidentifier:
    PHI_PATTERNS = {
        'patient_name': (
            r'\\b(?:Patient|Name):\\s*([A-Z][a-z]+[A-Z][a-z]+(?:\\s[A-Z][a-z]+)?)\\b',
            'FULL_NAME'
        ),
        'date_full': (
            r'\\b(\\d{1,2}[/\\-]\\d{1,2}[/\\-]\\d{4})\\b',
            'DATE'
        ),
        'mrn': (
            r'\\b(?:MRN|Medical Record):\\s*(\\d{6,10})\\b',
            'MRN'
        ),
        'phone': (
            r'\\b(\\+?[\\d\\s\\-\\(\\)]{10,15})\\b',
            'PHONE'
        ),
        'email': (
            r'\\b([a-zA-Z0-9._%+\\-]+@[a-zA-Z0-9.\\-]+\\.[a-zA-Z]{2,})\\b',
            'EMAIL'
        ),
    }

    PSEUDONYM_POOLS = {
        'FULL_NAME': ['James Anderson', 'Robert Williams', 'Maria Garcia',
                      'David Johnson', 'Emma Thompson', 'Michael Brown'],
        'PHYSICIAN_NAME': ['Dr. A. Smith', 'Dr. B. Jones', 'Dr. C. Davis'],
        'DATE': None,
        'MRN': None,
    }

    def __init__(self, replacement_strategy='pseudonymize', consistent_replacement=True):
        self.replacement_strategy = replacement_strategy
        self.consistent_replacement = consistent_replacement
        self._replacement_cache: Dict[str, str] = {}
        self._pseudonym_counters: Dict[str, int] = {}

    def deidentify(self, text: str) -> DeidentificationResult:
        self._replacement_cache.clear()
        self._pseudonym_counters.clear()
        deidentified = text
        phi_found = []
        replacement_map = {}
        for phi_category, (pattern, phi_type) in self.PHI_PATTERNS.items():
            matches = list(re.finditer(pattern, deidentified))
            for match in reversed(matches):
                original = match.group(1) if match.lastindex else match.group(0)
                pseudonym = self._get_pseudonym(phi_type, original)
                phi_found.append({
                    'category': phi_category,
                    'phi_type': phi_type,
                    'start': match.start(),
                    'end': match.end(),
                    'replacement': pseudonym
                })
                full_match = match.group(0)
                replaced_match = full_match.replace(original, pseudonym)
                deidentified = deidentified[:match.start()] + replaced_match + deidentified[match.end():]
                replacement_map[pseudonym] = original
        return DeidentificationResult(
            deidentified_text=deidentified,
            phi_found=phi_found,
            replacement_map=replacement_map
        )

5. The Complete Hybrid NLP Pipeline

5.1 Integrating All Three Paradigms

The most robust clinical NLP systems do not rely on a single paradigm but orchestrate multiple approaches, each contributing where it performs best. The HybridMedicalNLPPipeline class integrates all techniques from this series.

Pipeline architecture:

  • De-identification (if enabled)
  • Layer 1: Regex — fast structured extraction of dates, values, medications (high precision for formatted fields)
  • Layer 2: scispaCy NER — biomedical entity recognition, abbreviation detection, UMLS linking
  • Layer 3: LLM (optional, highest cost) — complex inference, guideline comparison, structured extraction from free text
  • Reconciliation — merge results, resolve conflicts, validate against clinical constraints
  • Structured Output (JSON)

The pipeline is configurable: the LLM layer can be disabled for cost-sensitive batch processing, falling back to regex + scispaCy.

from dataclasses import dataclass, field
from typing import Optional
import time

@dataclass
class HybridPipelineResult:
    document_id: str
    deidentified: bool
    regex_data: Dict = field(default_factory=dict)
    nlp_entities: Optional[Any] = None
    llm_structured: Optional[Dict] = None
    llm_complications: Optional[Dict] = None
    llm_guidelines: Optional[Dict] = None
    reconciled: Dict = field(default_factory=dict)
    conflicts: List[str] = field(default_factory=list)
    processing_time_ms: float = 0.0
    errors: List[str] = field(default_factory=list)

class HybridMedicalNLPPipeline:
    def __init__(self, enable_deidentification=True, enable_nlp=True,
                 enable_llm=True, enable_guideline_check=False):
        self.enable_deidentification = enable_deidentification
        self.enable_nlp = enable_nlp
        self.enable_llm = enable_llm
        self.enable_guideline_check = enable_guideline_check
        self._deidentifier = None
        self._nlp_extractor = None
        self._llm_analyzer = None

    def process_document(self, text, document_id="unknown", run_guideline_check=False):
        start_time = time.time()
        result = HybridPipelineResult(document_id=document_id,
                                      deidentified=self.enable_deidentification)
        working_text = text

        if self.enable_deidentification:
            try:
                deid_result = self._get_deidentifier().deidentify(text)
                working_text = deid_result.deidentified_text
                result.regex_data['phi_removed'] = len(deid_result.phi_found)
            except Exception as e:
                result.errors.append(f"De-identification error: {e}")

        if self.enable_nlp:
            try:
                nlp_extractor = self._get_nlp_extractor()
                result.nlp_entities = nlp_extractor.analyze(working_text)
            except Exception as e:
                result.errors.append(f"NLP extraction error: {e}")

        if self.enable_llm:
            try:
                llm = self._get_llm_analyzer()
                llm_result = llm.extract_structured_data(working_text)
                result.llm_structured = llm_result.structured_data
                comp_result = llm.analyze_complications_severity(working_text)
                result.llm_complications = comp_result.structured_data
                if run_guideline_check or self.enable_guideline_check:
                    guide_result = llm.compare_to_guidelines(working_text)
                    result.llm_guidelines = guide_result.structured_data
            except Exception as e:
                result.errors.append(f"LLM extraction error: {e}")

        result.reconciled = self._reconcile_outputs(result)
        result.processing_time_ms = (time.time() - start_time) * 1000
        return result

6. Validation and Performance Metrics

6.1 Why Validation Is Non-Negotiable in Medical NLP

A classifier that achieves 95% accuracy on its training set but 70% on real-world data causes harm. In clinical applications, the stakes of this performance gap are not academic: incorrect complication detection, missed drug interactions, or erroneous guideline assessments directly affect patient management. Medical NLP validation must therefore be:

External: Evaluated on documents from a different time period or institution than the training data, testing true generalizability rather than memorization.

Task-specific: Aggregate accuracy conceals clinically important failures. A system might achieve 95% overall accuracy while failing on the 5% of cases involving rare but severe complications, exactly the cases where automated support is most valuable.

Entity-level: For NER tasks, evaluation should be at the entity span level (exact match of start position, end position, and label), not at the document level.

The standard metrics are:

\text{Precision} = \frac{TP}{TP + FP}

\text{Recall} = \frac{TP}{TP + FN}

\text{F1} = 2 \times \frac{\text{Precision} \times \text{Recall}}{\text{Precision} + \text{Recall}}

For binary classification with class imbalance (rare complications), the Area Under the Precision-Recall Curve (AUPRC) is more informative than ROC-AUC.

6.2 LLM Evaluation: Beyond Numeric Metrics

Evaluating LLM extraction quality requires additional considerations compared to classical classifiers. When a logistic regression misclassifies a case, the error is binary (correct or incorrect). When an LLM extracts the wrong medication dose, the error exists on a spectrum from trivial (50mg vs. 50.0mg) to critical (amiodarone 200mg vs. 2000mg). This asymmetry in error severity is something aggregate metrics simply fail to capture, and it matters enormously in cardiac surgery contexts.

For structured data extraction tasks, the recommended evaluation framework includes:

  • Field-level accuracy: Proportion of extracted fields exactly matching the gold standard annotation
  • Semantic equivalence: Cases where the LLM extracts a semantically equivalent but differently worded value (e.g., “twice daily” vs. “BID”) should be scored as correct
  • Hallucination rate: Proportion of extracted values not present in the source document, the most dangerous failure mode in clinical applications
  • Abstention quality: Does the model correctly return null for absent fields rather than inferring or fabricating values?

7. Production Considerations and Deployment

7.1 Architecture Considerations for Clinical Settings

Deploying any of the tools described in this series in a clinical environment requires addressing infrastructure considerations that go well beyond code quality.

Data residency: GDPR Article 44-49 restricts transfer of health data outside the EEA. If using cloud LLM APIs (including the Anthropic API), ensure contractual data processing agreements are in place and verify that data is not used for model training. For maximum compliance, consider on-premises deployment of open-source models (Llama 3, Mistral Medical).

Audit logging: Every document processed, every extraction performed, and every human validation decision must be logged with timestamps and user identifiers. This is both a regulatory requirement and a prerequisite for system improvement.

Human-in-the-loop validation: No automated NLP system should modify clinical records without human review. The appropriate architecture positions NLP output as proposed values subject to clinician validation, a co-pilot rather than an autopilot.

Model versioning and drift monitoring: Clinical language evolves. New procedures, drugs, and documentation conventions emerge continuously. Regular re-evaluation of system performance on recent documents is essential to detect performance degradation before it affects clinical use.

Fail-safe design: System failures must degrade gracefully. If the LLM API is unavailable, the system should fall back to regex + scispaCy extraction rather than blocking clinical workflow.

7.2 Local LLMs: On-Premises Deployment as a Privacy Solution

A natural and clinically important question arises from the privacy constraints discussed in Section 4: if sending clinical documents to a cloud API introduces GDPR/HIPAA compliance complexity, could deploying an LLM on an institutional server (entirely within the hospital’s own infrastructure) resolve these issues?

The answer is yes, in large measure, but with important technical and regulatory caveats that must be understood before committing to this architecture.

The Privacy Advantage of On-Premises LLMs

When an LLM runs on a server physically located within the hospital’s own data center or private cloud, clinical documents never leave the institutional perimeter. This eliminates the primary privacy risks associated with cloud API usage:

  • No data transfer to third-party processors, removing GDPR Article 44-49 obligations regarding cross-border transfers
  • No contractual Data Processing Agreement required with an external vendor
  • No risk of patient data being used for model training by a third party
  • Full institutional control over data retention, access logging, and deletion policies
  • Alignment with the data residency requirements increasingly mandated by national healthcare regulations (e.g., Italy’s Codice Privacy and AGENAS guidelines on health data sovereignty)

In practice, this means that on-premises LLM deployment can make de-identification before inference optional rather than mandatory, since the data never leaves the environment where it is already authorized to reside. This is a significant operational simplification, as de-identification pipelines (Section 4) introduce their own risks of information loss and require ongoing validation.

Available Open-Source Models for Clinical Use

The open-source LLM ecosystem has matured substantially and now includes models specifically oriented toward biomedical and clinical applications.

General-purpose models suitable for clinical NLP:

  • Llama 3.1 / Llama 3.3 (Meta AI): state-of-the-art open-source models with strong reasoning capabilities; the 70B parameter version approaches GPT-4 performance on many clinical reasoning benchmarks
  • Mistral 7B / Mixtral 8x7B (Mistral AI): excellent performance-to-compute ratio; the Mixtral mixture-of-experts architecture delivers strong results with lower inference cost than comparably-performing dense models

Biomedical fine-tuned models:

  • BioMistral (Labrak et al., 2024): Mistral 7B fine-tuned on PubMed Central and MIMIC-III clinical notes; demonstrates improved performance on biomedical NER, relation extraction, and clinical question answering compared to the base model
  • Meditron (EPFL, Chen et al., 2023): Llama 2 fine-tuned on a curated corpus of PubMed abstracts, medical textbooks, and clinical practice guidelines; specifically designed for guideline-adherent clinical reasoning
  • ClinicalCamel: fine-tuned on clinical conversation datasets; optimized for patient-provider dialogue rather than document processing
  • OpenBioLLM (Saama AI Research): fine-tuned on diverse biomedical datasets with strong performance on clinical entity extraction tasks

For the specific use case of this series (extraction from cardiac surgery discharge summaries), BioMistral or a Llama 3.1 70B base model with a well-engineered system prompt are the most appropriate starting points as of early 2026. The former benefits from clinical domain adaptation; the latter offers superior general reasoning for complex inference tasks such as guideline comparison. That said, the right choice ultimately depends on the institution’s GPU resources and the acceptable latency threshold.

Infrastructure Requirements

On-premises LLM deployment is not without cost. The computational requirements are substantial:

Model sizeGPU VRAM requiredSuitable hardwareApproximate throughput
7B parameters (fp16)~14 GBSingle NVIDIA A100 40GB~50 documents/min
13B parameters (fp16)~26 GBSingle A100 80GB~30 documents/min
70B parameters (fp16)~140 GB2x A100 80GB~8 documents/min
70B parameters (4-bit quantized)~35 GBSingle A100 80GB~20 documents/min

Note: The figures above are indicative only; real throughput depends on document length, batching, quantization, context window, and the serving stack (e.g., vLLM vs. llama.cpp) as well as the specific GPU model and configuration.

For most hospital settings, 4-bit quantization (using libraries such as bitsandbytes or llama.cpp) offers a practical compromise: a 70B model quantized to 4 bits fits on a single high-end GPU with performance degradation of approximately 2-5% on clinical benchmarks relative to the full-precision version.

The deployment stack for an institutional LLM server typically consists of the model weights (downloaded once from Hugging Face or a private registry), an inference server (vLLM or Ollama are the current standards for production throughput), and an API layer that exposes an OpenAI-compatible endpoint, allowing existing code written against the Anthropic or OpenAI API to switch to the local model by changing a single endpoint URL.

# Switching from Claude API to a local LLM server requires
# minimal code changes when the local server exposes an
# OpenAI-compatible endpoint (as vLLM and Ollama do)

# Original: Anthropic Claude API
# client = anthropic.Anthropic()
# response = client.messages.create(model="claude-opus-4-6", ...)

# Local LLM server (vLLM serving BioMistral or Llama 3)
from openai import OpenAI

local_client = OpenAI(
    base_url="<http://your-hospital-llm-server:8000/v1>",
    api_key="not-required-for-local"
)

response = local_client.chat.completions.create(
    model="BioMistral-7B",
    messages=[
        {"role": "system", "content": system_prompt},
        {"role": "user", "content": user_message}
    ],
    temperature=0.0,
    max_tokens=2048
)

extracted_text = response.choices[0].message.content

Important Caveats: Privacy Is Necessary But Not Sufficient

On-premises deployment resolves the data transfer dimension of compliance, but several obligations remain regardless of where the model runs.

Validation requirements are unchanged. A local LLM requires the same rigorous clinical validation as a cloud model. The fact that data stays on-site does not make the model’s outputs more reliable. Hallucination rates, extraction accuracy, and complication detection sensitivity must be measured on institution-specific test sets before clinical use.

Audit and logging obligations persist. GDPR and national healthcare regulations require that every automated processing of health data be documented, with identifiable records of who authorized the processing, when it occurred, and what decisions were made based on it. This applies equally to on-premises systems.

Security of the server itself. “On-premises” and “secure” are not synonyms. The LLM inference server must be integrated into the hospital’s information security framework: network segmentation, access controls, intrusion detection, and patch management. A poorly secured on-premises GPU server may represent a greater risk than a well-managed cloud API.

Performance gap relative to frontier models. As of early 2026, the best open-source models (Llama 3.1 70B, Mixtral 8x7B) remain below GPT-4 and Claude Opus on complex clinical reasoning tasks, particularly guideline comparison and multi-step inference. For straightforward extraction tasks (Section 3.2), the gap is small and often acceptable. For tasks requiring nuanced clinical judgment, the performance difference should be empirically measured on the target task before choosing local deployment over a cloud API with appropriate safeguards.

The privacy-performance tradeoff is real and context-dependent. The optimal architecture depends on the specific task, the sensitivity of the documents, the institutional IT capabilities, and the acceptable performance floor. A pragmatic approach adopted by several academic medical centers is a tiered system: de-identified documents processed via cloud API for maximum performance; identified documents processed by a local model for maximum privacy; with de-identification quality itself validated to determine when the cloud tier is safe to use.


8. Conclusion: The Complete NLP Landscape for Medical Text

This three-part series has traced a complete progression through the methodological landscape of clinical NLP, from character-level pattern matching to AI-powered clinical reasoning. The progression is not merely technical. It reflects a deepening of what “understanding language” means computationally.

Regular expressions (Article 1) encode understanding as explicit human-designed rules. They are precise, fast, explainable, and appropriate for structured fields in consistent formats. They cannot generalize beyond their patterns and require constant maintenance as documentation conventions change.

spaCy and scispaCy (Article 2) encode understanding as statistical models of linguistic structure. They recognize medical concepts regardless of surface form, detect grammatical relationships between entities, and link terms to standardized vocabularies. They require labeled training data for new domains and struggle with documentation-specific language patterns.

Machine learning classifiers (this article, Section 2) encode understanding as learned associations between text features and clinical outcomes. They are interpretable, computationally efficient, and appropriate for well-defined classification tasks with sufficient annotated training data.

Large Language Models (this article, Section 3) encode understanding as distributed representations of language structure, semantics, and world knowledge. They can perform complex clinical inference, guideline comparison, and reasoning from context, capabilities that no rule-based or feature-engineering approach can replicate. They introduce new risks: hallucination, non-determinism, and limited interpretability that require careful validation before clinical deployment.

The ideal clinical NLP system combines all paradigms strategically: regex for structured fields, scispaCy for entity recognition and normalization, classical ML for classification with interpretable explanations, and LLMs for complex semantic tasks requiring clinical reasoning, all integrated in a validated, audited, privacy-compliant pipeline.

The tools and code patterns presented across this series provide a foundation for building such systems. The next step, as with any medical technology, is rigorous validation on real clinical data before any patient care application.


References

  1. Mikolov T, et al. “Efficient Estimation of Word Representations in Vector Space.” arXiv 2013:1301.3781.
  2. Devlin J, et al. “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.” NAACL-HLT 2019:4171-4186.
  3. Rasmy L, et al. “Med-BERT: pretrained contextualized embeddings on large-scale structured electronic health records for disease prediction.” NPJ Digital Medicine 2021;4(1):86.
  4. Singhal K, et al. “Large language models encode clinical knowledge.” Nature 2023;620:172-180.
  5. Shickel B, et al. “Deep EHR: A Survey of Recent Advances in Deep Learning Techniques for Electronic Health Record (EHR) Analysis.” IEEE Journal of Biomedical and Health Informatics 2018;22(5):1589-1604.
  6. Uzuner O, et al. “Evaluating the state-of-the-art in automatic de-identification.” JAMIA 2007;14(5):550-563.
  7. Yang X, et al. “A large language model for electronic health records.” NPJ Digital Medicine 2022;5(1):194.
  8. Agrawal M, et al. “Large language models are few-shot clinical information extractors.” EMNLP 2022.
  9. Lyu Q, et al. “Translating Radiology Reports into Plain Language using ChatGPT and GPT-4.” arXiv 2023:2303.09038.
  10. Salinas Alvarado MA, et al. “A Systematic Review of Natural Language Processing in Clinical Medicine.” ACM Computing Surveys 2023.
  11. European Parliament. “General Data Protection Regulation (GDPR).” Official Journal of the European Union 2016:L119/1.
  12. U.S. Department of Health and Human Services. “Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule.” 45 CFR Parts 160 and 164, 2002.
  13. Labrak Y, et al. “BioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains.” arXiv 2024:2402.10373.
  14. Chen Z, et al. “MEDITRON-70B: Scaling Medical Pretraining for Large Language Models.” arXiv 2023:2311.16079.
  15. Touvron H, et al. “Llama 2: Open Foundation and Fine-Tuned Chat Models.” arXiv 2023:2307.09288.
  16. Kwon W, et al. “Efficient Memory Management for Large Language Model Serving with PagedAttention.” SOSP 2023. [vLLM]
  17. Jiang AQ, et al. “Mixtral of Experts.” arXiv 2024:2401.04088.
People climb a staircase with uneven steps

Non-Parametric Statistics in Medicine

Posted on May 19, 2026August 9, 2026 by Michele Danilo Pierri

Non-Parametric Statistics in Medicine:

A Practical Guide for Skewed, Ordinal, and Small Clinical Datasets

Last updated: May 2026

Author: Michele D. Pierri

Reading time: 15–20 minutes


Introduction

Clinical data rarely behave like the textbook examples we encounter in introductory statistics courses. Length of stay tends to be right-skewed. Biomarkers such as C-reactive protein, ferritin, D-dimer, troponin, or NT-proBNP frequently show extreme values, sometimes by an order of magnitude. Pain scores and functional classes are bounded, discrete, partly ordinal. And pilot studies, which most of us run at some point in our careers, often involve small samples where a single outlier can drag the mean in a misleading direction.

For these reasons, non-parametric statistics should not be viewed as a secondary or somehow inferior form of analysis. In many clinical scenarios, they are simply the most coherent statistical choice available.

The key idea is straightforward:

Non-parametric tests are useful when the clinical variable is ordinal, markedly skewed, affected by outliers, or when the assumptions required by classical parametric tests are not clinically or statistically convincing.

That said, non-parametric statistics should not be reduced to the simplistic rule:

“If the Shapiro-Wilk test is significant, use a non-parametric test.”

This rule is incomplete and, in our clinical experience, sometimes misleading. The real question is not only whether the data are normally distributed. The real question is whether the statistical method matches the clinical structure of the data, the study design, and the interpretation we want to make.

In this article we will walk through four simulated clinical scenarios:

  1. Pain before and after treatment: Wilcoxon signed-rank test
  2. Length of stay in two independent groups: Mann-Whitney U test
  3. Biomarker levels across three clinical groups: Kruskal-Wallis test
  4. Repeated ordinal clinical scores over time: Friedman test

The examples are intentionally simple. Their purpose is not to replace a full statistical analysis plan but to show how non-parametric thinking works in realistic medical situations.

Quick cheat sheet (clinical workflow)

  • Summarize skewed continuous variables as median (IQR); ordinal variables as median (IQR) or category counts when appropriate.
  • Match design to test: paired → Wilcoxon; 2 independent groups → Mann-Whitney; >2 independent groups → Kruskal-Wallis; repeated measures (ordinal/ranked) → Friedman.
  • Report beyond p-values: an effect size with confidence interval (or at least an effect size), and a clinically interpretable difference (e.g., median change).
  • Plan post-hoc: if Kruskal-Wallis or Friedman is significant, use corrected pairwise comparisons (e.g., Holm or Benjamini-Hochberg).

1. What Does “Non-Parametric” Mean?

Parametric tests, such as the t-test or ANOVA, usually rely on assumptions about the distribution of the data or the distribution of model residuals. They typically focus on means and standard deviations.

Non-parametric tests, by contrast, require fewer distributional assumptions. Many of them work by replacing raw values with their ranks.

Consider, instead of comparing the exact length of stay values:

5, 6, 7, 8, 35

A rank-based method asks where each patient stands relative to the others. This simple reframing makes the method far less sensitive to extreme values.

It does not mean, however, that non-parametric tests have no assumptions at all. They still require appropriate study design, independent observations when required, and correct pairing when the data are paired. They also answer specific statistical questions that are not always identical to their parametric counterparts.

One particularly important point deserves emphasis:

The Mann-Whitney U test is often described as a test comparing medians, but this is strictly valid only under specific distributional conditions. More generally, it tests whether values from one group tend to be larger or smaller than values from another group.

This distinction matters a great deal in clinical writing, even if it is routinely glossed over.


Scenario 1: Pain Before and After Treatment

Clinical question

A group of patients with chronic low back pain receives a new analgesic treatment. Pain is measured using a Visual Analogue Scale (VAS) from 0 to 10, both before and after treatment.

The question is:

Did pain decrease after treatment?

Why a non-parametric test may be appropriate

VAS scores are bounded between 0 and 10. They are frequently not normally distributed. Anyone who has collected VAS data in clinic knows the typical clustering at specific values (0, 5, 8, or 10). The data are also paired: each patient has a before-treatment and after-treatment value.

The appropriate non-parametric test is the Wilcoxon signed-rank test.

It is the non-parametric analogue of the paired t-test, although the interpretation is not exactly the same. It evaluates whether paired differences are symmetrically distributed around zero.

Practical assumptions to check

  • Correct pairing (each “before” matches the same patient “after”).
  • Differences are roughly symmetric around a central value (often 0). If differences are extremely skewed, consider alternative approaches such as the sign test or a model-based strategy.

Basic Python example

import numpy as np
import pandas as pd
from scipy.stats import wilcoxon

np.random.seed(42)

n = 40

pain_before = np.clip(np.random.normal(loc=7.0, scale=1.4, size=n), 0, 10)
improvement = np.random.gamma(shape=2.0, scale=0.8, size=n)
pain_after = np.clip(pain_before - improvement, 0, 10)

data_pain = pd.DataFrame({
    "patient_id": range(1, n + 1),
    "pain_before": pain_before,
    "pain_after": pain_after
})

stat, p_value = wilcoxon(data_pain["pain_before"], data_pain["pain_after"])

print(data_pain.head())
print(f"Wilcoxon statistic:{stat:.3f}")
print(f"p-value:{p_value:.4f}")
Graph showing paired VAS pain scores before and after treatment

Interpretation

If the p-value falls below the chosen significance threshold, we conclude that the post-treatment pain scores differ from the pre-treatment scores in a statistically meaningful way.

Statistical significance, however, is not enough. A reduction of 0.3 points on a VAS scale may be statistically significant in a large sample yet clinically trivial. For this reason the analysis should always report the median change together with its interquartile range.

A reasonable reporting sentence might read:

Pain decreased after treatment, with median VAS changing from 7.1 before treatment to 5.3 after treatment. The reduction was statistically significant according to the Wilcoxon signed-rank test.


Scenario 2: Length of Stay in Two Independent Groups

Clinical question

A cardiac surgery team wants to compare postoperative length of stay between patients with preoperative anemia and patients without preoperative anemia.

The question is:

Is postoperative length of stay different between the two groups?

Why a non-parametric test may be appropriate

Length of stay is rarely normally distributed. Most patients are discharged after a relatively short period. A handful of patients, however, remain hospitalized for considerably longer, often because of complications, frailty, infections, or prolonged rehabilitation needs.

The net result is a right-skewed distribution, sometimes severely so.

In this scenario the two groups are independent. The appropriate non-parametric test is the Mann-Whitney U test.

Practical assumptions to check

  • Independence between groups (no repeated measurements of the same patient across groups).
  • The variable is at least ordinal or continuous, and comparable across groups.

Basic Python example

import numpy as np
import pandas as pd
from scipy.stats import mannwhitneyu

np.random.seed(42)

n_non_anemic = 55
n_anemic = 50

los_non_anemic = np.random.lognormal(mean=1.8, sigma=0.35, size=n_non_anemic)
los_anemic = np.random.lognormal(mean=2.05, sigma=0.45, size=n_anemic)

data_los = pd.DataFrame({
    "group": ["Non-anemic"] * n_non_anemic + ["Anemic"] * n_anemic,
    "length_of_stay": np.concatenate([los_non_anemic, los_anemic])
})

stat, p_value = mannwhitneyu(
    data_los.loc[data_los["group"] == "Non-anemic", "length_of_stay"],
    data_los.loc[data_los["group"] == "Anemic", "length_of_stay"],
    alternative="two-sided"
)

print(data_los.groupby("group")["length_of_stay"].median())
print(f"Mann-Whitney U statistic:{stat:.3f}")
print(f"p-value:{p_value:.4f}")
Graph showing length of stay by anemia status

Interpretation

The Mann-Whitney U test evaluates whether observations in one group tend to be larger than observations in the other group. In our example, if the anemic group has longer length of stay and the p-value is significant, the result supports the hypothesis that preoperative anemia is associated with prolonged hospitalization. This is, incidentally, consistent with several registry studies on cardiac surgery cohorts.

The result should always be reported with medians and interquartile ranges, not just a p-value.

A suitable reporting sentence:

Length of stay was longer in anemic patients than in non-anemic patients. Data are reported as median and interquartile range, and between-group comparison was performed using the Mann-Whitney U test.


Scenario 3: Biomarker Levels Across Three Clinical Groups

Clinical question

Suppose we want to compare postoperative C-reactive protein levels across three groups of patients:

  1. No complication
  2. Minor complication
  3. Major complication

The question is:

Do postoperative inflammatory marker levels differ across complication severity groups?

Why a non-parametric test may be appropriate

Biomarkers such as CRP are notoriously skewed. Some patients show extremely high values because of infection, systemic inflammation, tissue injury, or postoperative complications.

When we have more than two independent groups, the non-parametric analogue of one-way ANOVA is the Kruskal-Wallis test.

Practical assumptions to check

  • Independent groups (each patient contributes to one group only).
  • The test detects distributional differences across groups; interpreting it as a comparison of “median differences” is safest only under additional shape assumptions.

Basic Python example

import numpy as np
import pandas as pd
from scipy.stats import kruskal

np.random.seed(42)

n_no = 45
n_minor = 40
n_major = 35

crp_no = np.random.lognormal(mean=3.2, sigma=0.35, size=n_no)
crp_minor = np.random.lognormal(mean=3.6, sigma=0.40, size=n_minor)
crp_major = np.random.lognormal(mean=4.0, sigma=0.45, size=n_major)

data_crp = pd.DataFrame({
    "group": (
        ["No complication"] * n_no +
        ["Minor complication"] * n_minor +
        ["Major complication"] * n_major
    ),
    "crp": np.concatenate([crp_no, crp_minor, crp_major])
})

stat, p_value = kruskal(
    data_crp.loc[data_crp["group"] == "No complication", "crp"],
    data_crp.loc[data_crp["group"] == "Minor complication", "crp"],
    data_crp.loc[data_crp["group"] == "Major complication", "crp"]
)

print(data_crp.groupby("group")["crp"].median())
print(f"Kruskal-Wallis statistic:{stat:.3f}")
print(f"p-value:{p_value:.4f}")
Graph of CRP across complication severity groups

Interpretation

A significant Kruskal-Wallis test tells us that at least one group tends to differ from the others. It does not, on its own, tell us which groups differ.

When the global test is significant, post-hoc pairwise comparisons may be performed, commonly using Mann-Whitney U tests with correction for multiple comparisons. The plan, however, should be prespecified and reported transparently. Adding tests after seeing the data is a recipe for false-positive findings.

A reasonable reporting sentence:

CRP levels differed across complication groups according to the Kruskal-Wallis test. Post-hoc pairwise comparisons were then performed with correction for multiple testing.


Scenario 4: Repeated Ordinal Clinical Scores Over Time

Clinical question

A group of patients is followed after a rehabilitation program. Functional limitation is measured at baseline, 1 month, 3 months, and 6 months using an ordinal score from 1 to 5, where higher values indicate worse functional limitation.

The question is:

Does functional limitation improve over time?

Why a non-parametric test may be appropriate

The score is ordinal. The distance between score 1 and score 2 is not necessarily the same as the distance between score 4 and score 5; in fact, in clinical experience, the jump from 4 to 5 often represents a far more substantial functional deterioration. In addition, the same patients are measured repeatedly over time.

The appropriate non-parametric test in this situation is the Friedman test, the non-parametric analogue of repeated-measures ANOVA for ranked data.

Practical assumptions to check

  • Repeated measurements are on the same subjects across time points (complete blocks are ideal).
  • The ordinal scale is appropriate for rank-based comparisons. If missingness is substantial, dedicated longitudinal methods may be preferable.

Basic Python example

import numpy as np
import pandas as pd
from scipy.stats import friedmanchisquare

np.random.seed(42)

n = 35

baseline = np.random.choice([3, 4, 5], size=n, p=[0.25, 0.45, 0.30])
month_1 = np.clip(baseline - np.random.choice([0, 1], size=n, p=[0.45, 0.55]), 1, 5)
month_3 = np.clip(month_1 - np.random.choice([0, 1], size=n, p=[0.40, 0.60]), 1, 5)
month_6 = np.clip(month_3 - np.random.choice([0, 1], size=n, p=[0.55, 0.45]), 1, 5)

data_function = pd.DataFrame({
    "patient_id": range(1, n + 1),
    "baseline": baseline,
    "month_1": month_1,
    "month_3": month_3,
    "month_6": month_6
})

stat, p_value = friedmanchisquare(
    data_function["baseline"],
    data_function["month_1"],
    data_function["month_3"],
    data_function["month_6"]
)

print(data_function.head())
print(f"Friedman statistic:{stat:.3f}")
print(f"p-value:{p_value:.4f}")

Median ordinal functional score over time

Interpretation

A significant Friedman test suggests that the repeated measurements are not all drawn from the same distribution. In our example, it would support the presence of a change in functional limitation over time.

As with Kruskal-Wallis, the global test does not identify which time points differ. When needed, post-hoc paired comparisons can be performed with appropriate correction.

A reasonable reporting sentence:

Functional limitation improved over time. The overall change across follow-up visits was significant according to the Friedman test.


Practical Decision Table

Clinical scenarioData structureTypical variableParametric testNon-parametric test
Pain before and after therapyPaired observationsVAS scorePaired t-testWilcoxon signed-rank test
Length of stay in two groupsTwo independent groupsDays of hospitalizationIndependent t-testMann-Whitney U test
Biomarker across severity groupsMore than two independent groupsCRP, ferritin, D-dimerOne-way ANOVAKruskal-Wallis test
Ordinal score over timeRepeated measuresFunctional score, NYHA-like scoreRepeated-measures ANOVAFriedman test

Common Mistakes in Medical Papers

1. Using non-parametric tests automatically after a significant normality test

Normality tests can be overly sensitive in large samples and underpowered in small samples. The decision should also rest on histograms, Q-Q plots, clinical plausibility, outliers, and the actual scale of measurement. Mechanical reliance on a single test is rarely a good idea.

2. Saying that Mann-Whitney always compares medians

This is a common oversimplification. When the two distributions have a similar shape, the Mann-Whitney U test can be interpreted as a test of location shift. When the distributions differ in shape or spread, the interpretation is broader: one group tends to have larger or smaller values than the other.

3. Reporting mean and standard deviation for markedly skewed variables

For strongly skewed variables, median and interquartile range are usually more clinically informative.

4. Reporting only p-values

A p-value does not quantify clinical relevance. Whenever possible, the report should include effect sizes, median differences, confidence intervals, and clinically interpretable summaries.

5. Ignoring multiple comparisons

After a significant Kruskal-Wallis or Friedman test, post-hoc comparisons must account for multiplicity. Otherwise, the probability of false-positive findings rises quickly.

6. Confusing statistical significance with clinical importance

A statistically significant reduction in pain score may still be clinically irrelevant when the magnitude is small. Conversely, a clinically relevant difference can easily fail to reach statistical significance in a small pilot study.


How to Report Non-Parametric Analyses

A concise reporting style could be:

Continuous skewed variables were summarized as median and interquartile range. Between-group comparisons were performed using the Mann-Whitney U test for two independent groups and the Kruskal-Wallis test for more than two independent groups. Paired before-after comparisons were performed using the Wilcoxon signed-rank test. Repeated ordinal measurements were analyzed using the Friedman test. A two-sided p-value < 0.05 was considered statistically significant.

For a more complete report, effect sizes and confidence intervals should be added where possible. Practical options include:

  • Wilcoxon signed-rank: rank-biserial correlation or an r-type effect size (when available).
  • Mann-Whitney: rank-biserial correlation or Cliff’s delta.
  • Kruskal-Wallis: epsilon-squared (or a similar rank-based η²) to quantify the global effect.
  • Friedman: Kendall’s W as an overall effect size.

Conclusion

Non-parametric statistics are particularly useful in medicine, where clinical data are routinely skewed, ordinal, bounded, or affected by outliers. These methods are not a fallback for weak data analysis. In many cases they represent the most appropriate way to analyze real-world clinical variables.

The practical message is straightforward:

Choose the test according to the clinical question, the study design, and the structure of the variable, not only according to a mechanical normality test.

In clinical research the goal is not to deploy the most sophisticated method available. The goal is to use a method that respects the data and produces an interpretation that is clinically meaningful.


Suggested References

Hollander M, Wolfe DA, Chicken E. Nonparametric Statistical Methods. Wiley.

Altman DG. Practical Statistics for Medical Research. Chapman & Hall/CRC.

Bland M. An Introduction to Medical Statistics. Oxford University Press.

Conover WJ. Practical Nonparametric Statistics. Wiley.


See also on this site: Nonparametric statistics

A doctor translates text

Natural Language Processing for Medical Report Analysis – Part 2

Posted on May 9, 2026May 25, 2026 by Michele Danilo Pierri

Introduction to Natural Language Processing for Medical Report Analysis: Advanced Linguistic Analysis with spaCy and scispaCy for Medical Reports


Article authored by Michele D. Pierri, MD

Cardiac Surgeon & Medical Technology Researcher

Last updated: May 2025

Reading time: 30 minutes


Abstract

Regular expressions give us a solid foundation for pulling structured data out of clinical documents. They struggle, though, with the semantic complexity that defines medical language. This article introduces two complementary libraries (spaCy and scispaCy) that bring genuine linguistic intelligence to clinical text. We work through detailed examples on a post-CABG discharge summary, covering tokenization, part-of-speech tagging, dependency parsing, and Named Entity Recognition (NER) with biomedical models. UMLS entity linking is also explored, since it ties extracted clinical terms to standardized medical ontologies and enables interoperability across health information systems. By the final section the reader will know not only how to implement these techniques in Python, but when to prefer them over regex; and when, in the daily reality of clinical NLP, the two should be combined.

Article Series:

  1. Introduction to NLP and Regex for Medical Reports
  2. Advanced Linguistic Analysis with spaCy and scispaCy
  3. Machine Learning and Large Language Models for Clinical Text Analysis

1. Beyond Pattern Matching: Why Linguistic Analysis?

1.1 The Limits of Regular Expressions

The first article of this series built a regex parser able to extract patient demographics, surgical parameters, medications, and complications from a post-CABG discharge summary. That implementation worked because our sample document followed predictable formatting conventions. Real-world clinical documentation, in our experience reading hundreds of operative reports across centers, is far less disciplined.

Consider these semantically equivalent phrases that all describe the same complication:

"Postoperative atrial fibrillation was noted on POD 3"
"Patient developed new-onset AF in the immediate post-surgical period"
"Rhythm monitoring revealed paroxysmal atrial fibrillation 72 hours after bypass"
"POD 3: AF, cardioverted successfully"

A regex pattern designed to catch the first phrasing will miss the others. We could write extra patterns for each variant, but this becomes an arms race against the variability of natural language. Regex also cannot:

  • Determine that “no wound infections” and “wound infection noted” have opposite clinical meanings (negation handling)
  • Understand that “CABG,” “bypass surgery,” and “coronary revascularization” refer to the same procedure
  • Extract the relationship between a drug and its indication (“amiodarone for atrial fibrillation”)
  • Handle abbreviation disambiguation (“AF” could mean atrial fibrillation or aortic flow depending on context)

These limitations are what push us toward NLP tools built on computational linguistics, the scientific study of language structure.

1.2 The Linguistic Approach

Natural language processing at the linguistic level rests on a simple premise: language has internal structure, and that structure can be computationally modeled. Rather than matching character sequences, linguistic NLP:

  1. Tokenizes text into meaningful units (words, punctuation, medical abbreviations)
  2. Tags each token with its grammatical role (noun, verb, adjective)
  3. Parses the syntactic relationships between tokens (subject, object, modifier)
  4. Recognizes named entities, that is, spans of text referring to real-world concepts (diseases, drugs, procedures)
  5. Links recognized entities to external knowledge bases (UMLS, SNOMED-CT, RxNorm)

The pipeline turns raw text not into matched strings, but into a structured representation of linguistic meaning.

1.3 The spaCy Ecosystem for Biomedical Text

spaCy is an industrial-strength NLP library for Python, designed for production use rather than academic experimentation. It delivers fast, accurate linguistic analysis through pre-trained statistical models built on large corpora. For general English text, the en_core_web_sm/md/lg models are the standard starting point.

Medical language presents a domain-specific challenge. General-purpose NLP models are trained predominantly on news articles, books, and web content, corpora that share little with clinical documentation. Medical text is characterized by:

  • High density of Latin and Greek terminology
  • Extensive use of abbreviations and acronyms (POD, CABG, LVEF, LIMA)
  • Domain-specific syntactic patterns
  • Specialized named entity types (diseases, procedures, anatomical structures, medications)
  • Negation patterns with clinical significance (“no evidence of”, “without complications”)

This is where scispaCy earns its place. Developed by the Allen Institute for AI, scispaCy provides spaCy-compatible models trained on biomedical scientific literature. The performance gap on clinical text, compared to general-purpose models, is substantial.


2. spaCy Fundamentals for Medical Text

2.1 Installation and Setup

Before proceeding, ensure you have the required libraries installed. spaCy and scispaCy require careful version management:

# Install spaCy
pip install spacy

# Install scispaCy
pip install scispacy

# Install a scispaCy biomedical model
# en_core_sci_md offers a good balance of accuracy and performance
pip install <https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/en_core_sci_md-0.5.4.tar.gz>

# For UMLS entity linking (requires additional setup)
pip install scispacy[linker]

Note on model selection: scispaCy provides several pre-trained models. en_core_sci_sm is the smallest and fastest; en_core_sci_md includes word vectors for better similarity comparisons; en_core_sci_lg offers maximum accuracy. For production medical applications, en_core_sci_md or en_core_sci_lg is recommended. For research requiring UMLS linking specifically, en_core_sci_lg combined with the UmlsEntityLinker is the gold standard.

2.2 The spaCy Document Object Model

When spaCy processes a text string, it returns a Doc object: a rich data structure that exposes all linguistic annotations. Understanding this object model is essential before doing anything serious with the library.

import spacy

# Load the general English model first to understand the basics
# (we will switch to scispaCy models for medical text shortly)
nlp = spacy.load("en_core_web_sm")

sample_text = """
The patient was extubated on POD 1 without complications.
Postoperative atrial fibrillation was noted on POD 3,
successfully cardioverted with amiodarone loading followed by maintenance dose.
"""

doc = nlp(sample_text)

# The Doc object contains Token objects
print(f"Number of tokens:{len(doc)}")

# Each token exposes multiple attributes
for token in doc[:10]:  # First 10 tokens
    print(f"  Token:{token.text!r:20} | POS:{token.pos_:10} | Lemma:{token.lemma_:20} | Stop:{token.is_stop}")

Output:

Number of tokens: 40
Token: ‘The’ | POS: DET | Lemma: ‘the’ | Stop: True
Token: ‘patient’ | POS: NOUN | Lemma: ‘patient’ | Stop: False
Token: ‘was’ | POS: AUX | Lemma: ‘be’ | Stop: True
Token: ‘extubated’ | POS: VERB | Lemma: ‘extubate’ | Stop: False
Token: ‘on’ | POS: ADP | Lemma: ‘on’ | Stop: True
Token: ‘POD’ | POS: PROPN | Lemma: ‘POD’ | Stop: False
Token: ‘1’ | POS: NUM | Lemma: ‘1’ | Stop: False

2.3 Tokenization in Medical Text

Tokenization (splitting text into individual tokens) sounds trivial, but in clinical NLP it is anything but. Take “Amiodarone 200mg”: should this be one token or two? What about “POD 3”, or “LVEF 40%”? Medical abbreviations punctuated with periods (“M.D.”, “i.v.”) add another layer of trouble.

spaCy’s tokenization rules are based on the Penn Treebank standard but can be customized. Let us look at how tokenization handles our discharge summary:

import spacy
from typing import List, Dict

def analyze_tokenization(text: str, nlp) -> List[Dict]:
    """
    Analyze how spaCy tokenizes medical text.

    This function reveals tokenization decisions that are clinically
    significant - for example, whether "100mg" is treated as one token
    or split into "100" and "mg".

    Args:
        text: Clinical text to tokenize
        nlp: Loaded spaCy model

    Returns:
        List of token dictionaries with linguistic attributes
    """
    doc = nlp(text)
    tokens = []

    for token in doc:
        tokens.append({
            'text': token.text,
            'lemma': token.lemma_,
            'pos': token.pos_,          # Coarse-grained POS (NOUN, VERB, etc.)
            'tag': token.tag_,           # Fine-grained POS (NNS, VBD, etc.)
            'dep': token.dep_,           # Syntactic dependency role
            'is_alpha': token.is_alpha,
            'is_digit': token.is_digit,
            'is_stop': token.is_stop,    # Common words with little semantic value
            'is_punct': token.is_punct,
            'shape': token.shape_,       # Character shape: "100mg" → "dddxx"
        })

    return tokens

# Test on the medication section of our discharge summary
medication_text = """
DISCHARGE MEDICATIONS:
1. Aspirin 100mg daily
2. Clopidogrel 75mg daily (for 12 months)
3. Atorvastatin 80mg daily
4. Metoprolol 50mg twice daily
5. Ramipril 5mg daily
6. Amiodarone 200mg daily (for 6 weeks)
"""

# We'll use scispaCy for actual medical processing
nlp_sci = spacy.load("en_core_sci_md")
tokens = analyze_tokenization(medication_text, nlp_sci)

# Print clinically relevant tokens (non-stop, non-punctuation)
print("Clinically Relevant Tokens:")
print(f"{'Text':<20}{'Lemma':<20}{'POS':<10}{'Shape':<15}")
print("-" * 70)
for t in tokens:
    if not t['is_stop'] and not t['is_punct'] and t['text'].strip():
        print(f"{t['text']:<20}{t['lemma']:<20}{t['pos']:<10}{t['shape']:<15}")

Clinical Insight on Tokenization: scispaCy handles medical abbreviations more gracefully than general models, although the picture is not uniform. “100mg” may be split or kept together depending on the model and configuration. For downstream extraction tasks, custom tokenization rules for domain-specific patterns are often beneficial:

from spacy.lang.char_classes import ALPHA, ALPHA_LOWER, ALPHA_UPPER
from spacy.lang.en import English

def add_medical_tokenization_rules(nlp):
    """
    Add custom tokenization rules for medical text.

    Medical documents contain patterns that standard tokenizers
    handle poorly: drug-dose combinations, medical acronyms,
    and clinical measurement notations.
    """
    # Prevent splitting on medical dose patterns like "100mg", "5mg"
    # These should remain as single tokens for accurate extraction
    infixes = nlp.Defaults.infixes

    # Add special cases for common medical abbreviations
    special_cases = {
        "POD": [{"ORTH": "POD"}],           # Postoperative Day
        "CPB": [{"ORTH": "CPB"}],           # Cardiopulmonary Bypass
        "CABG": [{"ORTH": "CABG"}],         # Coronary Artery Bypass Grafting
        "LIMA": [{"ORTH": "LIMA"}],         # Left Internal Mammary Artery
        "SVG": [{"ORTH": "SVG"}],           # Saphenous Vein Graft
        "LVEF": [{"ORTH": "LVEF"}],         # Left Ventricular Ejection Fraction
        "LAD": [{"ORTH": "LAD"}],           # Left Anterior Descending
        "RCA": [{"ORTH": "RCA"}],           # Right Coronary Artery
    }

    for text, case in special_cases.items():
        nlp.tokenizer.add_special_case(text, case)

    return nlp

2.4 Part-of-Speech Tagging

Part-of-speech (POS) tagging assigns each token its grammatical role. In medical NLP, POS tags are useful for identifying:

  • Nouns (NOUN, PROPN): Disease names, anatomical structures, drugs
  • Verbs (VERB): Procedures performed, clinical events (“extubated,” “cardioverted”)
  • Adjectives (ADJ): Qualitative descriptions (“satisfactory,” “stable,” “reduced”)
  • Numbers (NUM): Dosages, durations, laboratory values
def extract_clinical_verbs(text: str, nlp) -> List[Dict]:
    """
    Extract action verbs from clinical text.

    In cardiac surgery reports, verbs often encode critical procedural
    and clinical events: "performed," "extubated," "cardioverted,"
    "noted," "removed." This function identifies these for downstream
    event extraction.

    Args:
        text: Clinical narrative text
        nlp: scispaCy model

    Returns:
        List of clinical verbs with their lemmas and surrounding context
    """
    doc = nlp(text)
    clinical_verbs = []

    for token in doc:
        if token.pos_ == "VERB" and not token.is_stop:
            # Get 3-word context window around verb
            start = max(0, token.i - 3)
            end = min(len(doc), token.i + 4)
            context = doc[start:end].text

            clinical_verbs.append({
                'verb': token.text,
                'lemma': token.lemma_,
                'tense': token.tag_,    # VBD = past tense, VBN = past participle
                'context': context,
                'negated': _check_negation(token, doc)  # See negation section
            })

    return clinical_verbs

def _check_negation(token, doc) -> bool:
    """
    Simple negation detection using dependency parsing.

    Checks if a token has a negation modifier in its dependency tree.
    More sophisticated negation handling (like NegEx algorithm) is
    covered in the discussion section.
    """
    for child in token.children:
        if child.dep_ == "neg":
            return True
    return False

# Apply to the postoperative course section
postop_text = """
The patient was extubated on POD 1 without complications.
Chest tubes were removed on POD 2.
Postoperative atrial fibrillation was noted on POD 3,
successfully cardioverted with amiodarone loading.
No wound infections or sternal instability noted.
"""

nlp_sci = spacy.load("en_core_sci_md")
verbs = extract_clinical_verbs(postop_text, nlp_sci)

print("Clinical Verbs Detected:")
print(f"{'Verb':<20}{'Lemma':<20}{'Tense':<8}{'Negated':<10}{'Context'}")
print("-" * 90)
for v in verbs:
    print(f"{v['verb']:<20}{v['lemma']:<20}{v['tense']:<8}{str(v['negated']):<10}{v['context']}")

2.5 Dependency Parsing: Understanding Relationships

Dependency parsing is where linguistic analysis starts to look like genuine clinical intelligence. A dependency parser builds a tree showing how words in a sentence relate to each other grammatically. In medical text, this opens the door to extracting structured relationships:

  • Subject → Verb → Object: “amiodarone cardioverted atrial fibrillation”
  • Modifier → Head noun: “postoperative atrial fibrillation” (adjective modifying noun)
  • Prepositional relationships: “cardioverted with amiodarone” (instrument of action)
def extract_drug_event_relations(text: str, nlp) -> List[Dict]:
    """
    Extract drug-event relationships using dependency parsing.

    Dependency parsing allows us to identify sentences like:
    "successfully cardioverted with amiodarone" →
      drug: amiodarone, action: cardioverted, outcome: successful

    This is a level of extraction impossible with regex alone,
    as it requires understanding sentence structure, not just
    character patterns.

    Args:
        text: Clinical text containing drug and event mentions
        nlp: Loaded spaCy/scispaCy model

    Returns:
        List of extracted drug-event relationships
    """
    doc = nlp(text)
    relations = []

    # Medical procedure verbs relevant to cardiac surgery
    PROCEDURE_VERBS = {
        'cardiovert', 'extubate', 'remove', 'perform',
        'complete', 'administer', 'initiate', 'discontinue'
    }

    # Common cardiac drug keywords
    CARDIAC_DRUGS = {
        'amiodarone', 'aspirin', 'clopidogrel', 'atorvastatin',
        'metoprolol', 'ramipril', 'furosemide', 'metformin'
    }

    for token in doc:
        # Find procedure verbs
        if token.lemma_.lower() in PROCEDURE_VERBS:
            relation = {
                'verb': token.text,
                'verb_lemma': token.lemma_,
                'drug': None,
                'patient_subject': False,
                'outcome_modifier': None,
                'sentence': token.sent.text.strip()
            }

            # Traverse dependency tree for related tokens
            for child in token.children:
                # Check for drug as instrument (prepositional: "with amiodarone")
                if child.dep_ == "prep":
                    for grandchild in child.children:
                        if grandchild.lemma_.lower() in CARDIAC_DRUGS:
                            relation['drug'] = grandchild.text

                # Check for patient as subject ("patient was extubated")
                if child.dep_ in ("nsubj", "nsubjpass"):
                    if child.lemma_.lower() in ("patient",):
                        relation['patient_subject'] = True

                # Check for outcome adverbs ("successfully cardioverted")
                if child.dep_ == "advmod":
                    relation['outcome_modifier'] = child.text

            # Also check the token's own head for drugs
            for sibling in token.head.children:
                if sibling.lemma_.lower() in CARDIAC_DRUGS and sibling != token:
                    if not relation['drug']:
                        relation['drug'] = sibling.text

            if relation['verb']:
                relations.append(relation)

    return relations

# Test on the postoperative section
postop_text = """
The patient was extubated on POD 1 without complications.
Chest tubes were removed on POD 2.
Postoperative atrial fibrillation was noted on POD 3,
successfully cardioverted with amiodarone loading followed by maintenance dose.
"""

relations = extract_drug_event_relations(postop_text, nlp_sci)

print("Drug-Event Relationships Extracted:")
for r in relations:
    print(f"\\n  Verb:{r['verb']}")
    print(f"  Drug:{r['drug'] or 'N/A'}")
    print(f"  Patient subject:{r['patient_subject']}")
    print(f"  Outcome modifier:{r['outcome_modifier'] or 'N/A'}")
    print(f"  Sentence:{r['sentence'][:80]}...")

The dependency parser thus enables extraction of structured relationships: not just entity mentions, but how entities interact within clinical events. This capability has direct applications in pharmacovigilance, complication tracking, and automated population of surgical registries.


3. scispaCy: Biomedical NLP

3.1 Why Standard Models Fail for Medical Text

To appreciate scispaCy’s contribution, it helps to compare how a general-purpose spaCy model and a scispaCy model handle the same clinical sentence. Take: “LIMA to LAD anastomosis was completed with satisfactory flow measurements.”

General spaCy (en_core_web_sm):

  • “LIMA” → labeled as ORG (organization), incorrect
  • “LAD” → labeled as ORG, incorrect
  • “anastomosis” → not recognized as a named entity at all

scispaCy (en_core_sci_md):

  • “LIMA” → recognized with ENTITY label (biomedical entity)
  • “LAD” → recognized as biomedical entity
  • “anastomosis” → recognized as biomedical entity

The difference is training data. scispaCy models are trained on the CRAFT corpus (Colorado Richly Annotated Full-Text Corpus) and PubMed abstracts, scientific biomedical text that includes the vocabulary and entity patterns of clinical medicine. The NER annotations in these corpora rely on the BC5CDR schema for chemical and disease entities, plus broader entity annotations from ontologies like UMLS.

3.2 Named Entity Recognition with scispaCy

Named Entity Recognition (NER) is the automatic identification and classification of named entities in text; in our case, medical concepts such as diseases, drugs, procedures, and anatomical structures.

import spacy
from typing import List, Dict, Tuple
import json

# Load scispaCy model
nlp_sci = spacy.load("en_core_sci_md")

# Full discharge summary for NER analysis
DISCHARGE_SUMMARY = """
CARDIAC SURGERY DISCHARGE SUMMARY

Patient: John Smith
MRN: 12345678
DOB: 15/03/1958
Admission Date: 10/01/2024
Discharge Date: 18/01/2024

DIAGNOSIS:
1. Triple vessel coronary artery disease
2. Reduced left ventricular function (LVEF 35%)
3. Hypertension
4. Type 2 Diabetes Mellitus

PROCEDURE PERFORMED:
Coronary Artery Bypass Grafting x3 (CABG)
- LIMA to LAD
- SVG to OM1
- SVG to RCA
Date of surgery: 11/01/2024

OPERATIVE DETAILS:
The procedure was performed via median sternotomy under general anesthesia.
Cardiopulmonary bypass time: 98 minutes
Aortic cross-clamp time: 67 minutes
All anastomoses were completed with satisfactory flow measurements.

POSTOPERATIVE COURSE:
The patient was extubated on POD 1 without complications. Chest tubes were removed on POD 2.
Postoperative atrial fibrillation was noted on POD 3, successfully cardioverted with amiodarone
loading followed by maintenance dose. Patient remained in sinus rhythm thereafter.

No wound infections or sternal instability noted. Echocardiography on POD 5 showed LVEF 40%
with good biventricular function and no pericardial effusion.

DISCHARGE MEDICATIONS:
1. Aspirin 100mg daily
2. Clopidogrel 75mg daily (for 12 months)
3. Atorvastatin 80mg daily
4. Metoprolol 50mg twice daily
5. Ramipril 5mg daily
6. Amiodarone 200mg daily (for 6 weeks)
7. Metformin 1000mg twice daily
8. Furosemide 40mg daily (for 2 weeks)

FOLLOW-UP:
- Outpatient cardiology clinic in 2 weeks
- Cardiac rehabilitation program enrollment
- INR monitoring not required (patient on dual antiplatelet therapy)

The patient was discharged home in stable condition with appropriate wound care instructions.

Dr. Sarah Johnson, MD
Cardiac Surgery Department
"""

def extract_biomedical_entities(text: str, nlp) -> Dict[str, List[str]]:
    """
    Extract all biomedical entities from clinical text using scispaCy NER.

    scispaCy's NER model identifies spans of text corresponding to
    biomedical concepts: diseases, chemicals, genes, cell types,
    species, and more. The entity label set depends on the model used:

    - en_core_sci_*: Uses a single 'ENTITY' label for all biomedical concepts
    - en_ner_bc5cdr_md: Distinguishes DISEASE and CHEMICAL entities
    - en_ner_jnlpba_md: Tags GENE, PROTEIN, RNA, DNA, CELL_TYPE, CELL_LINE

    For general clinical use, en_core_sci_md with UMLS linking provides
    the most complete semantic annotation.

    Args:
        text: Clinical text
        nlp: Loaded scispaCy model

    Returns:
        Dictionary grouping entities by their labels
    """
    doc = nlp(text)

    entities_by_label: Dict[str, List[str]] = {}

    for ent in doc.ents:
        label = ent.label_
        if label not in entities_by_label:
            entities_by_label[label] = []

        # Deduplicate while preserving order
        entity_text = ent.text.strip()
        if entity_text and entity_text not in entities_by_label[label]:
            entities_by_label[label].append(entity_text)

    return entities_by_label

# Extract entities from the discharge summary
entities = extract_biomedical_entities(DISCHARGE_SUMMARY, nlp_sci)

print("Biomedical Entities Extracted by scispaCy:")
print("=" * 60)
for label, ents in sorted(entities.items()):
    print(f"\\n[{label}] ({len(ents)} entities):")
    for entity in ents[:15]:  # Show first 15 per category
        print(f"  -{entity}")

3.3 Specialized NER Models: BC5CDR for Diseases and Chemicals

For clinical NLP requiring a precise distinction between disease entities and chemical/drug entities, the en_ner_bc5cdr_md scispaCy model trained on the BioCreative V CDR corpus is the appropriate choice. The model was trained to distinguish:

  • DISEASE: Pathological conditions (“atrial fibrillation,” “coronary artery disease,” “hypertension”)
  • CHEMICAL: Drugs and chemical compounds (“amiodarone,” “aspirin,” “atorvastatin”)
# Note: requires installation of en_ner_bc5cdr_md
# pip install <https://s3-us-west-2.amazonaws.com/ai2-s2-scispacy/releases/v0.5.4/en_ner_bc5cdr_md-0.5.4.tar.gz>

def extract_diseases_and_chemicals(text: str) -> Dict[str, List[str]]:
    """
    Use the BC5CDR model to separately identify diseases and drugs.

    The BC5CDR model was trained on 1,500 PubMed articles with
    expert annotations of disease and chemical mentions. It achieves
    F1 scores >85% on the BC5CDR test set, making it suitable for
    clinical applications where chemical-disease distinction matters.

    Clinical application: Automatically extracting the complication
    profile and pharmacological treatment from discharge summaries
    for pharmacovigilance databases.

    Args:
        text: Clinical document text

    Returns:
        Dictionary with 'DISEASE' and 'CHEMICAL' entity lists
    """
    try:
        nlp_cdr = spacy.load("en_ner_bc5cdr_md")
    except OSError:
        print("Model en_ner_bc5cdr_md not installed. Using en_core_sci_md instead.")
        return extract_biomedical_entities(text, spacy.load("en_core_sci_md"))

    doc = nlp_cdr(text)

    result = {"DISEASE": [], "CHEMICAL": []}

    for ent in doc.ents:
        entity_text = ent.text.strip().lower()
        label = ent.label_
        if label in result and entity_text not in result[label]:
            result[label].append(ent.text.strip())

    return result

# Test on our discharge summary
entities_bc5cdr = extract_diseases_and_chemicals(DISCHARGE_SUMMARY)

print("Diseases detected:")
for disease in entities_bc5cdr.get("DISEASE", []):
    print(f"  ✓{disease}")

print("\\nChemicals/Drugs detected:")
for chemical in entities_bc5cdr.get("CHEMICAL", []):
    print(f"  ✓{chemical}")

3.4 Abbreviation Detection and Expansion

Medical text is dense with abbreviations, a major source of ambiguity for NLP systems. “AF” in a cardiology report almost certainly means atrial fibrillation; in a different context, it might mean acid-fast, amniotic fluid, or audio frequency. scispaCy provides an AbbreviationDetector component that identifies abbreviation-definition pairs within the same document, enabling in-context disambiguation.

import spacy
import scispacy
from scispacy.abbreviation import AbbreviationDetector

def setup_abbreviation_detector(model_name: str = "en_core_sci_md") -> spacy.language.Language:
    """
    Configure scispaCy pipeline with abbreviation detection.

    The AbbreviationDetector implements the algorithm from Schwartz & Hearst (2003),
    which identifies abbreviation-definition pairs based on character matching
    heuristics. When a document defines "CABG" as "Coronary Artery Bypass Grafting",
    the detector enables automatic expansion of all subsequent CABG mentions.

    Returns:
        spaCy language model with abbreviation detection configured
    """
    nlp = spacy.load(model_name)

    # Add abbreviation detector to the pipeline
    nlp.add_pipe("abbreviation_detector")

    return nlp

def extract_abbreviations(text: str, nlp) -> Dict[str, str]:
    """
    Extract and map abbreviations found in clinical text.

    Medical abbreviations are a significant challenge for NLP systems:
    the same abbreviation can have different meanings in different
    specialties or even different sections of the same document.

    This function identifies abbreviation-definition pairs where
    both the abbreviation and its expansion appear in the same document.

    Args:
        text: Clinical document text
        nlp: scispaCy model with AbbreviationDetector

    Returns:
        Dictionary mapping abbreviation → full form
    """
    doc = nlp(text)
    abbreviation_map = {}

    for abbreviation in doc._.abbreviations:
        abbr_text = abbreviation.text
        # The long form is the identified full expansion
        long_form = abbreviation._.long_form.text if abbreviation._.long_form else "Unknown"
        abbreviation_map[abbr_text] = long_form

    return abbreviation_map

# Note: Our discharge summary uses abbreviations but may not always define them inline
# Let's create a version with explicit definitions to demonstrate the capability
annotated_summary = """
The patient underwent Coronary Artery Bypass Grafting (CABG) using Left Internal
Mammary Artery (LIMA) to Left Anterior Descending (LAD) and Saphenous Vein Graft (SVG)
to Obtuse Marginal (OM1) and Right Coronary Artery (RCA).
Cardiopulmonary Bypass (CPB) time was 98 minutes.
Postoperative atrial fibrillation (AF) was noted and treated with amiodarone.
Left Ventricular Ejection Fraction (LVEF) improved from 35% to 40%.
"""

nlp_with_abbrev = setup_abbreviation_detector()
abbreviations = extract_abbreviations(annotated_summary, nlp_with_abbrev)

print("Abbreviations detected and expanded:")
for abbr, expansion in sorted(abbreviations.items()):
    print(f"{abbr} →{expansion}")

Expected output:

Abbreviations detected and expanded:
  AF → atrial fibrillation
  CABG → Coronary Artery Bypass Grafting
  CPB → Cardiopulmonary Bypass
  LAD → Left Anterior Descending
  LIMA → Left Internal Mammary Artery
  LVEF → Left Ventricular Ejection Fraction
  OM1 → Obtuse Marginal
  RCA → Right Coronary Artery
  SVG → Saphenous Vein Graft

A caveat from clinical practice: the Schwartz & Hearst algorithm only catches abbreviations that are explicitly defined inside the same document. Most operative reports we read in our daily work assume the reader already knows what CABG, LIMA or LVEF mean and never spell them out. For real deployment, this detector almost always needs to be paired with a curated cardiac-surgery glossary as a fallback.


4. UMLS Entity Linking: Connecting to Medical Ontologies

4.1 The Importance of Standardized Medical Vocabulary

Named entity recognition tells us that a text span refers to a medical concept; UMLS linking tells us which standardized concept it refers to. The distinction matters for interoperability.

Consider: “coronary artery disease,” “CAD,” “ischemic heart disease,” and “arteriosclerotic heart disease” are all strings pointing to the same underlying condition. In the Unified Medical Language System (UMLS), this concept has a single Concept Unique Identifier: C0010068. By linking extracted entities to UMLS CUIs we get:

  1. Semantic normalization: All equivalent surface forms map to the same concept
  2. Cross-database interoperability: UMLS bridges SNOMED-CT, ICD-10, MeSH, RxNorm, and thousands of other vocabularies
  3. Knowledge graph integration: CUIs can be used to query external medical knowledge bases
  4. Research reproducibility: Studies using CUI-based cohort definitions are unambiguous

The UMLS Metathesaurus is maintained by the U.S. National Library of Medicine and covers over 3.5 million concepts across more than 200 biomedical vocabularies. Key semantic type hierarchies relevant to cardiac surgery include:

  • T047, Disease or Syndrome (coronary artery disease, atrial fibrillation)
  • T121, Pharmacologic Substance (amiodarone, aspirin)
  • T061, Therapeutic or Preventive Procedure (CABG, cardioversion)
  • T023, Body Part, Organ, or Organ Component (left ventricle, coronary artery)

4.2 Implementing UMLS Entity Linking with scispaCy

import spacy
import scispacy
from scispacy.linking import EntityLinker

def setup_umls_pipeline(model_name: str = "en_core_sci_md") -> spacy.language.Language:
    """
    Configure scispaCy pipeline with UMLS entity linking.

    The EntityLinker component compares extracted entity spans against
    a local copy of the UMLS knowledge base using approximate string
    matching and concept embeddings. It returns ranked candidate CUIs
    with confidence scores.

    IMPORTANT: The UMLS linker requires ~2GB RAM for the knowledge base
    and has a one-time download of ~1.5GB. Performance is substantially
    better with en_core_sci_lg than with the smaller models.

    Configuration parameters:
    - resolve_abbreviations: Use AbbreviationDetector before linking
    - linker_name: "umls" uses full UMLS; "mesh" uses MeSH only (smaller)
    - threshold: Minimum similarity score (0.7-0.85 recommended for clinical use)
    - max_entities_per_mention: Candidates per entity span

    Returns:
        Configured spaCy pipeline with UMLS linking
    """
    nlp = spacy.load(model_name)

    # Add abbreviation detection (must come before linker)
    nlp.add_pipe("abbreviation_detector")

    # Add UMLS entity linker
    nlp.add_pipe(
        "scispacy_linker",
        config={
            "resolve_abbreviations": True,       # Expand abbreviations before linking
            "linker_name": "umls",               # Full UMLS knowledge base
            "threshold": 0.80,                   # Minimum similarity threshold
            "max_entities_per_mention": 3,       # Top-k candidates
        }
    )

    return nlp

def extract_umls_entities(text: str, nlp) -> List[Dict]:
    """
    Extract entities with full UMLS annotations.

    For each recognized biomedical entity, this function retrieves:
    - The entity text as it appears in the document
    - The UMLS Concept Unique Identifier (CUI)
    - The canonical (preferred) name for the concept
    - The UMLS semantic type (T047 = Disease, T121 = Drug, etc.)
    - The similarity score (confidence of the linking)
    - The definition from UMLS (if available)

    Args:
        text: Clinical document text
        nlp: Pipeline with UMLS linker configured

    Returns:
        List of entities with complete UMLS annotations
    """
    doc = nlp(text)

    # Access the linker component for knowledge base queries
    linker = nlp.get_pipe("scispacy_linker")

    annotated_entities = []

    for ent in doc.ents:
        entity_info = {
            'text': ent.text,
            'start_char': ent.start_char,
            'end_char': ent.end_char,
            'label': ent.label_,
            'umls_candidates': []
        }

        # Get UMLS linking results
        for umls_entity in ent._.kb_ents:
            cui = umls_entity[0]         # UMLS Concept Unique Identifier
            score = umls_entity[1]       # Similarity score (0-1)

            # Query the knowledge base for this CUI
            if cui in linker.kb.cui_to_entity:
                kb_entity = linker.kb.cui_to_entity[cui]
                candidate = {
                    'cui': cui,
                    'canonical_name': kb_entity.canonical_name,
                    'aliases': list(kb_entity.aliases[:5]),        # First 5 aliases
                    'types': list(kb_entity.types),                # Semantic type codes
                    'definition': kb_entity.definition or "No definition available",
                    'similarity_score': round(score, 4)
                }
                entity_info['umls_candidates'].append(candidate)

        if entity_info['umls_candidates']:
            annotated_entities.append(entity_info)

    return annotated_entities

def format_umls_report(entities: List[Dict]) -> str:
    """
    Format UMLS entity linking results as a readable clinical report.

    Args:
        entities: Output from extract_umls_entities()

    Returns:
        Formatted string report
    """
    lines = ["UMLS Entity Linking Report", "=" * 60]

    # Semantic type descriptions for readability
    SEMANTIC_TYPES = {
        "T047": "Disease or Syndrome",
        "T121": "Pharmacologic Substance",
        "T061": "Therapeutic/Preventive Procedure",
        "T023": "Body Part, Organ, Component",
        "T116": "Amino Acid/Peptide/Protein",
        "T048": "Mental/Behavioral Dysfunction",
        "T033": "Finding",
        "T184": "Sign or Symptom",
        "T060": "Diagnostic Procedure",
    }

    for entity in entities:
        lines.append(f"\\nEntity: '{entity['text']}'")
        lines.append(f"  Character span: [{entity['start_char']}:{entity['end_char']}]")

        if entity['umls_candidates']:
            best = entity['umls_candidates'][0]  # Top candidate
            lines.append(f"  Best UMLS match:")
            lines.append(f"    CUI:{best['cui']}")
            lines.append(f"    Canonical name:{best['canonical_name']}")
            lines.append(f"    Similarity:{best['similarity_score']:.2%}")

            # Describe semantic types
            type_descriptions = [
                SEMANTIC_TYPES.get(t, t) for t in best['types']
            ]
            lines.append(f"    Semantic types:{', '.join(type_descriptions)}")

            if best['aliases']:
                lines.append(f"    Known aliases:{', '.join(best['aliases'][:3])}")

    return "\\n".join(lines)

# Example usage
# Note: requires ~2GB RAM for UMLS knowledge base
# nlp_umls = setup_umls_pipeline()
# entities = extract_umls_entities(DISCHARGE_SUMMARY, nlp_umls)
# print(format_umls_report(entities))

Expected output for key entities (illustrative, actual CUIs from UMLS 2024AB):

UMLS Entity Linking Report

Entity: ‘coronary artery disease’
Character span: [150:172]
Best UMLS match:
CUI: C0010068
Canonical name: Coronary Arteriosclerosis
Similarity: 96.80%
Semantic types: Disease or Syndrome
Known aliases: CAD, Ischemic Heart Disease, Coronary Artery Disease

Entity: ‘atrial fibrillation’
Character span: [478:496]
Best UMLS match:
CUI: C0004238
Canonical name: Atrial Fibrillation
Similarity: 99.10%
Semantic types: Disease or Syndrome, Finding
Known aliases: AF, A-fib, Auricular Fibrillation

Entity: ‘amiodarone’
Character span: [521:531]
Best UMLS match:
CUI: C0002598
Canonical name: Amiodarone
Similarity: 99.90%
Semantic types: Pharmacologic Substance
Known aliases: Cordarone, Nexterone, Pacerone

Entity: ‘median sternotomy’
Best UMLS match:
CUI: C0185792
Canonical name: Median Sternotomy
Similarity: 98.50%
Semantic types: Therapeutic/Preventive Procedure


5. Temporal Information Extraction

5.1 The Clinical Importance of Temporal Relations

In cardiac surgery, the timing of events carries weight comparable to the events themselves. “Atrial fibrillation on POD 3” and “atrial fibrillation on POD 14” may imply different etiologies, different risks, and different management approaches. A complete NLP pipeline therefore needs to capture not only what happened, but when.

spaCy’s dependency parser is well suited to extracting temporal relationships:

def extract_temporal_clinical_events(text: str, nlp) -> List[Dict]:
    """
    Extract clinical events with their temporal anchors.

    Uses dependency parsing to identify relationships between
    clinical events and their associated postoperative day (POD),
    time expressions, or date references.

    Temporal patterns in cardiac surgery discharge summaries:
    - Explicit POD references: "extubated on POD 1"
    - Implicit timing: "chest tubes removed subsequently"
    - Duration: "amiodarone for 6 weeks"
    - Relative timing: "48 hours after surgery"

    Args:
        text: Clinical narrative text
        nlp: scispaCy model

    Returns:
        List of events with temporal information
    """
    doc = nlp(text)

    # Pattern for "POD N", postoperative day
    pod_pattern = re.compile(r'POD\\s*(\\d+)', re.IGNORECASE)

    events_with_timing = []

    for sent in doc.sents:
        sent_text = sent.text

        # Find POD references in this sentence
        pod_matches = pod_pattern.findall(sent_text)
        pod_days = [int(d) for d in pod_matches]

        # Find clinical events (non-stop verbs and nouns)
        clinical_tokens = []
        for token in sent:
            if (token.pos_ in ("VERB", "NOUN", "PROPN") and
                not token.is_stop and
                len(token.text) > 2):
                clinical_tokens.append(token.lemma_)

        if pod_days and clinical_tokens:
            events_with_timing.append({
                'sentence': sent_text.strip(),
                'pod_days': pod_days,
                'earliest_pod': min(pod_days),
                'clinical_tokens': clinical_tokens[:8],  # Top 8 content words
            })

    # Sort by earliest POD
    events_with_timing.sort(key=lambda x: x['earliest_pod'])

    return events_with_timing

# Build clinical timeline from the postoperative course
postop_text = """
The patient was extubated on POD 1 without complications. Chest tubes were removed on POD 2.
Postoperative atrial fibrillation was noted on POD 3, successfully cardioverted with amiodarone
loading followed by maintenance dose. Patient remained in sinus rhythm thereafter.
Echocardiography on POD 5 showed LVEF 40% with good biventricular function.
"""

timeline = extract_temporal_clinical_events(postop_text, nlp_sci)

print("Clinical Timeline Reconstruction:")
print("=" * 60)
for event in timeline:
    print(f"\\n  POD{event['earliest_pod']}:")
    print(f"{event['sentence']}")

Output:

Clinical Timeline Reconstruction:
============================================================

  POD 1:
  The patient was extubated on POD 1 without complications.

  POD 2:
  Chest tubes were removed on POD 2.

  POD 3:
  Postoperative atrial fibrillation was noted on POD 3, successfully cardioverted with amiodarone loading.

  POD 5:
  Echocardiography on POD 5 showed LVEF 40% with good biventricular function.

This kind of structured temporal reconstruction has direct applications in automated quality metrics reporting, particularly for surgical safety indicators like time-to-extubation and time-to-complication onset.


6. Comparing spaCy/scispaCy with Regex: A Decision Framework

6.1 Quantitative Performance Considerations

Neither regex nor linguistic NLP is universally better. The optimal choice depends on the specific extraction task, the consistency of the source documents, and the computational resources you actually have at your disposal. The table below summarizes the practical considerations:

CriterionRegexspaCy/scispaCy
Setup complexityMinimalModerate (model installation)
Processing speedVery fast (~1ms/doc)Moderate (100ms–5s/doc)
Memory requirementsNegligible500MB–2GB
Accuracy for structured fieldsHigh (when patterns are known)Moderate
Accuracy for free-text narrativeLowHigh
Negation handlingManual, brittleBuilt-in via dependency parsing
Semantic normalizationNoneUMLS linking
Maintenance burdenHigh (pattern updates)Low (model updates)
ExplainabilityCompletePartial (neural models)
Regulatory complianceStraightforwardRequires validation

6.2 When to Use Regex

Regex is still the preferred approach for:

  • Highly structured fields with predictable formats: dates, MRN numbers, medication dosages in formatted lists
  • Real-time applications where processing latency is critical
  • Resource-constrained environments (edge computing, embedded systems)
  • Deterministic requirements where output must be fully explainable and reproducible
  • Simple extraction rules with few variations (e.g., “extract all numbers followed by ‘mg’”)

6.3 When to Use spaCy/scispaCy

Linguistic NLP earns its place when you need:

  • Free-text narratives with high syntactic variability (postoperative course sections)
  • Entity recognition where surface forms vary (“AF,” “atrial fib,” “atrial fibrillation”)
  • Relationship extraction between medical concepts (drug → indication, procedure → complication)
  • Negation detection at scale (well above what rule-based approaches can deliver)
  • Interoperability requirements where UMLS/SNOMED-CT standardization is needed
  • Research applications requiring semantic search or concept-based cohort selection

6.4 The Hybrid Pipeline (Preview)

In practice, the most robust systems combine both. A hybrid pipeline for cardiac surgery discharge summaries might follow this logic:

  1. Regex layer: Extract structured fields (dates, numeric values, section headers) with high precision
  2. scispaCy NER layer: Identify medical entities in free-text sections
  3. UMLS linking layer: Normalize entities to standard vocabularies
  4. Validation layer: Cross-validate regex and NLP results; flag discrepancies

The architecture is explored in depth in Article 3, where we layer in machine learning classifiers for complication severity scoring and LLM-based reasoning for the harder inference tasks.


7. Conclusion: From Syntax to Semantics

This article has shown how spaCy and scispaCy push medical text analysis past character-level pattern matching, into something closer to genuine linguistic intelligence. The key advances over regex-based extraction are:

Semantic awareness: scispaCy NER recognizes medical concepts regardless of their surface form, handling the terminological variability inherent to clinical language. “Triple vessel coronary artery disease,” “3-vessel CAD,” and “severe multivessel coronary disease” all refer to the same clinical condition.

Relational extraction: Dependency parsing exposes the grammatical relationships between tokens, enabling extraction of drug-event associations, procedural outcomes, and temporal sequences that regex simply cannot reach.

Ontological grounding: UMLS entity linking standardizes extracted concepts to universal identifiers, supporting interoperability across electronic health record systems, clinical registries, and research databases. Once a CUI is assigned to a concept, that annotation is unambiguous across institutions and across time.

Negation sensitivity: Medical NLP that cannot distinguish affirmed from negated findings is, plainly, clinically dangerous. The dependency-based negation detection illustrated here, while not exhaustive, handles the majority of common patterns in discharge documentation.

Limitations of the Linguistic Approach

scispaCy, biomedical training notwithstanding, is not without limitations that need to be acknowledged before any clinical use:

  • Out-of-vocabulary terms: Novel drug names, rare procedures, or institution-specific abbreviations may not be recognized
  • Domain shift: Models trained on PubMed abstracts may underperform on clinical notes, which have different linguistic characteristics
  • Negation complexity: Advanced negation patterns (“unlikely to represent,” “cannot exclude”) require specialized components (MedSpaCy, NegEx) beyond standard dependency parsing
  • Numerical relationship extraction: With “LVEF 35%”, scispaCy recognizes LVEF as an entity but does not inherently link the numerical value; that gap is best filled by regex or a post-processing layer
  • Validation requirements: Clinical deployment of NLP systems requires rigorous validation on institution-specific documents, with performance metrics (precision, recall, F1) computed on expert-annotated gold standards

8. Looking Ahead: Machine Learning and LLMs

The third and final article in this series will tackle what neither regex nor linguistic NLP can do: learn from data, handle unrestricted text variation, and perform complex clinical inference.

Machine learning classifiers trained on annotated discharge summaries can pick up complication risk from subtle linguistic patterns. Large Language Models (LLMs), the Claude API among them, can extract arbitrarily complex information through natural language querying, compare findings against clinical guidelines, and generate structured summaries from narrative text. The article will also address the non-negotiable topics: GDPR/HIPAA compliance, de-identification pipelines, and deployment considerations for production medical AI systems.

Preview: scispaCy might recognize “atrial fibrillation” as a disease entity. A properly prompted LLM, by contrast, can determine whether the documented management (amiodarone loading, cardioversion) was consistent with current ACC/AHA guidelines, a level of clinical reasoning that pattern-based NLP systems cannot reach.

References

  1. Neumann M, et al. “ScispaCy: Fast and Robust Models for Biomedical Natural Language Processing.” Proceedings of the 18th BioNLP Workshop and Shared Task, 2019:319–327.
  2. Honnibal M, Montani I. “spaCy 2: Natural language understanding with Bloom embeddings, convolutional neural networks and incremental parsing.” Unpublished, 2017.
  3. Bodenreider O. “The Unified Medical Language System (UMLS): integrating biomedical terminology.” Nucleic Acids Research 2004;32(suppl_1):D267–D270.
  4. Schwartz AS, Hearst MA. “A simple algorithm for identifying abbreviation definitions in biomedical text.” Proceedings of the Pacific Symposium on Biocomputing, 2003:451–462.
  5. Li J, et al. “BioCreative V CDR task corpus: a resource for chemical disease relation extraction.” Database 2016:baw068.
  6. Lample G, et al. “Neural Architectures for Named Entity Recognition.” Proceedings of NAACL-HLT 2016, 2016:260–270.
  7. Chapman WW, et al. “A simple algorithm for identifying negated findings and diseases in discharge summaries.” Journal of Biomedical Informatics 2001;34(5):301–310. [NegEx algorithm]
  8. Soldaini L, Goharian N. “QuickUMLS: a fast, unsupervised approach for medical concept extraction.” MedIR Workshop, SIGIR, 2016.
  9. Peng Y, et al. “Transfer Learning in Biomedical Natural Language Processing.” Proceedings of the 18th BioNLP Workshop, 2019.
  10. Savova GK, et al. “Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): architecture, component evaluation and applications.” JAMIA 2010;17(5):507–513.
  11. Uzuner O, et al. “2010 i2b2/VA challenge on concepts, assertions, and relations in clinical text.” JAMIA 2011;18(5):552–556.
  12. Stenetorp P, et al. “BRAT: a Web-based Tool for NLP-Assisted Text Annotation.” Proceedings of EACL 2012, 2012:102–107.
A doctor in a railwais stations in rural russia

Doctor Zhivago and the Reality of Medicine in Revolutionary Russia

Posted on May 4, 2026August 9, 2026 by Michele Danilo Pierri

Doctor Zhivago turns the daily work of a physician into a lens for what care becomes when institutions break. War medicine, infectious disease, rural scarcity. Through these threads Pasternak delivers a lesson that still cuts close to contemporary, technology-driven healthcare: clinical judgment and ethical independence carry the most weight precisely when the system around them collapses.


1. Introduction and author background

Doctor Zhivago is more than a historical or political novel. In its pages medicine works as an ethical and epistemological anchor in a world that is coming apart. Boris Pasternak (1890–1960) places a physician at the center of one of the most turbulent stretches of modern history: late Tsarist Russia, the Great War, the October Revolution, the Civil War.

Pasternak was no clinician. He moved, however, in an intellectual milieu shaped by science, philosophy, and a stubborn realism. His protagonist, Yuri Zhivago, is deliberately doctor and poet at once: two complementary ways of approaching reality, clinical observation on one side, humanistic interpretation on the other.

For a medical reader the choice matters. Zhivago is not a revolutionary leader, not a political theorist. He is a clinician forced to practice under extreme constraints, material, social, ethical. Through him Pasternak quietly puts a question on the table: what does medicine become when institutions fail and ideology tries to override biological reality?


2. A medical-oriented synopsis (without retelling the whole plot)

Rather than walking through the storyline, it is more useful (especially on a medical blog) to follow the clinically relevant trajectory of Yuri Zhivago.

Zhivago trains as a physician in pre-revolutionary Russia, in a period when medicine is moving from late 19th-century bedside empiricism toward an early scientific modernity. As war and revolution unfold he is repeatedly displaced. Urban hospitals, then wartime settings. Then academic life, then rural isolation. Across each shift, he keeps practicing, often under coercion and scarcity.

What he actually does is recognizable to any reader with a clinical background:

  • treating traumatic injuries and the complications of violence,
  • managing infectious diseases,
  • coping with malnutrition and exposure,
  • improvising care with whatever is at hand.

Medicine in Doctor Zhivago is never told as a story of progress or heroism. It is persistent, adaptive, frequently powerless. The physician survives not because events are controllable. He survives because suffering keeps surfacing, and someone still has to respond.


3. Medical scenes and clinical situations

3.1 War medicine and field care

The wartime sections expose Zhivago to ballistic trauma, amputations, sepsis, shortages of anesthesia and antiseptics. Clinical decisions are dictated by triage, not by optimization. There is no illusion of “best practice” here, only survival medicine. The psychological burden is constant. The physician operates under moral pressure, negotiating between what should be done and what can actually be done.

A particularly striking dimension is the strain placed on medical neutrality. The novel shows the erosion of the clinician’s protected role under armed groups and shifting authority. Instead of treating “neutral care” as a given, it documents how easily that principle breaks once coercion enters the clinical space.

3.2 Rural and isolated medical practice

Later, in remote rural areas (the Urals in particular), medicine reverts to its essentials:

  • clinical observation,
  • basic physical examination,
  • empirical decision-making.

No laboratory medicine. No imaging. A pharmacological arsenal that ranges from limited to symbolic. The doctor–patient relationship becomes direct and personal, stripped of institutional mediation. It is medicine that is pre-technological but not pre-scientific. The reasoning is still there. The tools are gone.

3.3 The body as counter-ideology

One of the strongest medical dimensions of the novel is the role of the suffering body. Pain, disease, hunger, death. They keep undoing political rhetoric. Ideological promises dissolve in front of biological vulnerability.

In this sense the physician is not merely a caregiver. He is a witness to an irreducible reality: the body does not obey slogans. The point is as clinical as it is philosophical.


Historical box (Russia 1900–1920): what “medicine under scarcity” realistically meant

  • Antisepsis existed, but implementation was uneven, especially outside major centers and during war.
  • No antibiotics: bacterial infections and wound complications could be lethal even when the “right” decision was made.
  • Anesthesia and analgesia were limited by supply chains, training, and infrastructure; surgery was possible but high-risk.
  • Diagnostics were mostly clinical: history and physical examination carried the weight, with few tests available.
  • Logistics were a clinical variable: transport disruption, cold exposure, malnutrition, and crowding changed outcomes as much as “medical skill”. </aside>

4. What Doctor Zhivago suggests about medicine of its time

4.1 Historical plausibility of clinical practice

The medicine described in the novel is broadly consistent with what was historically plausible in Russia between 1900 and 1920. Antisepsis is known, but applied unevenly. Antibiotics do not exist yet. Surgery is possible but risky. Diagnosis leans heavily on clinical reasoning. Zhivago sits at the intersection between classical bedside medicine and the first stirrings of modern clinical science.

4.2 Implicit medical ethics

Zhivago’s practice keeps pointing back to a few recurring principles:

  • individual responsibility,
  • moral resistance to ideological capture,
  • loyalty to the patient rather than to the institution.

In modern terms it anticipates a core ethical idea, namely that care must remain independent from political pressure. The novel does not romanticize the clinician as “pure”. It does insist, with some firmness, that medicine loses its meaning when it becomes only an instrument of power.

A caveat is worth stating. The text obviously does not provide a structured ethical framework, and reading too much theory into a literary work would be a stretch. Still, the orientation is clear enough.


5. Conclusion: why it still matters for technology-driven medicine

Read from today, with medicine increasingly shaped by technology, data, and algorithms, Doctor Zhivago offers a useful corrective:

  • medicine can function without technology,
  • it cannot function without judgment,
  • and it cannot survive if it becomes fully subordinated to ideology.

In an era of artificial intelligence, predictive models, and automated decision support, Zhivago stands for the irreducible human core of clinical practice. When systems fail, when data vanish, when infrastructures degrade, medicine returns to its most basic form: one human being interpreting another’s suffering.

For a blog focused on medical technology, the novel is a useful counterbalance. Progress matters. Clinical meaning still comes first.


References

  • “Between Killing and Curing: Doctors in Literary Depictions of the Russian Revolution and Civil War.” Synapsis: A Health Humanities Journal (Nov 26, 2024). https://medicalhealthhumanities.com/2024/11/26/between-killing-and-curing-doctors-in-literary-depictions-of-the-russian-revolution-and-civil-war/
  • “Tragic hero of the Russian Revolution.” Irish Medical Times (Apr 21, 2008). https://www.imt.ie/lifestyle/literature/tragic-hero-of-the-russian-revolution-21-04-2008/
  • “Doctor Zhivago | Novel, Themes & Russian Revolution.” Encyclopaedia Britannica (Feb 6, 2026). https://www.britannica.com/topic/Doctor-Zhivago-novel
A doctor tries to open a door to extraordinary tools.

OpenAI, Healthcare, and the Limits of Access

Posted on April 28, 2026August 16, 2026 by Michele Danilo Pierri

ChatGPT for Clinicians: A First Look from Italy

Article authored by Michele D. Pierri, MD

Cardiac Surgeon & Medical Technology Researcher

Last checked: April 28, 2026

Reading time: 7 minutes


OpenAI has rolled out a set of healthcare-oriented products built around ChatGPT. The headline reads, at first, like the launch of a single “medical GPT” available to physicians everywhere. The actual picture is more layered. What OpenAI is building is not one universal medical chatbot, but a differentiated healthcare ecosystem with distinct products for individual clinicians, healthcare organizations, and patients.

For individual physicians, the relevant product is ChatGPT for Clinicians. OpenAI describes it as a version of ChatGPT engineered to support clinical tasks: documentation, medical research, evidence review, clinical reasoning. The product was announced on April 22, 2026. Access, at present, is free for verified individual clinicians in the United States. Eligible users include physicians, nurse practitioners, physician assistants, and pharmacists.

Worth pausing here. ChatGPT for Clinicians includes access to advanced AI models for complex clinical questions, reusable clinical workflows organized as “skills,” trusted clinical search with citations, deep research across medical journals, and (in eligible situations) support for continuing medical education credits. OpenAI states that conversations within ChatGPT for Clinicians are not used to train its models. Optional HIPAA support may be available through a Business Associate Agreement for eligible accounts.

On one point OpenAI is explicit: ChatGPT for Clinicians is meant to support clinicians, not replace clinical judgment. The distinction is not cosmetic. A clinical AI assistant can retrieve evidence, summarize literature, draft documents, and structure reasoning. It cannot assume professional responsibility for diagnosis, treatment, informed consent, or patient safety.

The practical problem: access is currently U.S.-centered

For a physician practicing in Italy, the limiting factor surfaces almost immediately. According to OpenAI’s official documentation, signing up for ChatGPT for Clinicians requires a ChatGPT account, a valid National Provider Identifier, and a license verifiable through a third-party provider.

The National Provider Identifier (NPI) is a U.S. healthcare identifier. The Centers for Medicare & Medicaid Services define it as a unique 10-digit number used to identify healthcare providers in HIPAA standard transactions.

The result is a concrete barrier. An Italian physician registered with the Italian medical licensing system does not normally hold a U.S. NPI. The current self-service access route to ChatGPT for Clinicians is therefore not just difficult for a physician working in Italy. It is essentially not applicable.

A bit of a let-down.

Not because verification itself is unreasonable, of course. Quite the opposite: healthcare AI tools should require strict professional verification, regulatory safeguards, and clear accountability. The disappointment lies elsewhere. The first access route is built almost entirely around U.S. healthcare infrastructure. For clinicians outside the United States, including those of us working within European systems, the announcement reads as both exciting and frustrating. The product exists. The door is not yet really open.

OpenAI does mention plans to expand access to additional countries over time, with pilots for verified clinicians outside the United States in collaboration with the Better Evidence Network, where local regulations permit. As things stand at the time of writing, however, the individual access pathway remains U.S.-based.

ChatGPT for Healthcare: the institutional route

Separate from the clinician product is ChatGPT for Healthcare. This one is not primarily a self-service tool for individual doctors. It is an enterprise version of ChatGPT designed for healthcare organizations: clinicians, administrators, researchers. OpenAI describes it as a secure workspace with enterprise-grade governance, compliance controls, clinical search, centralized administration, and HIPAA support.

The distinction matters. ChatGPT for Clinicians targets individual verified clinicians in the United States. ChatGPT for Healthcare targets healthcare organizations that need scaled deployment, administrative controls, institutional data connections, and security governance.

For European hospitals, universities, or research centers, ChatGPT for Healthcare may eventually become the more realistic route. Eventually. In Europe, any clinical deployment would still need evaluation under GDPR, institutional privacy policies, professional responsibility standards, and local healthcare regulations. HIPAA is simply not the relevant legal framework in Italy. GDPR and national rules are.

ChatGPT Health: a different product for patients

OpenAI has also introduced ChatGPT Health, which should not be confused with ChatGPT for Clinicians. ChatGPT Health is designed for patients and general users. The company describes it as a dedicated health and wellness experience where users may connect medical records and wellness apps, with additional privacy protections. Per OpenAI, the product is meant to help people understand their health information and prepare for conversations with clinicians, not to provide diagnosis or treatment.

At launch, ChatGPT Health is available only to selected users. Not in the European Economic Area, Switzerland, or the United Kingdom.

The same broader pattern, again. OpenAI’s healthcare ecosystem is developing quickly, but access remains constrained by geography and regulation.

My access attempt from Italy

As an Italian cardiac surgeon, I had an obvious interest in testing whether ChatGPT for Clinicians could be accessed from Italy. The idea of a clinician-oriented AI workspace, with clinical search, structured evidence review, reusable workflows, and stronger safeguards, is directly relevant to current medical practice.

The limitation, as expected, surfaced before any real clinical testing could begin. OpenAI’s documentation states that the sign-up process requires a valid U.S. National Provider Identifier. Italian physicians do not normally have an NPI, so the current self-service access pathway cannot be completed by a physician practicing only in Italy.

Calling this a technical failure would be inaccurate. It is a structural access limitation. The system is built around U.S. professional verification, not around international medical licensing frameworks.

Step-by-step access check from Italy

The practical access check is straightforward:

  1. Visit the official ChatGPT for Clinicians page.
  2. Sign in with a ChatGPT account.
  3. Start the clinician verification process.
  4. Check whether the system asks for a U.S. NPI.
  5. Do not invent an NPI. Do not use another clinician’s identifier.
  6. If the process requires an NPI, stop the attempt and document that the current pathway is not applicable to Italian physicians.
  7. Do not upload patient data, clinical letters, reports, or identifiable medical information during the process.

That last point is non-negotiable. Until there is a clear legal, contractual, and institutional framework in place, real patient data should not be used for experimental access attempts.

What this means for European clinicians

For now, European clinicians have three more realistic options.

First, standard ChatGPT can be used for non-identifiable educational, scientific, and editorial work: literature summaries, teaching material, coding assistance, academic writing support, conceptual discussion.

Second, prototypes can be built through the OpenAI API using anonymized, synthetic, or non-identifiable data. Particularly relevant for research workflows, medical education, RAG systems, ICD/DRG coding tools, and clinical documentation experiments.

Third, institutions may explore ChatGPT for Healthcare as an enterprise route. Only after formal evaluation of privacy, governance, data processing agreements, GDPR compliance, clinical safety, and professional responsibility, however.

Conclusion

OpenAI’s healthcare products represent a meaningful step toward specialized AI tools for medicine. ChatGPT for Clinicians is particularly interesting because it acknowledges what clinical work actually is: not generic text generation, but a domain that requires evidence, citations, structured workflows, safety constraints, and professional accountability.

The current access model, though, reveals a clear asymmetry. U.S. clinicians are the first target population. Physicians in Italy and other European countries remain outside the self-service pathway.

Speaking as a clinician, I find the situation both promising and frustrating. Promising, because the direction is clearly the right one: medical AI tools need dedicated environments. Frustrating, because international clinicians who are already engaging seriously with AI, research, documentation, and healthcare innovation cannot yet access the clinician-specific product through their own professional credentials.

The crucial question is no longer whether AI will enter clinical practice. That question has already been answered. The real question is whether access to safe, verified, and professionally governed medical AI tools will expand beyond the U.S. in a way that respects the regulatory, ethical, and clinical realities of European healthcare systems.

LINKS

Official access page: ChatGPT for Clinicians

OpenAI Help Center page: ChatGPT for Clinicians – Help Center

A researcher observes a brain in a cage

Local AI: The LM Studio Surprise

Posted on April 26, 2026August 16, 2026 by Michele Danilo Pierri

Part 4: the conclusion of the “Local AI” series — The LM Studio surprise: same model, same hardware, 40-200x faster


In Part 3, we pushed gemma4:e4b through four demanding tests. The results were impressive — excellent clinical reasoning, comprehensive summaries, working code. But the wait times were brutal: 5 to 16 minutes per response on our CPU-only hardware.

I accepted this as the cost of running larger models locally. Then I tried LM Studio.

Same model. Same hardware. Same prompts.

TaskOllama + ChatboxLM StudioSpeedup
Coding5 min 30s6.9 seconds48x
Reasoning5 min 28s4.4 seconds75x
Summarization8 min 12s38.8 seconds13x
Differential Dx16 min 1s5.0 seconds192x

This isn’t a typo. The clinical differential diagnosis that took 16 minutes on Ollama completed in 5 seconds on LM Studio. Same gemma4:e4b model, same i7 laptop, same prompt.

This article explains what happened, why it matters, and when to use each tool.


What Is LM Studio?

LM Studio is a desktop application for running local LLMs. Unlike Ollama (which is primarily a backend/CLI tool that other interfaces connect to), LM Studio bundles everything together: model download, runtime engine, and chat interface in one application.

Download: lmstudio.ai

Key differences from Ollama:

AspectOllamaLM Studio
ArchitectureBackend service + separate UIAll-in-one application
Model formatOllama-specificStandard GGUF files
Model sourceOllama libraryHugging Face / direct download
APIOllama APIOpenAI-compatible API
ConfigurationLimitedExtensive parameters

The tradeoff: Ollama is simpler and has a larger ecosystem of compatible interfaces. LM Studio offers more control and, as we discovered, dramatically better performance in certain scenarios.


The Test Setup

To ensure a fair comparison, I used:

  • Model: gemma-4-E4B-it-GGUF (Q4_K_M quantization)
  • Hardware: Intel i7-1165G7, 20 GB RAM, no dedicated GPU
  • GPU Offload: Disabled (set to 0) — pure CPU inference
  • Prompts: Identical to Part 3

The model file is the same architecture and quantization level as Ollama’s gemma4:e4b. The only variable is the runtime software.


Test Results

Task 1: Coding

Prompt: Write a Python function that calculates average heart rate, excluding outliers, and returns min/max values.

MetricLM StudioOllama (Chatbox)
Time6.91 seconds5 min 30s
Tokens1,0032,089
Speed5.01 tok/s6.49 tok/s
Speedup48x faster—

The output quality was identical — professional code with type hints, edge case handling, and test examples. But LM Studio produced it in 7 seconds instead of 5.5 minutes.

Interestingly, LM Studio’s response was also more concise (1,003 tokens vs 2,089) while covering the same content. Less verbose, equally complete.


Task 2: Reasoning

Prompt: Hospital ward allocation problem (12 patients, room constraints, isolation requirements).

MetricLM StudioOllama (Chatbox)
Time4.39 seconds5 min 28s
Tokens2863,039
Answer✅ Correct (7 rooms)✅ Correct (7 rooms)
Speedup75x faster—

Both produced the correct answer with clear step-by-step reasoning. LM Studio did it in 4.4 seconds with 286 tokens. Ollama took 5.5 minutes and generated 3,039 tokens — over 10x more verbose for the same conclusion.


Task 3: Medical Summarization

Prompt: Summarize a CABG discharge letter highlighting diagnosis, procedure, complications, medication changes, and follow-up.

MetricLM StudioOllama (Chatbox)
Time38.83 seconds8 min 12s
Tokens2887,559
Speedup13x faster—

This was the most clinically impressive result. LM Studio’s summary was not just faster and shorter — it was more complete:

LM Studio captured:

  • ✅ All diagnoses
  • ✅ Detailed procedure (LIMA→LAD, SVG→diagonal, SVG→PDA)
  • ✅ Pneumothorax complication
  • ✅ Anemia with actual lab values (Hgb 7.9, Hct 21.9)
  • ✅ Specific medication changes (Metoprolol 50mg BID)
  • ✅ Complete follow-up instructions with timeframes

Ollama (Chatbox) missed:

  • ❌ The anemia lab values
  • ❌ Specific medication dosages

In 39 seconds and 288 tokens, LM Studio produced a more clinically useful summary than Ollama did in 8 minutes with 7,559 tokens. This was the result I found hardest to believe — and replicated multiple times to confirm.


Task 4: Differential Diagnosis

Prompt: 58-year-old male with STEMI presentation — generate differential diagnoses and immediate workup.

MetricLM StudioOllama (Chatbox)
Time4.98 seconds16 min 1s
Tokens8859,134
Speedup192x faster—

The response was excellent — STEMI correctly identified as top diagnosis with anatomical localization (anterior wall from V1-V4 distribution), aortic dissection included as critical rule-out, comprehensive workup with serial troponins and ECGs, “time is muscle” urgency conveyed.

One caveat: the response was truncated due to LM Studio’s default context length settings. This is configurable — increasing the context window in settings resolves it. Even truncated, the core clinical content was complete.


The Complete Comparison

TaskLM StudioOllama + ChatboxOllama + Open WebUI
Coding6.9 sec4-5 min5m 30s
Reasoning4.4 sec3 min5m 28s
Summarization38.8 sec5-6 min8m 12s
Differential Dx5.0 sec5-6 min16m 1s

LM Studio is 40-200x faster than Ollama-based interfaces while producing equal or better quality output.


Why Is LM Studio So Much Faster?

I don’t have definitive answers, but here are the likely factors:

1. No “Thinking” Overhead

On Ollama, gemma4:e4b shows an explicit “thinking” phase — visible chain-of-thought reasoning that takes 2-3 minutes before the actual response begins. LM Studio appears to skip or internalize this step, producing direct responses.

RuntimeBehaviorImpact
Ollama“Thinking…” visible for 2-3 minAdds minutes to every response
LM StudioDirect response, no visible thinkingNear-instant start

This may be due to different system prompts or model configuration. The output quality suggests the reasoning still happens — it’s just not displayed.

2. Architectural Differences

Ollama has more layers between you and the model:

User → Interface (Chatbox) → Ollama API → Ollama Server → llama.cpp → Model

LM Studio is more direct:

User → LM Studio → llama.cpp → Model

Fewer layers means less overhead, especially for the coordination between components.

3. Different llama.cpp Configuration

Both tools use llama.cpp as the underlying inference engine, but likely with different default parameters:

  • Batch size: LM Studio may use larger batches for more efficient processing
  • Thread allocation: Different CPU utilization strategies
  • Memory management: More efficient context caching

4. Response Verbosity

LM Studio consistently produced more concise responses — 288 tokens vs 7,559 for the same summarization task. Generating fewer tokens means faster completion, even at the same tokens-per-second rate.


LM Studio Setup

Installation

  1. Download from lmstudio.ai
  2. Run the installer (Windows, Mac, or Linux)
  3. Open LM Studio
LM Studio download page

Downloading a Model

Unlike Ollama, LM Studio downloads models directly from Hugging Face:

  1. Click the search icon (magnifying glass) in the left sidebar
  2. Search for “gemma-4-e4b”
  3. Select a GGUF version (I used Q4_K_M for balance of quality and size)
  4. Click Download

The model downloads to LM Studio’s local storage — separate from any Ollama models you may have.

Running the Model

  1. Click the chat icon in the left sidebar
  2. Select your downloaded model from the dropdown
  3. Start chatting

Configuring for CPU-Only

If you have integrated graphics (like our Intel Iris Xe), ensure GPU offload is disabled:

  1. Click the sliders icon (settings) in the chat panel
  2. Set “GPU Offload” to 0
  3. Adjust context length if needed (default may truncate long responses)

Bonus Feature: OpenAI-Compatible API

LM Studio includes a built-in API server that’s compatible with the OpenAI client library. This enables powerful use cases:

Starting the Server

  1. Click the “Local Server” icon in the left sidebar
  2. Click “Start Server”
  3. Note the endpoint (default: http://localhost:1234)

Connecting from Code

import openai

client = openai.OpenAI(
    base_url="<http://localhost:1234/v1>",
    api_key="not-needed"  # LM Studio doesn't require an API key
)

response = client.chat.completions.create(
    model="gemma-4-E4B-it",
    messages=[{"role": "user", "content": "Hello!"}]
)

print(response.choices[0].message.content)

Practical Applications

  • Access from mobile: Query your local model from a phone or tablet on the same network
  • IDE integration: Connect VS Code, Cursor, or other tools that support OpenAI-compatible endpoints
  • Application development: Build apps that use local LLM inference
  • Multi-device: Keep the model running on a powerful desktop, query from a laptop

For a physician, this means: run LM Studio on your office workstation, access the model from a tablet on the ward. Complete privacy, no cloud, no subscription.


When to Use What

After all this testing, here’s my recommendation:

Use Ollama + Chatbox When:

  • You want the simplest possible setup
  • You’re using small models (gemma2:2b, phi3:mini) where speed is already adequate
  • You need multiple interfaces (Open WebUI for web access, Chatbox for desktop)
  • You’re building applications that specifically require the Ollama API

Use LM Studio When:

  • You’re running larger models (4B+ parameters) on CPU
  • Response time matters
  • You want OpenAI-compatible API access
  • You need fine-grained control over model parameters
  • You’re willing to manage separate model downloads

The Honest Tradeoff

LM Studio requires downloading models separately from Ollama. If you’re already invested in the Ollama ecosystem, this means duplicate storage. On my system:

  • Ollama’s gemma4:e4b: ~3 GB
  • LM Studio’s gemma-4-E4B-it-GGUF: ~3 GB
  • Total: ~6 GB for the “same” model

For the 40-200x speed improvement, I consider this worthwhile. Your mileage may vary.


Future Directions

The LM Studio results open interesting possibilities:

Local LLMs as Agent Backends

With 5-second response times instead of 5-minute waits, using gemma4:e4b in multi-agent systems becomes feasible. A workflow with 10 LLM calls:

  • Ollama: 10 × 5 min = 50 minutes
  • LM Studio: 10 × 30 sec = 5 minutes

This is still slower than cloud APIs, but usable for privacy-critical applications.

Hybrid Architectures

A practical approach: use local LLMs (via LM Studio) for simple queries and data processing, escalate to cloud APIs (Claude, GPT-4) for complex reasoning. Best of both worlds — privacy for routine tasks, capability when needed.

Clinical Decision Support

The summarization and differential diagnosis results suggest gemma4:e4b could support clinical workflows:

  • Summarize patient histories before rounds
  • Generate differential diagnoses for educational review
  • Extract key information from discharge letters

Always with physician oversight — but as an assistant rather than a bottleneck.


Series Conclusion

Four articles ago, we started with a question: can you run useful LLMs on a standard professional laptop with no dedicated GPU?

The answer is yes — with caveats:

  1. Hardware matters less than expected: Our i7 + 20GB RAM + integrated graphics handled models up to 4B parameters comfortably
  2. Model choice is critical: gemma2:2b for speed, gemma4:e4b for quality — pick based on your needs
  3. Software choice is even more critical: The same model runs 40-200x faster on LM Studio than Ollama. This was the biggest surprise of the entire project
  4. Local LLMs are ready for real work: Coding assistance, document summarization, clinical reasoning support — all feasible on consumer hardware
  5. They’re not replacements for cloud AI: Complex multi-step reasoning, very long contexts, and cutting-edge capabilities still favor Claude, GPT-4, and similar services

The practical takeaway: install LM Studio, download gemma4:e4b, and you have a capable, private, offline AI assistant — on the laptop you already own.


Quick Start Summary

For readers who want to skip to the end:

  1. Download LM Studio from lmstudio.ai
  2. Search and download gemma-4-E4B-it-GGUF (Q4_K_M version)
  3. Set GPU Offload to 0 if you don’t have a dedicated GPU
  4. Start chatting — expect responses in seconds, not minutes

That’s it. Local AI on consumer hardware, without the wait.


This concludes the “Local AI” series. All tests were conducted on an Intel i7-1165G7 laptop with 20 GB RAM and no dedicated GPU.

Previous articles:

  • *Part 1: Running LLMs Without Dedicated Graphics — Setup and installation*
  • *Part 2: Choosing the Right Model — Benchmarks and model selection*
  • *Part 3: Stress-Testing on Real Tasks — Coding, reasoning, and medical applications*

  • Previous
  • 1
  • 2
  • 3
  • 4
  • 5
  • …
  • 8
  • Next
© 2024–2026 micheledpierri.com · Privacy Policy · Impressum