- Design
- systematic review and random-effects meta-analysis (Hartung-Knapp-Sidik-Jonkman) with prediction intervals and dataset-overlap sensitivity analysis
- Population
- 34 studies of machine learning and deep learning sepsis prediction models in hospitalised adults, mostly retrospective, with overlapping public cohorts
- Primary outcome
- pooled area under the receiver operating characteristic curve
- Effect
- pooled AUROC 0.913 (95% CI 0.887-0.933) with 95% prediction interval 0.660-0.983; within 4 h 0.894, beyond 4 h 0.926, intervals overlapping
A systematic review pooled 34 studies of machine learning and deep learning models predicting sepsis in hospitalised adults, deliberately reporting prediction intervals alongside confidence intervals. The pooled area under the receiver operating characteristic curve was 0.913 (95% CI 0.887-0.933) — the figure that gets quoted. The 95% prediction interval was 0.660 to 0.983.
That second number is the one that matters for a unit considering buying such a tool. A confidence interval describes how precisely the average performance across these studies is known; a prediction interval describes what the next hospital should expect. An interval reaching down to 0.66 means a model performing barely better than a coin-weighted guess is entirely compatible with this evidence. Subgroup estimates by prediction window overlapped completely (0.894 within 4 hours, 0.926 beyond 4 hours), establishing no advantage to either horizon.
The review also names the structural problem: several of the 34 studies reused MIMIC- or PhysioNet-derived cohorts, so they are not 34 independent patient populations, and most were retrospective model-development work rather than prospective evaluation. Nothing here says these models do not work. It says the published evidence cannot tell you whether one will work in your hospital, and that the only thing which can is a prospective evaluation in your own population.
- Ask any vendor for prospective performance in a population like yours, not a pooled AUROC
- Ask which dataset the model was developed on — MIMIC and PhysioNet appear repeatedly
- Discrimination is not utility: a model can rank patients well and still change no decision
- Plan the alert workflow before the model — who receives it, what they are expected to do, and what happens when they ignore it
- Audit your own alert performance after deployment; published figures will not transfer
Why it matters
It gives a specific question to ask a vendor, and a reason to distrust the single number on the brochure.
Don't overread it
Poor generalisability in the published literature is not evidence that these models fail in practice — it is evidence that nobody has shown they succeed outside their development data.
The statistics, in plain English
The gap between a confidence interval of 0.887-0.933 and a prediction interval of 0.660-0.983 is the whole finding. The first says the average is well estimated; the second says individual settings vary enormously. Reviews that report only the first make heterogeneous evidence look settled — which is exactly what the authors are warning against.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for emergency & critical care, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free