- Design
- systematic review and random-effects meta-analysis with subgroup and dataset-overlap analysis
- Population
- 34 studies of machine-learning and deep-learning sepsis prediction models in hospitalised adults
- Primary outcome
- pooled area under the receiver operating characteristic curve for sepsis prediction
- Effect
- pooled AUROC 0.913 (95% CI 0.887–0.933), 95% prediction interval 0.660–0.983
Thirty-four studies of machine-learning and deep-learning sepsis prediction in hospitalised adults were pooled, most of them retrospective model development or validation work. The headline figure is a pooled area under the receiver operating characteristic curve of 0.913 (95% CI 0.887 to 0.933) — the sort of number that gets a product bought.
The 95% prediction interval is 0.660 to 0.983. That is the number that matters, because it describes what performance to expect in a new setting rather than in the studies already done. The authors also found that several of the 34 reports reused MIMIC- or PhysioNet-derived cohorts, so they do not represent 34 independent patient populations. Subgroup estimates by prediction window — 0.894 for within four hours, 0.926 for more than four hours ahead, 0.858 where the window was unclear — had overlapping intervals and established no advantage for any horizon.
The practical reading for anyone being offered a sepsis early-warning tool is to ask two questions: which dataset was it developed on, and has it been evaluated prospectively in a hospital like this one. The answer to the second is usually no. A separate Korean claims analysis in the same paper described longer stays and higher costs for sepsis-related episodes; the authors are explicit that it is context, not validation.
- Ask which cohort a sepsis algorithm was developed and validated on before adopting it
- Treat a reported AUROC from retrospective data as a ceiling, not an expectation
- Require local prospective evaluation with defined alert thresholds and an escalation pathway
- Do not let an algorithm displace bedside review — none of these models has been shown to improve outcomes
Why it matters
A pooled AUROC from retrospective data says nothing about how a tool will behave in your hospital.
Don't overread it
This synthesises model discrimination only — no included study showed that using one improves patient outcomes.
The statistics, in plain English
A confidence interval describes uncertainty about the average across studies; a prediction interval describes what a new study would likely find. Here they diverge sharply — 0.887 to 0.933 against 0.660 to 0.983 — which means the pooled average is precise but the next hospital's experience is not predictable from it. Reuse of the same public datasets inflates apparent consistency, because the same patients appear repeatedly.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for infectious diseases, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free