- Design
- Systematic review and random-effects meta-analysis
- Population
- 34 studies of ML/DL sepsis prediction in hospitalised adults, mostly retrospective
- Primary outcome
- Discrimination (AUROC) for sepsis prediction
- Effect
- Pooled AUROC 0.913 (95% CI 0.887–0.933); 95% prediction interval 0.660–0.983
This systematic review and meta-analysis included 34 studies of machine learning and deep learning models predicting sepsis in hospitalised adults. Most were retrospective, and several reused the same public intensive care datasets (MIMIC and PhysioNet), so the studies were not independent.
The pooled area under the curve was 0.913 (95% CI 0.887 to 0.933), but the 95% prediction interval, which estimates performance in a new setting, ran from 0.660 to 0.983. Performance was similar for models predicting sepsis within four hours (0.894) and further ahead (0.926). Few prospective or randomised studies exist.
The gap between the confidence interval and the prediction interval is the message: a model that performs well in one hospital's data may perform only modestly in yours. Any AI alert introduced locally needs local validation and monitoring of alert fatigue.
- Do not assume a sepsis prediction tool will perform in your unit as it did in published studies.
- Ask vendors for external and prospective validation, not only development data.
- Monitor false alarms after introduction; alert fatigue can cancel any benefit.
- Clinical screening with NEWS2 (or your local early warning score) and lactate remains the baseline.
Why it matters
Hospitals are buying these tools now, and the published figures overstate how they are likely to perform.
Don't overread it
Most included studies were retrospective and shared datasets; the pooled AUC is not an estimate of real-world performance.
The statistics, in plain English
The confidence interval (0.887 to 0.933) describes the average across studies; the prediction interval (0.660 to 0.983) describes what to expect in a new setting, and it is much wider. An AUC of 0.66 would be only modestly useful.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for emergency & critical care, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free