DailyDoctor Archive Specialties Get app
Back to the 15 September 2026 edition

Clinical update · 01 of 05

AI sleep apnoea screening: read the prediction interval, not the confidence interval

Ask for the prediction interval and local validation data before adopting any AI sleep apnoea screening tool; the pooled figures do not predict your clinic's performance.

Design
systematic review and bivariate random-effects diagnostic meta-analysis, QUADAS-2 and GRADE assessed
Population
60 studies of adults with suspected obstructive sleep apnoea or population cohorts; 47 in meta-analysis
Primary outcome
sensitivity and specificity against polysomnography at AHI thresholds of 5, 15 and 30 per hour
Effect
sensitivity 0.94/0.87/0.83, specificity 0.77/0.81/0.91; prediction interval for specificity at AHI 5 was 0.30-0.96

Sleep apnoea is common and substantially underdiagnosed, and polysomnography capacity is the bottleneck everywhere. A systematic review searched four databases for studies of artificial-intelligence screening models validated against polysomnography, including 60 studies with 47 in the meta-analysis, and synthesised accuracy separately at apnoea-hypopnoea index thresholds of 5, 15 and 30 events per hour using bivariate random-effects models.

Pooled sensitivity was 0.94 (95% CI 0.92 to 0.96) at a threshold of 5, 0.87 at 15 and 0.83 at 30. Specificity moved the other way: 0.77, 0.81 and 0.91. Models built from polysomnography-derived signals outperformed those using other inputs at every threshold - sensitivity 0.96 against 0.92 at a threshold of 5, specificity 0.82 against 0.70 - which is unsurprising and matters, because the tools intended to reduce demand for sleep studies are precisely the ones that do not use sleep study signals.

The number that should govern the decision is the prediction interval. For specificity at a threshold of 5, the 95% confidence interval was 0.69 to 0.84, but the prediction interval ran from 0.30 to 0.96. The confidence interval describes how well the average is estimated; the prediction interval describes what a new setting should expect. A tool that might deliver 30% specificity in your population will generate three false positives for every true one, and the sleep laboratory it was bought to unclog will fill with them. Certainty of evidence was low or very low.

  • Do not procure an AI screening tool on pooled sensitivity and specificity alone; ask for the prediction interval.
  • Insist on local validation against polysomnography before a tool is allowed to prioritise referrals.
  • Note that the tools needing no sleep study performed worst - that is the trade being sold.
  • Match the threshold to purpose: high sensitivity at an index of 5 for ruling out, higher specificity at 30 for prioritising.
  • Where sleep laboratory capacity is scarce, a poorly specific screen consumes the capacity it was meant to protect.

Why it matters

These tools are being sold on pooled accuracy figures that the review's own prediction intervals say will not transfer to an individual service.

Don't overread it

Certainty of evidence was graded low to very low with limited external validation - this is not evidence that any specific product works in practice.

The statistics, in plain English

A confidence interval answers how precisely the average across studies has been measured; a prediction interval answers what the next study or setting is likely to see. Here they diverge enormously - specificity at a threshold of 5 is 0.69 to 0.84 by confidence interval but 0.30 to 0.96 by prediction interval. That gap is a direct measure of between-study heterogeneity, and it means the pooled number describes no particular population. A summary area under the curve of 0.94 sits on top of that same heterogeneity and should not reassure.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

sleeplungcancer

Tomorrow morning, before your first patient

One edition a day for pulmonology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app