- Design
- Systematic review and Bayesian bivariate diagnostic meta-analysis
- Population
- 13 studies, 9,983 adults tested against conventional sleep studies
- Primary outcome
- Diagnostic accuracy of PPG-based AI for obstructive sleep apnoea
- Effect
- Sensitivity 79.6% (95% CrI 55.5–93.8%); specificity 76.5% (48.2–94.0%)
This systematic review and Bayesian meta-analysis pooled 13 studies of 9,983 adults in which artificial intelligence models read photoplethysmography (PPG), the optical pulse signal recorded by smartwatches, rings and pulse oximeters, and were compared with conventional sleep testing.
Pooled sensitivity was 79.6% (95% credible interval 55.5% to 93.8%) and specificity 76.5% (48.2% to 94.0%). At higher apnoea-hypopnoea index thresholds, specificity rose (85.1% for AHI ≥30) and sensitivity fell (87.2% at AHI ≥5 to 76.7% at AHI ≥30). Deep learning models were more specific (82.9%) than older machine learning (63.6%). Evidence quality was rated moderate.
Patients increasingly arrive with a watch report suggesting sleep apnoea. At these figures, a positive result is a reason to take a proper history and arrange sleep testing, not a diagnosis; a negative result does not exclude apnoea in a patient with typical symptoms.
- Treat a wearable 'sleep apnoea' alert as a prompt for history and validated testing, not as a diagnosis.
- A negative wearable result does not rule out sleep apnoea when snoring, witnessed apnoeas or daytime sleepiness are present.
- Ask about drowsy driving in anyone with suspected sleep apnoea.
- Where polysomnography is scarce, home sleep apnoea testing is the usual next step.
Why it matters
Patients are arriving with device reports, and the GP is the one who decides what they mean.
Don't overread it
These models were tested in research settings; accuracy of consumer devices in everyday use may be lower.
The statistics, in plain English
Sensitivity of 80% means about 1 in 5 people with apnoea would be missed; specificity of 77% means about 1 in 4 without it would be flagged. The credible intervals are very wide (roughly 50% to 94%), so accuracy in any particular device or population could be much better or much worse.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for family medicine, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free