- Design
- Retrospective deep learning development study, two cohorts, no external validation
- Population
- 122 patients (207 eyes) with intermediate or converting age-related macular degeneration, mean age 82 to 84
- Primary outcome
- Agreement with clinician decisions to start anti-VEGF treatment
- Effect
- AUROC 0.73 (95% CI 0.59 to 0.87), sensitivity 0.69 (0.54 to 0.83), specificity 0.92 (0.84 to 1.00)
A deep learning framework was trained on serial optical coherence tomography and visual acuity from 122 patients and 207 eyes, some with intermediate dry age-related macular degeneration that never converted and some that progressed to neovascular disease and were treated. One task was to flag the point at which anti-VEGF treatment was needed; the other was to estimate outer retinal thickness from Bruch's membrane to the outer plexiform layer as a biomarker.
Against the clinicians' own decisions the model reached an area under the receiver operating curve of 0.73 (95% CI 0.59 to 0.87) and an area under the precision-recall curve of 0.51 (0.35 to 0.66). Accuracy was 0.87 (0.76 to 0.97) and specificity 0.92 (0.84 to 1.00), but sensitivity was 0.69 (0.54 to 0.83). Outer retinal thickness estimation gave a normalised mean absolute error of 0.12.
Read the sensitivity, not the accuracy. In a cohort where most visits do not need an injection, a model can look accurate while missing three eyes in ten that did need one, and a missed conversion in neovascular disease is exactly the error that costs vision permanently. This is research-stage work, useful as a prompt for review rather than as permission to extend an interval, and the cohort is small enough that all the intervals are wide.
- Judge any conversion-detection model on sensitivity and on the width of its interval, not on headline accuracy.
- Treat automated flags as a reason to look at the scan yourself, never as a reason to defer review.
- Note that the comparator here was clinician decisions, not an independent ground truth of conversion.
- Ask for external validation on a different scanner and population before changing any monitoring interval.
- Keep counselling patients on intermediate disease to report distortion promptly; that remains the most sensitive detector available.
The statistics, in plain English
An area under the curve of 0.73 sits between coin-toss and useful, and its interval reaching down to 0.59 means the data cannot exclude near-chance performance. High specificity with modest sensitivity in a mostly untreated cohort inflates overall accuracy while leaving the clinically dangerous errors - the missed conversions - undercounted. With 207 eyes, every estimate here should be read as provisional.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for ophthalmology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free