DailyDoctor Archive Specialties Get app
Back to the 9 September 2026 edition

Research · 04 of 06

A model that decides when to start anti-VEGF is not yet good enough to let one wait

This automated treat-now model matched clinicians only moderately, with sensitivity 0.69, so it is a prompt to review a scan rather than a basis for lengthening monitoring intervals.

Design
Retrospective deep learning development study, two cohorts, no external validation
Population
122 patients (207 eyes) with intermediate or converting age-related macular degeneration, mean age 82 to 84
Primary outcome
Agreement with clinician decisions to start anti-VEGF treatment
Effect
AUROC 0.73 (95% CI 0.59 to 0.87), sensitivity 0.69 (0.54 to 0.83), specificity 0.92 (0.84 to 1.00)

A deep learning framework was trained on serial optical coherence tomography and visual acuity from 122 patients and 207 eyes, some with intermediate dry age-related macular degeneration that never converted and some that progressed to neovascular disease and were treated. One task was to flag the point at which anti-VEGF treatment was needed; the other was to estimate outer retinal thickness from Bruch's membrane to the outer plexiform layer as a biomarker.

Against the clinicians' own decisions the model reached an area under the receiver operating curve of 0.73 (95% CI 0.59 to 0.87) and an area under the precision-recall curve of 0.51 (0.35 to 0.66). Accuracy was 0.87 (0.76 to 0.97) and specificity 0.92 (0.84 to 1.00), but sensitivity was 0.69 (0.54 to 0.83). Outer retinal thickness estimation gave a normalised mean absolute error of 0.12.

Read the sensitivity, not the accuracy. In a cohort where most visits do not need an injection, a model can look accurate while missing three eyes in ten that did need one, and a missed conversion in neovascular disease is exactly the error that costs vision permanently. This is research-stage work, useful as a prompt for review rather than as permission to extend an interval, and the cohort is small enough that all the intervals are wide.

  • Judge any conversion-detection model on sensitivity and on the width of its interval, not on headline accuracy.
  • Treat automated flags as a reason to look at the scan yourself, never as a reason to defer review.
  • Note that the comparator here was clinician decisions, not an independent ground truth of conversion.
  • Ask for external validation on a different scanner and population before changing any monitoring interval.
  • Keep counselling patients on intermediate disease to report distortion promptly; that remains the most sensitive detector available.

The statistics, in plain English

An area under the curve of 0.73 sits between coin-toss and useful, and its interval reaching down to 0.59 means the data cannot exclude near-chance performance. High specificity with modest sensitivity in a mostly untreated cohort inflates overall accuracy while leaving the clinically dangerous errors - the missed conversions - undercounted. With 207 eyes, every estimate here should be read as provisional.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

paedophthcornearetinacataractsurg

Tomorrow morning, before your first patient

One edition a day for ophthalmology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app