- Design
- retrospective crossover reader study using an anatomically derived quantitative zone reference map on deep learning-registered RetCam panoramic images
- Population
- 85 retinopathy of prematurity examinations assessed by four ophthalmologists with and without the reference map
- Primary outcome
- reader-reference and interobserver agreement (kappa) for zone classification and zone-dependent ETROP treatment decisions
- Effect
- zone classification reader-reference kappa 0.58 to 0.92, interobserver 0.47 to 0.86; Zone III vs non-Zone III 0.40 to 0.88
Retinopathy of prematurity zone classification decides who gets treated under ETROP criteria, and it is done by eye, on a wide-field image, with the peripheral retina the hardest part to judge. This group built an anatomically derived quantitative zone reference map from optic disc-to-fovea distance and multimodal anatomical data, applied it to RetCam panoramic images assembled by a deep learning registration pipeline, and tested it in a crossover reader study — four ophthalmologists reading 85 examinations with and without the map.
Agreement improved substantially. Reader-to-reference agreement for zone classification rose from kappa 0.58 to 0.92, and agreement between observers from 0.47 to 0.86. The hardest distinction improved most: for Zone III versus non-Zone III, reader-reference agreement went from 0.40 to 0.88. Across 20 zone-dependent ETROP decisions, agreement rose from 0.85 to complete concordance.
Two caveats keep this honest. It is a retrospective reader study on stored images, not prospective use at the cotside, and the ETROP concordance rests on 20 observations — too few to carry much weight on its own. But the starting figures are the ones worth sitting with: interobserver kappa of 0.47 means two ophthalmologists looking at the same baby's retina agreed on zone barely more than half the time beyond chance. In high-volume Indian ROP screening programmes, where images are frequently graded remotely and by rotating readers, standardising that judgement is worth more than most incremental imaging advances.
- Recognise that unaided zone classification agrees poorly between readers — kappa 0.47 in this study
- Zone III versus non-Zone III is where disagreement concentrates and where the map helped most
- Consider a reference overlay for tele-ROP programmes where images are graded by rotating readers
- Retrospective reader study on stored images; prospective cotside performance is untested
- The ETROP decision concordance rests on 20 observations and should not be quoted as the headline
The statistics, in plain English
Kappa measures agreement beyond what chance alone would produce: 0.47 is moderate at best and 0.86 is close to the level expected of a reliable clinical measurement. Because this was a crossover design, the same readers assessed the same cases both ways, which controls for reader ability but means they may have remembered cases between conditions. The ETROP decision result went from 0.85 to perfect agreement across only 20 observations, where one changed decision moves the statistic a long way.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for ophthalmology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free