DailyDoctor Archive Specialties Get app
Back to the 19 September 2026 edition

Research · 03 of 05

GPT-5 graded hydrops at 25% from images and 87% from a structured description

Treat a general language model as a reasoning aid over your own structured findings, not as a reader of the image.

Design
Retrospective diagnostic accuracy study comparing prompting strategies and human-AI workflows
Population
436 patients with Ménière's disease (872 ears) imaged with delayed gadolinium-enhanced 3D-FLAIR MRI
Primary outcome
Accuracy, AUC, F1 and agreement for cochlear and vestibular endolymphatic hydrops grading
Effect
Basic prompting: 25% cochlear, 39% vestibular accuracy. Enhanced prompting: 87% and 80%. Human-first assistance raised junior readers from 36% to 59% (cochlear) and 42% to 53% (vestibular)

This retrospective study used delayed gadolinium-enhanced 3D-FLAIR MRI from 436 patients with Ménière's disease — 872 ears — and asked GPT-5 to grade cochlear and vestibular endolymphatic hydrops under several prompting strategies, comparing it with junior and senior neuroradiologists reading independently and with assistance.

Given the images with basic prompting, performance was poor: 25% accuracy for cochlear grading and 39% for vestibular, and adding clinical history or few-shot examples did not rescue it. Given enhanced prompting — structured image descriptions, clinical context and representative examples — accuracy reached 87% for cochlear and 80% for vestibular grading, with better agreement and favourable expert ratings.

In the human-AI experiments, a human-first workflow raised junior physicians' cochlear grading accuracy from 36% to 59% and vestibular from 42% to 53%.

The finding worth carrying is the first one, not the second. The model was not reading the image; it was reasoning over a description a human produced. That is a genuinely useful thing — it is a structured second opinion for a reader who is unsure — but it is not image interpretation, and the 87% figure should never be quoted without the condition attached to it.

  • The 87% figure depends on a human writing the structured description first
  • Twenty-five per cent accuracy from images alone is the number that describes autonomous performance
  • Human-first assistance helped junior readers but left them well short of senior performance
  • Useful as a prompt to reconsider, not as a grading tool
  • Any local evaluation must test the workflow you would actually use, not the best-prompted condition

Why it matters

It separates what a general model can do from what a human feeding it findings can do — a distinction most claims about AI reading elide.

Don't overread it

This is retrospective, single-condition testing on one sequence in one disease; it says nothing about performance on other MRI tasks.

The statistics, in plain English

Accuracy of 25% on a graded scale is close to or below what guessing would give, so the unassisted image-reading result is not a weak performance — it is no performance. The jump to 87% measures how much information the structured description carried, which is to say most of it. And junior readers reaching 59% with assistance is an improvement on 36% but is still wrong two times in five.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

nuclearimagingoncimagingmskimagingpaedimagingcontrastsafetycardiacimaging

Tomorrow morning, before your first patient

One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app