DailyDoctor Archive Specialties Get app
Back to the 9 September 2026 edition

Research · 02 of 05

GPT-5 graded endolymphatic hydrops well — once a radiologist had described the images to it

GPT-5 graded endolymphatic hydrops at 25% accuracy from images and 87% once a human supplied structured descriptions — it is an assistive reasoning tool for junior readers, not an image reader.

Design
Retrospective reader study comparing GPT-5 under several prompting strategies with junior and senior neuroradiologists, alone and in assisted workflows
Population
436 patients with Ménière's disease, 872 ears, on delayed gadolinium-enhanced 3D-FLAIR MRI
Primary outcome
Accuracy of cochlear and vestibular endolymphatic hydrops grading, with AUC, F1 and agreement
Effect
Basic prompting 25% cochlear and 39% vestibular; enhanced prompting 87% and 80%; junior readers assisted improved 36% to 59% and 42% to 53%

Four hundred and thirty-six patients with Ménière's disease, 872 ears, had delayed gadolinium-enhanced 3D-FLAIR MRI graded for cochlear and vestibular endolymphatic hydrops by GPT-5 under several prompting strategies, and by junior and senior neuroradiologists working alone and with the model. Given the images with basic prompting, including clinical history or few-shot examples, accuracy was 25% for cochlear and 39% for vestibular grading. Given structured image descriptions, clinical context and representative examples, accuracy rose to 87% and 80%. In the human-first collaborative mode, junior physicians improved from 36% to 59% for cochlear grading and 42% to 53% for vestibular.

The gap between 25% and 87% is the finding, and it is not a finding about the model's radiological ability. What changed between those two numbers was that a human converted the image into structured text. The model's contribution is reasoning over a description, not perception of a scan — which is exactly the capability boundary that gets blurred when a general-purpose model is described as reading an MRI.

Two consequences follow. First, no autonomous role: 25% accuracy on a graded scale is not a starting point that careful prompting rescues, it is evidence that the image input was not being used. Second, a real assistive one: junior readers improved substantially, and in a service where subspecialty neuroradiology cover is thin, a tool that raises a trainee's grading toward consultant level has obvious value. Treat this as a study of structured reporting with a language model attached, and note that the improvement was measured on grading agreement, not on any patient outcome.

  • Do not deploy a general-purpose language model to grade images directly; the unassisted accuracy here was 25%.
  • The performance came from structured image descriptions written by a human, not from image reading.
  • The assistive gain was largest for junior readers — that is where to evaluate it, if anywhere.
  • Outcomes measured were grading accuracy and agreement, not patient management or hearing outcomes.
  • Any local trial needs its own prompt and description template; performance was entirely prompt-dependent.

The statistics, in plain English

Accuracy figures on a multi-grade scale need a baseline to mean anything: with a small number of ordered grades, chance alone gets a fair share right, so 25% is not merely poor but close to uninformative. The jump to 87% under enhanced prompting cannot be attributed to the model in isolation, because the enhanced condition includes human-generated structured descriptions — the two are confounded by design, and no arm separates them. The collaborative results are cleaner: junior readers went from 36% to 59%, a within-reader comparison where each reader is their own control. That is the number worth carrying forward, and it is still an accuracy measure, not a clinical outcome.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

ultrasoundimagingabdominalimagingimagingaineuroimagingnuclearimagingpaedimaging

Tomorrow morning, before your first patient

One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app