- Design
- retrospective two-session reader study with generalised estimating equations; readers committed before seeing model output
- Population
- 93 thoracic imaging cases read by 5 thoracic radiologists and 5 radiology residents
- Primary outcome
- diagnostic accuracy before and after language model assistance, by reader expertise
- Effect
- radiologists 56.3% to 65.6%, residents 42.4% to 58.5%; switching to a wrong answer after a misleading output 32.4% vs 60.6%
Ten readers — five thoracic radiologists and five residents — interpreted 93 thoracic cases spanning radiography, CT, MRI and PET-CT, each with five candidate diagnoses. In the first session each reader chose a diagnosis and wrote a free-text description of the findings. A large language model then received only that description, not the images, ranked the five options and explained its top three. In the second session readers saw the output generated from their own description and chose again.
The model was better with words than with pictures: 63.9% accurate from reader descriptions against 52.7% from images. It was also better with better words — 67.3% from thoracic radiologists' descriptions against 60.4% from residents'. Readers improved in the second session, from 56.3% to 65.6% for the radiologists and from 42.4% to 58.5% for the residents, and the residents' gain was larger (16.1 against 9.2 percentage points).
The cost sits in the same numbers. Residents accepted the model's favoured diagnosis 73.2% of the time against 48.9% for radiologists, and when the output was misleading they switched to the wrong answer 60.6% of the time against 32.4%. The group that gained most from the assistance was the group least able to resist it when it was wrong.
That is automation bias measured rather than asserted, and it has a specific implication for how these tools get introduced. A text-based workflow of this kind is deployable now — it needs no regulatory clearance for image analysis and no integration beyond a text box. Deploying it to trainees, where the accuracy gain looks largest, is also deploying it where the error mode is worst. These were quiz cases with a fixed five-option differential, which flatters the model considerably against real reporting.
- Expect the largest accuracy gain and the largest harm from a misleading output in the same group — trainees
- Where a language model is used this way, require the reader to commit to a diagnosis before seeing the output, as this workflow did
- Judge the quality of the input description, not only the output — the model was more accurate from expert descriptions
- Do not treat a five-option quiz differential as equivalent to open-ended reporting; this design favours the model
- Build supervision around trainee use rather than restricting it, given the measured gain when the output is correct
Why it matters
The argument for giving these tools to trainees is the accuracy gain, and the same data show trainees are where a wrong output does most damage.
Don't overread it
Quiz cases with five predefined options are not reporting, and this cannot predict how the workflow performs on an unselected list.
The statistics, in plain English
These are accuracy percentages across 93 cases and 10 readers, so each reader contributes many observations that are not independent of each other — generalised estimating equations were used to account for that, which is the right approach. A difference of 60.6% against 32.4% in switching to a wrong answer is large and unlikely to be chance (P = 0.009), but rests on the subset of cases where the model was misleading. The fixed five-option differential means the model had a 20% floor by guessing, so its 52.7% from images alone is less impressive than it sounds, and real reporting has no such list.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free