DailyDoctor Archive Specialties Get app
Back to the 23 September 2026 edition

Research · 02 of 05

LLM help lifted residents most — and misled them most

If trainees use LLM tools for differentials, supervise them: they gained the most and were misled the most.

Design
Retrospective two-session reader study
Population
10 readers (5 thoracic radiologists, 5 residents); 93 thoracic imaging cases
Primary outcome
Diagnostic accuracy with and without LLM assistance
Effect
Residents 42.4% → 58.5%; consultants 56.3% → 65.6%; switch to wrong answer after misleading output 60.6% vs 32.4%

In a Korean reader study, 10 readers — five thoracic radiologists and five residents — interpreted 93 thoracic quiz cases, choosing from five diagnoses and writing a free-text description. An LLM then ranked the differential from each reader's own description, without seeing the images, and readers could revise.

The LLM was more accurate from text (63.9%) than from images (52.7%), and better from consultants' descriptions than residents'. Accuracy rose from 42.4% to 58.5% for residents and from 56.3% to 65.6% for thoracic radiologists. But residents accepted the LLM's favoured answer 73% of the time, and switched to a wrong diagnosis after misleading output 61% of the time, against 32% for consultants.

The gain depended on the quality of the reader's description and the reader's ability to reject bad advice. These were quiz cases with fixed options, not real reporting.

  • Treat LLM differentials as a prompt to think, not an answer to accept.
  • Supervise trainees who use LLM tools; the harm falls on the least experienced.
  • Teach trainees to describe findings well — it improves both their reports and any AI output.
  • Record when an AI tool influenced a diagnosis.

Why it matters

An assistant that helps the novice most also harms the novice most, which shapes how it should be deployed.

Don't overread it

Quiz cases with five fixed options and ten readers; this is not evidence about accuracy in clinical reporting.

The statistics, in plain English

The improvements are averages across fixed multiple-choice cases; real reporting has no five-option list, so accuracy gains may be smaller. The 61% switch rate applies only to the cases where the LLM output was misleading.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

neuroimagingchestimagingimagingainuclearimaging

Tomorrow morning, before your first patient

One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app