DailyDoctor Archive Specialties Get app
Back to the 14 September 2026 edition

Clinical update · 01 of 06

A language model amplifies whatever expertise is holding it

A text-based language model workflow improves diagnostic accuracy for both radiologists and residents, but trainees accept wrong answers at twice the rate — supervise use rather than assuming the tool is self-correcting.

Design
retrospective multi-reader study with two sessions, generalised estimating equations for group comparison
Population
93 thoracic cases (radiograph, CT, MRI, PET/CT) read by five thoracic radiologists and five radiology residents
Primary outcome
reader diagnostic accuracy before and after receiving language model output generated from the reader's own free-text description
Effect
radiologists 56.3% to 65.6%, residents 42.4% to 58.5%; residents accepted model-favoured diagnoses 73.2% vs 48.9% and switched to a wrong answer after misleading output 60.6% vs 32.4%

Image-input models face regulatory and technical obstacles, so a reader-mediated text workflow is the version most departments could actually run today. A retrospective study tested it on 93 thoracic cases — radiographs, CT, MRI and PET/CT — from the Korean Society of Thoracic Radiology quiz archive, each with the correct diagnosis and four distractors. Ten readers, five thoracic radiologists and five residents, first chose a diagnosis and wrote a free-text description of the findings. That description alone, without any image, went to the model, which ranked the five options with reasoning. The readers then chose again.

Accuracy rose in both groups: 56.3% to 65.6% for the radiologists, 42.4% to 58.5% for the residents. The model itself did better on text than on images (63.9% vs 52.7%), and better on text written by a specialist than by a resident (67.3% vs 60.4%). The residents improved more in absolute terms — 16.1 against 9.2 percentage points — but the reason is the problem. Residents accepted the model's favoured diagnosis in 73.2% of cases against 48.9% for radiologists, and when the model's output was misleading, they switched to the wrong answer 60.6% of the time against 32.4%.

So the same tool, used by two groups, produced a net gain in both and an unsafe pattern of use in one. The determinant was not the model but the reader: expertise shaped both the quality of the input description and the willingness to disagree with the output. That has a direct implication for how these tools are introduced into a department — as a second opinion for someone competent to reject it, not as a scaffold for someone who is not.

  • Treat model output as a differential to argue with, not a ranking to accept.
  • The quality of the free-text description determines the quality of the answer — write findings, not impressions.
  • Supervise trainee use specifically; the acceptance rate, not the accuracy gain, is the thing to monitor.
  • Record when a model output changed a report, so the switch rate can be audited locally.
  • These were quiz cases with five fixed options — real reporting has no answer list.

Why it matters

It locates the value of these tools in the radiologist rather than the model, which is the opposite of how they are usually sold.

The statistics, in plain English

The headline that residents gained more (16.1 vs 9.2 percentage points, P = .02) is the least useful number in the study, because the gain and the risk come from the same behaviour: accepting the model. The pair to read together is 73.2% acceptance against 48.9%, and 60.6% against 32.4% for switching to a wrong answer after a misleading output. A tool that improves the average while doubling the rate of induced error in the less experienced group is not simply 'better for trainees'. Accuracy here is measured against a single correct answer from five options, which flatters every participant including the model.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

chestimagingimagingaipaedimagingnuclearimagingheadneckimaging

Tomorrow morning, before your first patient

One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app