- Design
- prospective rubric development and evaluation study, February to December 2025
- Population
- 19 development participants, 6 research-team evaluators and 111 further lay participants and radiologists; 480 reports in field testing
- Primary outcome
- interreader agreement on quality grades and agreement between subjective and rule-based distribution decisions
- Effect
- radiologist interreader α = 0.65 and 95.8% decision agreement; AI-applied rubric κ = 0.44 with 88.1% agreement on distribution decisions
Patient-facing translations of radiology reports are being generated by large language models in a growing number of departments, and a mistranslation reaches the patient without passing a radiologist. This prospective study built and tested a rubric to decide whether such a report is safe to release.
Nineteen participants developed it through survey-workshop cycles. The final rubric grades five attributes — clarity, content, certainty, tone and verbosity — each on a three-point scale, with a decision rule that a grade 1 on any attribute except verbosity makes the report unsafe to distribute. Reports of deliberately varying quality were generated from a public dataset and evaluated by research-team members, 111 further lay participants and radiologists, and by a language model.
Agreement was good where it needed to be: almost perfect between lay and radiologist research-team members (α = 0.87), substantial among the 12 additional radiologists (α = 0.65), moderate among lay participants (α = 0.51), with 95.8 per cent agreement between radiologists' subjective distribution decisions and the rule-based ones. When the model applied the rubric itself across 480 reports, its grades agreed only moderately with the reference standard (κ = 0.44), but the rule-based distribution decisions it produced matched the reference decisions 88.1 per cent of the time.
- If your department releases AI-translated reports, put an explicit release rule between generation and the patient
- Grade certainty and tone, not just readability — those are where mistranslation causes harm
- Do not rely on the model to grade its own output unsupervised at κ = 0.44
- Audit a sample of released plain-language reports against the original impressions
- Decide in advance who is accountable when a translated report misleads a patient
Why it matters
These translations already reach patients directly, with no step where a radiologist reads what was sent.
Don't overread it
The rubric was tested on reports generated from a public dataset, not on live clinical output, and the authors say it needs further validation.
The statistics, in plain English
Krippendorff's alpha of 0.65 is substantial agreement — good enough for a screening rule, not for adjudicating individual reports. The gap between the model's moderate grade agreement (κ = 0.44) and its 88.1 per cent agreement on the release decision reflects that the decision rule is binary and forgiving: it can get the grade wrong and still reach the right call about distribution.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free