- Design
- Prospective rubric development and evaluation using survey-workshop cycles with lay participants and a multidisciplinary panel
- Population
- 19 development participants; 6 research-team members plus 111 lay participants and radiologists in evaluation; 480 reports in field testing
- Primary outcome
- Interreader agreement on rubric grades, and agreement between rule-based and subjective decisions to release a report
- Effect
- Research-team agreement α 0.87; lay α 0.51 and radiologist α 0.65; model grading κ 0.44 with 88.1% agreement on the rule-based release decision
Plain-language versions of radiology reports are increasingly generated by language models and released to patients, and a translation error in that pipeline reaches the patient with no clinician between. This prospective study built a grading rubric through survey-workshop cycles with lay participants and a multidisciplinary panel, then tested it. Reports of deliberately varied quality were generated by two commercial models from public-dataset impressions and graded by research-team members, by 111 further lay participants and radiologists, and by a language model.
The rubric has five attributes — clarity, content, certainty, tone, verbosity — each scored on three points, with a decision rule: a report graded 1, unsafe or unacceptable, on any attribute other than verbosity should be withheld. Agreement among the trained research team was almost perfect (α 0.87). Among wider participants it was moderate for lay readers (α 0.51) and substantial for radiologists (α 0.65), and the rule-based distribution decision matched their subjective judgement 91.2% and 95.8% of the time. In broader field testing with 80 lay participants across 480 reports, agreement with reference grades was moderate (κ 0.43). A language model applying the rubric also reached moderate agreement on grades (κ 0.44), but its rule-based distribution decisions matched the reference-standard rule-based decisions 88.1% of the time.
That last figure is the interesting one, and it needs reading carefully. Grade agreement was mediocre while the release decision agreed far more often, because the rule only asks whether any attribute fails — a question more robust to disagreement about whether something is a 2 or a 3. It suggests a workable checkpoint rather than a solved problem: a model can plausibly screen, but 88% agreement means roughly one report in eight is classified differently, and the failures that matter are the unsafe reports let through. Anyone releasing AI-generated reports to patients needs some gate of this kind, and needs to know its error rate before switching the human off.
- Do not release AI-generated patient reports without a defined quality gate and a named person who owns it.
- Score clarity, content, certainty and tone as blocking attributes; verbosity alone should not withhold a report.
- Measure your own agreement rate before trusting a model to run the check.
- Track the failure direction that matters — unsafe reports released, not safe ones withheld.
- The reports here were generated from impressions in a public dataset, not from live clinical reports.
The statistics, in plain English
Krippendorff's alpha of 0.87 among three trained lay and three trained radiologist raters shows the rubric can be applied consistently by people who have been taught it; α 0.51 among untrained lay participants shows what happens when they have not. That drop is the usual pattern for any structured scoring system and argues for training rather than against the rubric. The more important comparison is between grade agreement (κ 0.44) and decision agreement (88.1%): coarse decisions survive disagreement that fine grades do not. But a percentage agreement figure hides direction, and for a safety gate the two error types are not equivalent — withholding a good report costs a delay, releasing an unsafe one costs a patient. The study does not separate them.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free