DailyDoctor Archive Specialties Get app
Back to the 10 September 2026 edition

Research · 03 of 05

A rubric for deciding whether an AI-written patient report is safe to send

A five-attribute rubric with a simple fail rule gave 88.1% agreement when a model applied it, which makes it a usable checkpoint before AI-written reports reach patients — and not a reason to remove human review.

Design
Prospective rubric development and evaluation using survey-workshop cycles with lay participants and a multidisciplinary panel
Population
19 development participants; 6 research-team members plus 111 lay participants and radiologists in evaluation; 480 reports in field testing
Primary outcome
Interreader agreement on rubric grades, and agreement between rule-based and subjective decisions to release a report
Effect
Research-team agreement α 0.87; lay α 0.51 and radiologist α 0.65; model grading κ 0.44 with 88.1% agreement on the rule-based release decision

Plain-language versions of radiology reports are increasingly generated by language models and released to patients, and a translation error in that pipeline reaches the patient with no clinician between. This prospective study built a grading rubric through survey-workshop cycles with lay participants and a multidisciplinary panel, then tested it. Reports of deliberately varied quality were generated by two commercial models from public-dataset impressions and graded by research-team members, by 111 further lay participants and radiologists, and by a language model.

The rubric has five attributes — clarity, content, certainty, tone, verbosity — each scored on three points, with a decision rule: a report graded 1, unsafe or unacceptable, on any attribute other than verbosity should be withheld. Agreement among the trained research team was almost perfect (α 0.87). Among wider participants it was moderate for lay readers (α 0.51) and substantial for radiologists (α 0.65), and the rule-based distribution decision matched their subjective judgement 91.2% and 95.8% of the time. In broader field testing with 80 lay participants across 480 reports, agreement with reference grades was moderate (κ 0.43). A language model applying the rubric also reached moderate agreement on grades (κ 0.44), but its rule-based distribution decisions matched the reference-standard rule-based decisions 88.1% of the time.

That last figure is the interesting one, and it needs reading carefully. Grade agreement was mediocre while the release decision agreed far more often, because the rule only asks whether any attribute fails — a question more robust to disagreement about whether something is a 2 or a 3. It suggests a workable checkpoint rather than a solved problem: a model can plausibly screen, but 88% agreement means roughly one report in eight is classified differently, and the failures that matter are the unsafe reports let through. Anyone releasing AI-generated reports to patients needs some gate of this kind, and needs to know its error rate before switching the human off.

  • Do not release AI-generated patient reports without a defined quality gate and a named person who owns it.
  • Score clarity, content, certainty and tone as blocking attributes; verbosity alone should not withhold a report.
  • Measure your own agreement rate before trusting a model to run the check.
  • Track the failure direction that matters — unsafe reports released, not safe ones withheld.
  • The reports here were generated from impressions in a public dataset, not from live clinical reports.

The statistics, in plain English

Krippendorff's alpha of 0.87 among three trained lay and three trained radiologist raters shows the rubric can be applied consistently by people who have been taught it; α 0.51 among untrained lay participants shows what happens when they have not. That drop is the usual pattern for any structured scoring system and argues for training rather than against the rubric. The more important comparison is between grade agreement (κ 0.44) and decision agreement (88.1%): coarse decisions survive disagreement that fine grades do not. But a percentage agreement figure hides direction, and for a safety gate the two error types are not equivalent — withholding a good report costs a delay, releasing an unsafe one costs a patient. The study does not separate them.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

interventionalchestimagingnuclearimagingcardiacimagingimagingaibreastimagingultrasoundimaging

Tomorrow morning, before your first patient

One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app