- Design
- Technical workflow description with prompt development and consistency testing across ten repeated runs
- Population
- Radiology report cohorts for nine use cases within the Assess-AI national registry
- Primary outcome
- Agreement of language-model finding extraction with report-derived reference labels
- Effect
- 0.985 for intracranial haemorrhage and 0.997 for pulmonary embolism in stage 2 development cohorts, which were not independent validation sets
Assess-AI is the American College of Radiology's national imaging AI registry, and the problem it exists to solve is that an algorithm's performance in deployment drifts away from its validation figures without anyone noticing. Checking that manually does not scale: someone has to read the report, decide what the truth was, and compare it with what the algorithm said.
This paper describes the workflow that automates that step. Language model prompts, built jointly by data scientists and subspecialty radiologists and tuned on use-case-specific report cohorts, extract the clinically relevant finding from the radiology report so it can be compared with the algorithm's output. Prompts were developed for nine use cases and evaluated for consistency across ten repeated runs. In stage 2 development cohorts, agreement with hybrid report-derived reference labels was 0.985 for intracranial haemorrhage and 0.997 for pulmonary embolism.
Those numbers should be read with the caveat the authors themselves attach, twice. The same cohorts were used to refine the prompts and to construct the labels, so they are not independent validation sets, and the discussion states plainly that independent accuracy and clinical utility remain unestablished. What the paper establishes is that report-anchored monitoring at national scale is technically feasible. Whether the monitoring is accurate enough to act on is a separate question that has not been answered.
- Post-deployment monitoring, not validation, is the gap this addresses - they are different problems
- The reference standard here is the radiology report, so monitoring inherits the report's errors
- Departments running AI tools should know whether anyone is tracking performance drift locally
- Agreement figures of 0.985 and 0.997 come from development cohorts, not independent testing
- Nine use cases is a start; most deployed imaging AI is not covered by any registry
Why it matters
It turns monitoring deployed algorithms from an unfunded manual chore into something a registry could actually do.
Don't overread it
The reported agreement comes from development cohorts that shaped the prompts and labels - it is not independent validation.
The statistics, in plain English
Agreement of 0.985 and 0.997 looks close to perfect, and the reason to discount it is stated in the paper: the cohorts that produced these figures were also used to refine the prompts and to build the reference labels being compared against. A model tested on the data that shaped it will always look better than it is. The right comparison is a held-out set the developers never saw, which this paper does not report. The consistency check - ten repeated runs per prompt - is a different and useful measure, because language models can give different answers to the same question, and a monitoring tool that is inconsistent is worse than none.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free