DailyDoctor Archive Specialties Get app
Back to the 13 September 2026 edition

Research · 03 of 05

A national registry using language models to watch imaging algorithms

Ask who is monitoring the imaging algorithms running in your department, and against what reference - the answer is often nobody and nothing.

Design
Technical workflow description with prompt development and consistency testing across ten repeated runs
Population
Radiology report cohorts for nine use cases within the Assess-AI national registry
Primary outcome
Agreement of language-model finding extraction with report-derived reference labels
Effect
0.985 for intracranial haemorrhage and 0.997 for pulmonary embolism in stage 2 development cohorts, which were not independent validation sets

Assess-AI is the American College of Radiology's national imaging AI registry, and the problem it exists to solve is that an algorithm's performance in deployment drifts away from its validation figures without anyone noticing. Checking that manually does not scale: someone has to read the report, decide what the truth was, and compare it with what the algorithm said.

This paper describes the workflow that automates that step. Language model prompts, built jointly by data scientists and subspecialty radiologists and tuned on use-case-specific report cohorts, extract the clinically relevant finding from the radiology report so it can be compared with the algorithm's output. Prompts were developed for nine use cases and evaluated for consistency across ten repeated runs. In stage 2 development cohorts, agreement with hybrid report-derived reference labels was 0.985 for intracranial haemorrhage and 0.997 for pulmonary embolism.

Those numbers should be read with the caveat the authors themselves attach, twice. The same cohorts were used to refine the prompts and to construct the labels, so they are not independent validation sets, and the discussion states plainly that independent accuracy and clinical utility remain unestablished. What the paper establishes is that report-anchored monitoring at national scale is technically feasible. Whether the monitoring is accurate enough to act on is a separate question that has not been answered.

  • Post-deployment monitoring, not validation, is the gap this addresses - they are different problems
  • The reference standard here is the radiology report, so monitoring inherits the report's errors
  • Departments running AI tools should know whether anyone is tracking performance drift locally
  • Agreement figures of 0.985 and 0.997 come from development cohorts, not independent testing
  • Nine use cases is a start; most deployed imaging AI is not covered by any registry

Why it matters

It turns monitoring deployed algorithms from an unfunded manual chore into something a registry could actually do.

Don't overread it

The reported agreement comes from development cohorts that shaped the prompts and labels - it is not independent validation.

The statistics, in plain English

Agreement of 0.985 and 0.997 looks close to perfect, and the reason to discount it is stated in the paper: the cohorts that produced these figures were also used to refine the prompts and to build the reference labels being compared against. A model tested on the data that shaped it will always look better than it is. The right comparison is a held-out set the developers never saw, which this paper does not report. The consistency check - ten repeated runs per prompt - is a different and useful measure, because language models can give different answers to the same question, and a monitoring tool that is inconsistent is worse than none.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

headneckimaginginterventionalimagingainuclearimaging

Tomorrow morning, before your first patient

One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app