- Design
- retrospective quality-improvement study with structured record annotation against a radiologist gold standard
- Population
- 185 adults who had abdominal imaging in 2022 with a radiologist recommendation for additional imaging
- Primary outcome
- proportion of patients with at least one process-related diagnostic error
- Effect
- 14.6% (27 of 185; 95% CI 10.2-20.4), the leading factor being delay in performing the ordered test; annotator agreement with gold standard rose from 61.1% to 80%
An academic health system took 200 randomly selected adults who had abdominal imaging during 2022 in which the radiologist recommended additional imaging, and had trained annotators work through the medical records against a purpose-built taxonomy of process-related error. The final sample was 185 patients.
At least one diagnostic error was found in 14.6% (95% CI 10.2-20.4). The largest single contributory factor was delay in performing the test that had been ordered. Annotator agreement with a gold standard set by two abdominal radiologists rose from 61.1% at baseline to 80% after four rounds of training, which is the methodological finding - a structured annotation process can be taught well enough to be used for audit.
The errors here are not errors of interpretation. Nobody misread a scan; the recommended follow-up did not happen, or did not happen in time. That failure mode is invisible to every conventional radiology quality measure, all of which look at the report rather than at what happened next. It is also the failure mode most likely to be worse in a system where the follow-up scan is self-funded and the patient decides whether to return - which describes a large share of Indian practice. Counting it is the first step to reducing it.
- Treat 'recommended additional imaging' as an order that needs tracking, not as advice that ends at the report
- Build or ask for a report that lists recommendations made and whether the follow-up examination occurred
- Make the recommendation specific - which examination, what interval - so that a delay is detectable
- Audit a sample of follow-up recommendations rather than a sample of reports
- Where cost is the barrier to follow-up, record that; it is a system finding, not a patient failing
Why it matters
It names a category of harm that every existing radiology quality metric is blind to, because it happens after the report is signed.
Don't overread it
A single academic centre, a sample chosen because further imaging had been recommended - this is not the error rate of abdominal imaging in general.
The statistics, in plain English
A rate of 14.6% with a confidence interval of 10.2% to 20.4% comes from 27 events in 185 patients, so the interval is wide and the point estimate should not be quoted precisely. Note the sample was deliberately enriched: every patient had a recommendation for further imaging, which is exactly the population in which a process failure can occur, so this is not the error rate across all abdominal imaging. The agreement improvement from 61.1% to 80% is based on about 50 cases and should be read as showing the method is teachable rather than as a precise reliability figure.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free