Most departments running AI triage have no record of the times it was wrong. The algorithm flags a study, the radiologist reads it, and the report says what the radiologist thinks - which is correct practice and also means the disagreement leaves no trace. False positives vanish into a normal report and false negatives vanish into a positive one.
That matters because performance drift is invisible without it. A tool that was validated on one scanner, one protocol and one population will behave differently on yours, and will keep behaving differently as protocols change. The only people positioned to notice are the readers, and they are not being asked.
The fix is one line in the report or, better, one click in the worklist: a flag recording that the AI output and the final read differed, and in which direction. It costs seconds, it gives the department a denominator, and it converts an impression that the tool 'gets a lot of chest studies wrong' into a number someone can act on.
- Record AI-reader disagreement at the point of reporting, in the direction it occurred
- Do not let the flag change the report itself - the report is the clinical document
- Aggregate monthly; a rate is actionable where anecdotes are not
- Note protocol or scanner changes in the same log, since they are the usual cause of drift
- If no mechanism exists, a shared spreadsheet beats nothing and takes an afternoon to set up
Why it matters
The people who see an algorithm failing are the only people not currently recording it.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free