- Design
- retrospective standalone evaluation of an AI nodule analysis system with expert adjudication of failures
- Population
- all 361 UK Lung Cancer Screening participants with a 3-month follow-up low-dose CT; 339 persisting nodules at or above 100 mm3
- Primary outcome
- longitudinal nodule matching success rate
- Effect
- 83.5% matched (95% CI 79.2 to 87.1); 91.8% with one nodule vs 72.8% with more than five; 1.5% needed manual correction
Automated growth assessment in lung cancer screening depends on a step that is rarely measured: matching the same nodule between two scans. This study took all 361 UK Lung Cancer Screening trial participants who had a three-month follow-up low-dose CT and let the AI system work unsupervised - every nodule it detected at baseline above 100 mm3 solid volume went forward to automated matching with no human selection.
The system found 181 participants with 378 baseline nodules at or above the threshold; 39 had resolved by follow-up. Of the 339 persisting nodules, it matched 83.5% (95% CI 79.2 to 87.1). Performance split by burden: 91.8% where a participant had a single baseline nodule, which was 59.7% of the cohort, but 72.8% where there were more than five, which was 6.6%. Expert review of the 56 unmatched findings is the important part - 91.1% were not nodules at all, mostly pleural plaques - so only five discrete solid nodules, 1.5% of the persisting total, actually needed a human to intervene.
Read as a workflow finding rather than a diagnostic one, this says the matching step can largely be automated, and that the residual error is concentrated in two identifiable places: scans with many nodules, and non-nodular structures the detector should not have offered. Both are triageable. If you are implementing this, review the high-burden scans manually and expect to discard plaques rather than chase them.
- Keep manual review for scans with more than five candidate nodules, where matching fell to 72.8%
- Expect most unmatched findings to be pleural plaques, not missed nodules
- Verify the volume threshold your software uses matches your screening protocol
- Audit matching locally before trusting automated volume doubling times
- Do not remove the radiologist from growth assessment on the strength of a single-cohort study
Why it matters
The bottleneck in automated screening is tracking, not detection, and this puts a number on how much of it can be handed over.
Don't overread it
A retrospective evaluation in one trial cohort; prospective validation in a routine, diverse screening population has not been done.
The statistics, in plain English
The headline 83.5% understates the clinically relevant performance, because most failures were structures that were never nodules - correcting for that leaves 1.5% of real nodules needing intervention, with an interval of 0.6% to 3.5%. The subgroup figures come from small numbers: 97 participants with one nodule and just 24 with more than five, so the 72.8% is imprecise. This was a single trial cohort with a three-month interval, not a routine screening programme at scale.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free