This retrospective study covered 39,323 chest CT examinations in 19,433 patients at a tertiary centre between September 2021 and May 2024, split into pre-AI and post-AI phases around the introduction of commercial pulmonary nodule software. The outcome was radiology reporting time, modelled with adjustment for reader function, examination type, patient location and requesting specialty, clustered by radiologist.
Adjusted median reporting time fell from 21.3 minutes to 18.2, a 14.6% reduction, with an adjusted hazard ratio for report completion of 1.17. Three minutes per scan across a department is a substantial number.
The heterogeneity is the more interesting part. The largest reductions were for ECG-gated thoracic CT, at 41.1%, and for thoracic radiologists, at 25.0%. Emergency department work benefited least. So the software helped most where the reader was already a subspecialist and the study was already structured — not where the reading was hardest or the reader least experienced, which is where the intuitive case for AI usually starts.
The design is the limitation and it is a real one. This is before-and-after, not randomised: three years separate the phases, during which staffing, PACS, reporting templates and case mix all move. The adjustment is careful but it cannot rule out that something else changed.
- Expect roughly 15% off reporting time, not the larger figures vendors quote from reading-room studies.
- Benefit concentrated in subspecialist readers and structured examinations, not in emergency work.
- Audit your own before-and-after times: the effect here is centre-specific and design-limited.
- Do not read this as evidence on accuracy; the outcome was time, not detection.
- A before-and-after study across three years cannot separate the software from everything else that changed.
The statistics, in plain English
An adjusted hazard ratio of 1.17 here means reports were completed about 17% faster at any given moment — the survival model is being used to analyse time-to-report rather than time-to-death, which is why a hazard ratio above 1.0 is the good direction. The interval of 1.14 to 1.21 is tight because the sample is very large. Large samples give precision, not validity: the weakness is the design, not the numbers. Before-and-after comparisons attribute to the intervention everything that changed between the two periods, and a lot changes in a radiology department over three years. The formal test for heterogeneity across subgroups was significant, which means the variation between reader types is unlikely to be chance — a genuinely uneven effect rather than noise.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free