- Design
- deep learning segmentation model pre-trained on virtual CT data and fine-tuned on expert-annotated clinical masks, compared against LAA-950
- Population
- 20 virtual human models plus multicentre clinical cohorts of 101, 23 and 1,159 patients
- Primary outcome
- segmentation accuracy (Dice), measurement bias and limits of agreement, and correlation with Fleischner visual score and pulmonary function
- Effect
- Dice 48.4% vs 22.5% in the main annotated cohort; bias 1.0% ± 4.3 vs 2.1% ± 15.1; visual score correlation 0.77 vs 0.47; DLCO multivariable R² 0.31–0.32 vs 0.21–0.25
Emphysema on CT is quantified by counting voxels below −950 Hounsfield units, the LAA-950 metric. It is simple, which is its appeal, and it is sensitive to scanner, dose and reconstruction kernel, which is its weakness — the same lung can produce different numbers on different machines. A segmentation-based deep learning model was built to do the job properly, pre-trained on virtual data across controlled scanner and dose conditions and fine-tuned on expert-annotated emphysema masks.
It outperformed the threshold on every axis tested. Dice segmentation accuracy was 76.6% against 51.5% on virtual data, 48.4% against 22.5% in the first clinical cohort of 101 patients, and 64.8% against 32.7% in a second cohort of 23. Bias and limits of agreement tightened markedly (1.0% ± 4.3 against 2.1% ± 15.1). Correlation with radiologists' Fleischner visual scores rose from 0.47 to 0.77, and multivariable correlation with carbon monoxide transfer factor improved from an R² of 0.21–0.25 to 0.31–0.32 in a cohort of 1,159 patients.
The practically important number is the limits of agreement, not the Dice score. A metric with ±15% limits cannot be used to follow a patient over time or across scanners, which is what emphysema quantification is for; ±4.3% might be. This is a developed and internally evaluated model rather than a deployed product, so the question for a department is not whether to adopt it but whether to keep treating LAA-950 as a number with meaning across scans done on different machines.
- Do not compare LAA-950 values across scanners or reconstruction kernels — the metric is not stable across them.
- Record scanner, dose and kernel alongside any quantitative emphysema figure you issue.
- State in the report whether a change over time is larger than the measurement variability.
- Correlate quantitative output with the visual Fleischner score rather than replacing it.
- Read a Dice coefficient below 50% as a statement about how hard the segmentation task is, not only about the method.
Why it matters
If the emphysema percentage moves by 15% when only the scanner changed, following a patient with it is an illusion of measurement.
The statistics, in plain English
The Dice coefficients look poor in absolute terms — 48.4% in the largest annotated cohort — because agreeing on the exact boundary of emphysematous lung is genuinely hard, and both methods are judged against one expert annotation. The comparison, not the level, is the finding. Limits of agreement of ±4.3% against ±15.1% is the number that matters for practice: it describes how much the measurement can move without the lung changing. Correlations with lung function remain modest (R² about 0.3), meaning CT quantification explains roughly a third of the variation in transfer factor and is not a substitute for it.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free