- Design
- Independent external validation study (CATALINA) of two locked AI pipelines, pooling prospectively collected outcomes from seven randomised adjuvant trials
- Population
- 1,759 patients with early triple-negative breast cancer, 1,356 with complete clinicopathological, pathologist and computational scores
- Primary outcome
- Association of computational and pathologist TIL scores with invasive disease-free survival, distant disease-free survival and overall survival
- Effect
- Computational HR 0.80 (0.73-0.89) for invasive disease-free survival vs pathologist 0.73 (0.66-0.82); computational score lost significance when adjusted for pathologist score
Tumour-infiltrating lymphocytes are one of the few robust prognostic markers in triple-negative breast cancer, and scoring them by eye is slow and variable. CATALINA is the independent validation that computational scoring needed: two previously published artificial-intelligence pipelines were deployed locked and masked, without retraining, against long-term outcomes pooled from seven randomised adjuvant trials.
In 1,356 patients with complete data, computational scores were independently associated with invasive disease-free survival (HR 0.80, 95% CI 0.73-0.89), distant disease-free survival (0.77, 0.69-0.86) and overall survival (0.79, 0.70-0.88) after adjustment. Pathologist-scored stromal lymphocytes were stronger on every endpoint (0.73, 0.70 and 0.72 respectively). Both improved five-year discrimination over clinicopathological variables alone.
The finding that matters is the negative one. Once pathologist scores were in the model, the computational score no longer carried a significant independent association, and it did not improve the area under the curve further. Correlation between the two was modest at best - r 0.375 to 0.473 - so the algorithms are not simply reproducing what a pathologist sees, and what they are adding is not additional prognostic information.
So the use case is substitution, not augmentation. Where a trained pathologist scores stromal lymphocytes routinely, an algorithm adds cost without adding information. Where that expertise is not available - which describes a great many centres, including much of India - a locked model that generalised across seven trials without retraining offers a reproducible score where the alternative is none at all. That is a real contribution, and it is a narrower one than the technology is usually sold with.
- Do not commission an AI TIL tool for a centre that already scores stromal lymphocytes well
- Where no trained scorer is available, a validated locked model is a defensible substitute
- Ask any vendor for independent external validation on outcomes, not concordance with pathologists
- Modest correlation with pathologist scores means the two are not interchangeable measurements
- TIL scoring remains prognostic, not predictive - it does not by itself select a treatment
The statistics, in plain English
A hazard ratio of 0.80 per unit of score means higher lymphocyte scores went with better outcomes; the pathologist's 0.73 is a stronger association than the algorithm's 0.80, since both sit below 1.0 and lower is stronger here. 'Did not maintain a significant association' when adjusted for the pathologist score means the two measures largely carry the same information, and the pathologist carries more of it. A correlation of 0.375 to 0.473 is modest - roughly, the algorithm and the pathologist agree on less than a quarter of the variation between patients.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for oncology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free