DailyDoctor Archive Specialties Get app
Back to the 12 September 2026 edition

Research · 02 of 06

The frailty index's central assumption held up when it was tested properly

Build your frailty index from whatever deficits your records hold, provided you count around 30 or more - which items you use makes little difference, and short machine-learned indices do not generalise.

Design
secondary analysis of national survey data with Cox and random survival forest models, permutation importance ranking and 200 randomly constructed indices per item count
Population
3,669 adults living with cardiovascular disease, 1999-2018, with a 46-item frailty index
Primary outcome
prediction of all-cause and cardiovascular mortality
Effect
46 individual items outperformed any composite (C-index 0.71, 95% CI 0.69-0.74, against 0.64-0.67); randomly built indices converged with importance-ranked ones at ~35 items for all-cause and ~20 for cardiovascular mortality

The frailty index rests on a claim that sounds too convenient: deficits are interchangeable, so what matters is how many you count rather than which ones. Machine learning has been used repeatedly to build shorter, 'optimised' indices on the assumption that some items must matter more. This study tested the claim directly in 3,669 adults with cardiovascular disease from a national survey, using a 46-item index and predicting all-cause and cardiovascular mortality with Cox and random survival forest models.

At every item count from 1 to 46, three approaches were compared: the top-ranked items used as individual predictors, those same items combined into an index, and 200 indices each built from randomly chosen items. Indices built at random improved steadily as items were added and converged with the importance-ranked composites by about 35 items for all-cause mortality and 20 for cardiovascular mortality. Importance-ranked short indices peaked at about 10 items and then declined - which the authors interpret as those short indices benefiting from avoiding dilution rather than from having found the right items.

One result cuts the other way and deserves stating: keeping all 46 items as separate predictors outperformed any composite score (C-index 0.71, 95% CI 0.69-0.74, against 0.64-0.67). Collapsing items into a single number does discard prognostic information. The trade is that the single number transfers between populations and datasets, which an item-level model does not. The practical conclusion is reassuring for anyone using a locally assembled frailty index: provided you count enough deficits, roughly 30 or more, it does not much matter which ones you had available.

  • Use whichever 30 or more deficits your records actually contain - the specific items matter little at that count
  • Be sceptical of short 'optimised' frailty indices derived by machine learning; the selection is outcome-specific and sample-dependent
  • Do not expect a frailty index to outperform a full individual-item model for prediction - it trades accuracy for generalisability
  • Keep the index consistent within a service so that scores are comparable between patients and over time
  • Note this was tested in adults with cardiovascular disease predicting mortality; other outcomes may behave differently

Why it matters

It licenses using the frailty index you can actually assemble, rather than the one a paper recommends.

Don't overread it

One dataset of adults with cardiovascular disease predicting mortality - interchangeability was tested for this outcome in this population, not universally.

The statistics, in plain English

A C-index is the probability that the model ranks two patients in the correct order of risk, where 0.5 is a coin toss and 1.0 is perfect - so 0.71 against 0.64-0.67 is a real but modest advantage for the item-level model. The design here is the interesting part: building 200 random indices at each item count and watching them converge with the 'optimised' ones is a direct test of interchangeability, and a much stronger one than showing that a single chosen index works. That importance-ranked short indices peaked and then declined is a signal of overfitting to this sample, not of an optimal item set.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

healthyageingfrailtydementiafalls

Tomorrow morning, before your first patient

One edition a day for geriatrics, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app