- Design
- validation study using the National Health and Aging Trends Study linked to Medicare fee-for-service and Medicare Advantage claims
- Population
- 1,527 fee-for-service and 1,457 Medicare Advantage beneficiaries in Round 11
- Primary outcome
- agreement and discrimination of three claims-based frailty indices against assessment-based frailty index and frailty phenotype
- Effect
- claims-classified frailty 10–12% (Kim index) vs 53–71% (VA index, Hospital Frailty Risk Score); C-statistics 0.79–0.82, 0.77–0.82 and 0.63–0.77 against the assessment-based index
Claims-based frailty algorithms are increasingly used to target services and stratify risk, and they were built for Medicare fee-for-service data. This study calculated three of them — the Kim claims-based frailty index, the Veterans Affairs frailty index and the Hospital Frailty Risk Score — in the National Health and Aging Trends Study linked to both fee-for-service and Medicare Advantage data, and tested them against two reference standards: a comprehensive geriatric assessment-based frailty index and the frailty phenotype.
In 1,527 fee-for-service and 1,457 Medicare Advantage beneficiaries, reference-standard frailty prevalence was 34.2 and 35.5 per cent by the assessment-based index, and 13.3 and 16.2 per cent by the phenotype. The claims-based measures diverged enormously: the Kim index classified 10 to 12 per cent of the whole population as frail, while the VA index and Hospital Frailty Risk Score classified 53 to 71 per cent. Among those with a prior hospitalisation the spread was 30 to 36 per cent against 56 to 87 per cent. Discrimination was better against the assessment-based index (C-statistics 0.79 to 0.82 for the Kim index) than against the phenotype, was consistent by race and survey round, and was modestly better in women.
The reassuring part — that the algorithms behave similarly in Medicare Advantage and fee-for-service — is what the study set out to test. The alarming part is incidental: applying published thresholds produced a fivefold difference in who counts as frail. Anyone using these scores to allocate a service needs to know which index and which threshold produced the number.
- Always state which frailty index and threshold produced a reported prevalence
- Do not compare frailty rates across services using different claims algorithms
- Choose the index by purpose: the Kim index for specificity, the VA index or Hospital Frailty Risk Score for sensitivity
- Tailor the threshold to the intended use rather than adopting the published cut-point
- Treat a claims-based score as a triage signal, not a substitute for comprehensive geriatric assessment
Why it matters
Services are being commissioned against frailty prevalence figures that depend more on the algorithm than on the population.
Don't overread it
This validates claims algorithms against research reference standards in US Medicare data; it does not tell you which one should drive a clinical decision.
The statistics, in plain English
A C-statistic of 0.79 to 0.82 means the Kim index ranks a frail and non-frail pair correctly about four times in five — good for a claims algorithm and not good enough to classify an individual. The fivefold spread in prevalence is not a disagreement about who is frail so much as about where the line is drawn: one index is tuned for specificity, the others for sensitivity, and published thresholds carry those choices with them.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for geriatrics, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free