- Design
- population based cohort study with parallel derivation, external application and coefficient updating
- Population
- 47,958 New Zealand and 46,558 Chinese adults aged 30-74 with diabetes and no prior CVD
- Primary outcome
- five-year first cardiovascular event; calibration of transferred equations
- Effect
- age hazard ratio per decade in women 1.61 (1.51-1.71) New Zealand versus 2.51 (2.32-2.72) China; expected-to-observed after recalibration 0.938 men and 0.809 women, improving to 1.027 and 1.025 after updating the age coefficient
Two equivalent primary prevention cohorts of people aged 30 to 74 with diabetes and no cardiovascular disease - 47,958 in New Zealand with 5,622 first events, 46,558 in China with 3,650 - were used to derive five-year risk equations by identical methods, so that the coefficients could be compared directly rather than inferred.
Most of the 16 predictors behaved similarly in both countries. Age did not. In women, the hazard ratio per decade was 1.61 (95% CI 1.51 to 1.71) in New Zealand and 2.51 (2.32 to 2.72) in China. When the New Zealand equations were applied to the Chinese cohort, standard recalibration - the usual fix, which rescales predicted risk to observed incidence - failed to correct it, leaving expected-to-observed ratios of 0.938 in men and 0.809 in women. Replacing the age coefficient with the locally derived one brought those to 1.027 and 1.025.
This is directly relevant wherever a clinic uses an equation built somewhere else, which describes most of Indian practice. The practical implication is not that risk scores are useless but that the standard remedy is inadequate: rescaling the output cannot repair a predictor that behaves differently in the population. Where a locally derived equation exists, prefer it. Where it does not, the ranking of patients within your clinic remains more trustworthy than the absolute percentage, and a 20-percent threshold imported wholesale will misclassify systematically by age.
- Check what population your risk calculator was derived in before quoting an absolute percentage
- Treat the risk ranking within your own patients as more reliable than the number itself
- Be most sceptical at the extremes of age, where the coefficient mismatch bites hardest
- Where a locally derived score exists for your population, use it in preference to an imported one
- Do not let a borderline calculated risk override clear clinical indicators such as very high LDL or albuminuria
Why it matters
The standard fix for using a foreign risk equation - recalibrate to local event rates - is shown here to be insufficient.
The statistics, in plain English
An expected-to-observed ratio of 0.809 means the equation predicted about four events for every five that happened - a systematic under-estimate of a fifth, in women. That is a calibration failure, and it is different from the model being bad at ranking: the same model can sort patients correctly and still be wrong about everyone's absolute risk. This is also a comparison of two specific countries; it demonstrates that the transfer problem is real and that age drives it, not that the Chinese coefficients apply to any other population.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for top clinical updates, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free