- Design
- retrospective cohort study with decision-curve analysis of nine risk scores
- Population
- 853,131 US veterans with imaging-confirmed non-cirrhotic steatotic liver disease, 2008-2020
- Primary outcome
- first diagnosis of cirrhosis or hepatocellular carcinoma within ten years
- Effect
- SAFE best net benefit for cirrhosis, 0.019 at a 2.5% threshold; 0.0016 for hepatocellular carcinoma
Nine risk scores were calculated at the point of first imaging in 853,131 US veterans with steatotic liver disease and no viral or primary liver disease, then followed for up to ten years. Cirrhosis developed in 3.96% and hepatocellular carcinoma in 0.35%.
FIB-4, APRI and the steatosis-associated fibrosis estimator discriminated cirrhosis best. The important step is what the authors did next: decision-curve analysis, which asks not how well a score ranks people but whether acting on it does more good than harm at a threshold a clinician would actually use. On that test the SAFE score came out ahead, with a net benefit of 0.019 at a threshold corresponding to a 2.5% ten-year cirrhosis risk - 1.9 extra people correctly identified per 100 labelled at risk. For hepatocellular carcinoma, net benefit was 0.0016, or 1.6 per 1000, which is close to nothing.
So the usable conclusion is narrow and worth having: a score can help decide when to repeat fibrosis assessment in fatty liver, and no score currently justifies starting cancer surveillance in someone who does not have cirrhosis. The cohort was 92.7% male US veterans, which limits transfer to an Indian clinic where fatty liver presents younger and at lower BMI.
- Use a fibrosis score to time the next assessment, not to decide on cancer surveillance.
- SAFE performed best on net benefit for cirrhosis; FIB-4 and APRI discriminated similarly.
- No score supported hepatocellular carcinoma surveillance in non-cirrhotic fatty liver.
- Most patients with fatty liver never develop advanced disease - 3.96% reached cirrhosis in ten years.
- The cohort was 92.7% male veterans; Indian patients present younger and at lower BMI.
Why it matters
It separates a score that should change a follow-up interval from scores that should not start a surveillance pathway.
Don't overread it
A retrospective cohort in a near-entirely male veteran population, with outcomes from diagnosis codes rather than adjudicated review.
The statistics, in plain English
Discrimination and net benefit answer different questions, and this study is a good demonstration of why the second matters more. A score can rank patients well and still be useless for a decision if acting on it produces mostly false positives at the threshold you would use. Net benefit converts that into a single number: 1.9 extra true positives per 100 for cirrhosis is modest but real; 1.6 per 1000 for liver cancer is not worth a surveillance programme.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for internal medicine, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free