A diagnostic meta-analysis of 32 studies evaluated deep-learning models grading knee osteoarthritis on plain radiographs against the Kellgren-Lawrence (K-L) scale.
Performance was strongly grade-dependent. Pooled sensitivity across K-L grades 0 to 4 was 0.90, 0.66, 0.80, 0.87 and 0.88, with precision following a similar pattern. The models were weakest exactly where clinical value would be greatest — early disease, K-L grade 1, where sensitivity was only 0.66. Heterogeneity was high, and external validation was limited.
The honest conclusion is that these tools are not yet ready for routine use, and certainly not for early detection or screening. They may assist reporting of clear moderate-to-severe disease, but a normal or grade-1 automated read cannot yet be trusted to exclude early osteoarthritis.
- Sensitivity only 0.66 for early (K-L grade 1) disease — the grade that matters most
- Better for moderate-to-severe OA (sensitivity 0.80–0.88)
- High heterogeneity and limited external validation across the 32 studies
- Not ready for routine grading or screening; do not rely on a normal automated read to exclude early OA
The statistics, in plain English
A sensitivity of 0.66 at K-L grade 1 means the models miss about a third of genuinely early osteoarthritis — a false-negative rate too high to screen with. High heterogeneity means results varied widely between studies, so a single reported accuracy figure overstates how dependable these models are in a new setting.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for rheumatology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free