- Design
- Multicentre retrospective diagnostic accuracy study with external testing and reader study
- Population
- 2,576 adults with suspected hip trauma and same-episode CT or MRI
- Primary outcome
- Sensitivity, specificity and AUC for femoral neck fracture
- Effect
- Sensitivity 97.5%, specificity 98.8%, AUC 0.99; occult fractures 94.7% vs radiologists 86.2%
A multicentre retrospective study from four hospitals trained and externally tested a deep learning model on pelvic and hip radiographs from 2,576 adults with suspected hip trauma, all of whom had CT or MRI in the same episode as the reference standard.
In the pooled test sets, the model detected femoral neck fractures with 97.5% sensitivity and 98.8% specificity. For the hardest group, fractures that were radiograph-negative or indeterminate (189 cases), it found 94.7%, against 86.2% for musculoskeletal radiologists and 68.8% for emergency physicians. With AI assistance, radiologists' sensitivity rose from 93.7% to 97.2% and emergency physicians' from 84.3% to 95.6%, and reading times fell by 15–19%.
The design is retrospective and enriched with fractures, so real-world positive predictive value will be lower, and all patients had cross-sectional imaging, which is not how most emergency departments work. Prospective evaluation is needed before relying on it to rule out fracture. The largest gain was for emergency physicians, which matters where radiographs are read first by non-radiologists, as in many Indian emergency departments at night. It was published in September 2026.
- Treat this model as promising but not yet validated prospectively; do not use AI alone to exclude a hip fracture.
- The biggest benefit was for emergency physicians reading films without a radiologist.
- Keep MRI or CT for patients with a normal radiograph who still cannot bear weight.
- If your department adopts fracture AI, audit its missed fractures and false positives locally.
- Watch for automation bias: readers may accept an AI negative too readily.
Why it matters
Occult femoral neck fractures are the ones whose delay leads to displacement and bigger operations.
Don't overread it
A retrospective, fracture-enriched test set overstates how well the model will perform in everyday triage.
The statistics, in plain English
Sensitivity is the share of real fractures detected; specificity is the share of non-fractures correctly called negative. An AUC of 0.99 means the model almost always ranked a fracture above a non-fracture. In a test set with many fractures, false positives matter less than they will in a real emergency department, where most films are normal.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free