- Design
- multicentre retrospective development and external validation, with a 10-reader comparison and AI-assisted reading
- Population
- 2,576 adults with suspected hip trauma at four hospitals, 2009-2025; pooled test group 1,766, of whom 189 had radiograph-negative or indeterminate fractures
- Primary outcome
- sensitivity, specificity and area under the curve for femoral neck fracture detection, and reader performance with assistance
- Effect
- model sensitivity 97.5%, specificity 98.8%, AUC 0.99; in occult fractures 94.7% against radiologists 86.2% and emergency physicians 68.8%; with assistance readers rose to 97.2% and 95.6%, reading times fell 14.9% and 18.9%
The occult femoral neck fracture is a recognised failure point: the radiograph is read as normal or equivocal, the patient is sent home or mobilised, and the fracture displaces. This multicentre retrospective study assembled 2,576 adults with suspected hip trauma across four hospitals from 2009 to 2025, all of whom had radiographs and CT or MRI in the same episode, and trained a two-stage model called OccuNet, using contrastive pretraining on artefact-augmented radiographs before fine-tuning a fracture detector. It was tested on three external test sets.
On the pooled test group of 1,766 patients, the model reached 97.5% sensitivity and 98.8% specificity, with an area under the curve of 0.99. The important subgroup is the 189 radiograph-negative or indeterminate fractures - Garden I-II - where the model's sensitivity was 94.7% against 86.2% for five musculoskeletal radiologists and 68.8% for five emergency physicians (both P<0.001). With the model's assistance, radiologist sensitivity rose from 93.7% to 97.2% and emergency physician sensitivity from 84.3% to 95.6%, while mean reading times fell by 14.9% and 18.9%.
The emergency physician figure is the one that should change something. A sensitivity of 68.8% for occult fracture on a plain radiograph, read by the doctor who decides whether the patient is admitted, is the real-world starting point in most hospitals at 2 am - and it rose to 95.6% with assistance. This remains a retrospective study with an enriched population: everyone had cross-sectional imaging, so the prevalence of fracture is far higher than in an unselected emergency department, and specificity in practice will be worse. No prospective deployment has been reported. But the size of the gap it identifies, between what is expected of a plain radiograph and what it delivers, is not contingent on the model at all.
- Where clinical suspicion of hip fracture persists after a normal or equivocal radiograph, proceed to CT or MRI - that judgement is the current safeguard and this study reinforces it
- Do not mobilise a patient with hip pain and inability to weight-bear on the strength of a normal radiograph alone
- If you are evaluating an AI tool for this, ask for prospective data in an unselected emergency population, not retrospective figures from a cross-sectional-imaging cohort
- Note that reading time fell rather than rose - the usual objection to AI assistance did not hold here
- Audit your own department's rate of delayed-diagnosis femoral neck fracture before and after any deployment
Why it matters
It puts a number on how often the plain radiograph, read by the person deciding on admission, misses the fracture.
Don't overread it
Retrospective, with a cohort enriched for fracture by requiring same-episode CT or MRI - real-world specificity will be lower and no prospective deployment has been reported.
The statistics, in plain English
An area under the curve of 0.99 looks close to perfect, but it was measured in a population where everyone had cross-sectional imaging - that is, where a fracture was already suspected enough to justify it. In an unselected emergency department the prevalence would be far lower, and at low prevalence even a specificity of 98.8% generates a large number of false positives relative to true ones. The subgroup result is the durable finding: comparing model, radiologist and emergency physician sensitivity on the same 189 occult fractures is a fair internal comparison regardless of how the cohort was assembled.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free