- Design
- Case-level randomised, adjudicator-blinded simulation study
- Population
- 13 family physicians, 260 simulated consultations
- Primary outcome
- Top-3 diagnostic accuracy
- Effect
- 74.6% vs 62.3%; adjusted OR 2.68 (95% CI 1.24 to 6.12)
Thirteen board-certified family physicians managed 260 simulated, voice-based primary care consultations of high difficulty, randomised case by case to a real-time AI diagnostic assistant or no aids. Blinded adjudicators scored whether a correct diagnosis appeared in the physician's top three. No real patients were involved. It was published in JMIR Formative Research on 22 September.
Top-three accuracy was 74.6% with AI and 62.3% without (adjusted OR 2.68, 95% CI 1.24 to 6.12). Top-one accuracy was 57% versus 47.8%, not statistically significant. When the AI's suggestion was wrong, physicians agreed with it more often (55.8% vs 38.8%), a non-significant sign of over-reliance. Consultations took 10.7% longer.
This is a stress test, not evidence of benefit in practice. The comparator had no resources at all, which exaggerates the gap against a real clinic.
- Treat AI diagnostic suggestions as a differential prompt, not an answer.
- Be alert to anchoring on an AI's first suggestion.
- Wait for evaluations in real consultations before adopting such tools.
- Expect consultations to lengthen, not shorten.
Why it matters
It gives an early signal of both the benefit and the anchoring risk of AI in the consulting room.
Don't overread it
Simulated cases against a no-resource comparator cannot show benefit in real practice.
The statistics, in plain English
The top-three difference of 12.3 percentage points has a wide interval (2.7 to 22.6), reflecting only 13 physicians. The over-reliance comparison did not reach significance, so it is a warning rather than a finding. Simulated cases with an unassisted, resource-restricted comparator favour the AI.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for family medicine, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free