DailyDoctor Archive Specialties Get app
Back to the 13 September 2026 edition

Practice changer · 05 of 05

Language models could not reproduce a rank order list, and displaced the same groups doing it

Keep language models out of rank order list generation; use them for retrieval and audit, and check any ranking tool for subgroup displacement before it goes near a decision.

Design
Single-institution pilot study; seven model configurations, ten lists each, under two information conditions
Population
148 applicants interviewed for 7 diagnostic radiology residency positions, 2025-2026 cycle
Primary outcome
Agreement with the committee's final rank order list, by truth-anchored weighted Kendall's tau, and rank displacement by subgroup
Effect
Median tau 0.15-0.36 from application data alone vs 0.84-0.93 with interviewer scores; interview score alone 0.83, Step 2 CK alone 0.12; female, international and non-MD applicants systematically displaced downwards (q < 0.05)

A single programme took the 148 applicants it interviewed in the 2025-2026 cycle for 7 diagnostic radiology positions and asked whether a language model could reproduce its final rank order list. Seven model configurations - three proprietary models at medium and high reasoning levels, plus one open-weight model - each generated ten independent lists under two conditions: pre-interview application data alone, and the same data with interviewer scores added. That is 140 lists, compared with the committee's using a truth-anchored weighted Kendall's tau.

From application data alone, agreement was poor across every configuration: median tau 0.15 to 0.36. Adding the interview scores lifted it to 0.84 to 0.93. For context, simply sorting applicants by summed interviewer score reproduced the final list at 0.83, while sorting by USMLE Step 2 CK score alone managed 0.12. Raising the reasoning level from medium to high did not reliably help. Producing the committee's list by hand had taken about 390 faculty-hours.

The finding that makes this a practice question rather than a curiosity is the displacement. In the application-only condition, the models systematically placed female, international and non-MD applicants below where the human committee put them, at q < 0.05, while placing programme signallers higher. The authors are careful that this does not establish which list was better - but a tool that moves specific groups down, consistently, on data that excludes the interview, is not one to put at the front of a selection process. Their conclusion, and the right one, is that these models belong in data retrieval and retrospective auditing, not in generating the list.

  • Do not use a language model to generate a rank order list from application data
  • Retrospective auditing of an existing list is a defensible use; generating one is not
  • If your programme is piloting such a tool, test for subgroup displacement before anything else
  • Interview score alone reproduced the committee list at tau 0.83 - the interview is carrying the judgement
  • USMLE Step 2 CK alone reached 0.12, which is worth remembering when it is treated as a screening threshold

Why it matters

A tool that saves 390 faculty-hours is going to be adopted, and this is the evidence about what it does on the way.

Don't overread it

Single-institution pilot; it does not establish that the committee's list was the better one, only that the models' differed systematically.

The statistics, in plain English

Kendall's tau measures how well two orderings agree, from 0 for unrelated to 1 for identical. The application-only figures of 0.15 to 0.36 mean the model lists and the committee list have little in common. The jump to 0.84 to 0.93 once interviewer scores were supplied is the informative comparison, because it shows the models can rank perfectly well when given the information that actually drove the decision - the failure is in the input, not the arithmetic. The subgroup result carries a q value rather than a p value, meaning it survives correction for testing several subgroups at once, which makes a chance finding much less likely. With 148 applicants at one programme, the direction is the finding and the magnitude is not transferable.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

headneckimaginginterventionalimagingainuclearimaging

Tomorrow morning, before your first patient

One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app