- Design
- randomised controlled trial, 1:1 allocation, linear mixed-effects analysis, Level II
- Population
- 40 consecutive adult primary cochlear implant operations by 7 attending and 5 trainee surgeons
- Primary outcome
- intraoperative cognitive load on the Surgery Task Load Index
- Effect
- attendings 3.8 vs 5.5 (P < 0.01); residents 7.2 vs 4.8 (P = 0.02); fellows 8.2 vs 8.0 (P = 0.88)
Forty consecutive adult primary cochlear implant operations, performed by seven attending and five trainee surgeons at a high-volume centre, were randomised to standard imaging-based planning or to that plus intraoperative inspection of a patient-specific 3D-printed temporal bone. The primary outcome was the surgeon's intraoperative cognitive load on the Surgery Task Load Index.
The effect reversed by experience. Attending surgeons reported lower overall load with the model (3.8 against 5.5, P < 0.01), driven by reduced physical fatigue. Residents reported higher load (7.2 against 4.8, P = 0.02), with both mental effort and physical fatigue increased. Fellows sat in between with no difference in overall load (8.2 against 8.0) but rated case complexity lower while reporting more physical fatigue. Everyone rated the models as anatomically accurate and useful, particularly for education and for anticipating difficulty.
The reading that fits is that a reference object offloads cognition only for someone who already knows what to do with it. For an experienced surgeon the model answers a question quickly; for a trainee it adds a second representation to reconcile with the imaging and the operative field, which is more work, not less. That the same intervention rated as useful while measurably increasing load is worth noticing — subjective usefulness and cognitive cost are not the same measurement.
With 40 cases split across three experience levels, each subgroup is small and these are subgroup findings within a modest trial. The authors' recommendation — selective use matched to experience rather than routine adoption — is appropriately modest, and the same caution applies to any planning adjunct evaluated on how useful people say it is.
- Do not assume a planning adjunct that helps experienced surgeons helps trainees; here the effect reversed
- Judge such tools on measured cognitive load or performance, not on whether users rate them useful
- Consider selective use of patient-specific models for complex anatomy rather than routine adoption
- Use the models for preoperative teaching, where they were rated most valuable, rather than intraoperatively for trainees
- Read these as subgroup findings in a 40-case trial, not as established effects
Why it matters
The assumption that a visual aid helps a novice most is the opposite of what was measured.
Don't overread it
Cognitive load was self-reported and no surgical outcome was measured, so this says nothing about whether the models improve operations.
The statistics, in plain English
Forty cases divided across attendings, fellows and residents leaves each group small, and these are subgroup comparisons rather than the trial's overall result — the kind of analysis most likely to produce a finding by chance. Linear mixed-effects modelling correctly accounts for the same surgeons contributing several cases. The Surgery Task Load Index is a self-reported scale, so it measures how loaded the surgeon felt rather than how they performed; no operative outcome was compared.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for ent & head and neck surgery, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free