DailyDoctor Archive Specialties Get app
Back to the 20 September 2026 edition

Research · 03 of 06

A machine learning tool for neonatal hearing loss, built on 270 infants

Prioritise follow-up by NICU duration and family history; the model itself is not ready for use.

Design
retrospective model development with 80/20 train-test split and stratified 5-fold cross-validation
Population
270 high-risk neonates who passed initial screening — 105 with hearing loss, 165 with normal hearing
Primary outcome
prediction of hearing loss from routinely recorded clinical risk factors
Effect
XGBoost accuracy 85.2%, AUC 87.1%; NICU stay duration and family history the strongest predictors on SHAP analysis

Two hundred and seventy infants who had passed initial hearing screening but carried clinical risk factors — 105 with hearing loss, 165 with normal hearing — were analysed retrospectively. Variables were the ones already recorded: prematurity, low birth weight, hyperbilirubinaemia, phototherapy, NICU stay duration and family history. Five model types were trained with an 80/20 split and stratified five-fold cross-validation.

XGBoost performed best, with 85.2 per cent accuracy and an AUC of 87.1 per cent. SHAP analysis identified NICU stay duration and positive family history as the dominant predictors. The team also built a web application for real-time risk scoring.

The authors are clear that this complements rather than replaces universal screening, and that framing is the right one. Two things temper it: a cohort of 270 with a 39 per cent event rate is far richer in hearing loss than any real newborn population, so the accuracy figure will not transfer; and there is no external validation. What the SHAP finding supports directly is simpler and needs no model — NICU duration and family history are the two variables that should drive follow-up priority.

  • Prioritise diagnostic follow-up by NICU stay duration and family history, which drove the model
  • Do not use this or any published model in place of universal newborn screening
  • Treat the 85 per cent accuracy figure as cohort-specific until externally validated
  • Note the cohort was enriched for hearing loss — real-world performance will be lower
  • Where screening coverage is patchy, risk-based prioritisation of ABR slots is the practical use

Why it matters

The two variables the model relies on are already in every neonatal record, and can be used without the model.

Don't overread it

Retrospective, single-cohort model development with no external validation — the published web tool has not been prospectively tested.

The statistics, in plain English

Accuracy of 85.2 per cent sounds strong but depends heavily on how common the condition is in the sample. Here 39 per cent of infants had hearing loss, far above any real population, which inflates apparent performance. Without external validation on a separate cohort, an AUC of 0.87 from a train-test split within one dataset is a development figure, not a performance one.

Read the rest in the app

You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

QR code to install Daily Doctor
Get Daily Doctor — free

Scan to keep reading on your phone. No account needed to start.

otologysleepairwaylaryngologyheadneckoncpaedent

Tomorrow morning, before your first patient

One edition a day for ent & head and neck surgery, written by the desk, every claim tied to its paper. Six minutes.

Get the app — free
Daily Doctor All 27 specialties, every morning. Free.
Get the app