- Design
- Retrospective deep learning development and external validation (MobileNetV2, sequential forward feature selection)
- Population
- 1844 patients with gliomas or glioneuronal/neuronal tumours, preoperative multiparametric MRI
- Primary outcome
- Classification across six integrated histological and molecular categories
- Effect
- External AUC 0.92 for IDH status, 0.83 for 1p/19q, 0.95 for PA vs PXA; integrated per-category AUCs 0.69 to 0.94
Preoperative multiparametric MRI from 1844 patients was used to build a three-level classifier separating six categories of glioma and glioneuronal tumour, with a segmentation step reaching a mean Dice of 0.92. Each classification step was given its own optimal sequence combination rather than all five sequences everywhere.
On external testing the system reached an AUC of 0.92 for separating adult-type diffuse gliomas from circumscribed astrocytic gliomas and glioneuronal tumours, 0.92 for IDH status, 0.83 for 1p/19q codeletion and 0.95 for pilocytic astrocytoma against pleomorphic xanthoastrocytoma. Integrated across the whole pipeline, per-category AUCs ranged from 0.69 to 0.94.
That range is the finding. The best steps are strong enough to influence how a case is discussed before surgery — which operation, which centre, what to consent for. The weakest, at 0.69, is not, and an integrated pipeline inherits the errors of every step above it. The output was formatted as a structured virtual pathology report, which is exactly where the risk of over-reading sits.
- Treat a predicted molecular subtype as a hypothesis for the multidisciplinary meeting, not a result.
- Note which sequences each step needs; protocols that omit ADC or FLAIR will not support it.
- Ask for per-category performance, not headline AUC, when any vendor offers this capability.
- Histology and sequencing remain the diagnosis; nothing here shortens that pathway.
- Indian centres with long waits for molecular testing are where triage value would be greatest — and where the error cost is highest.
Why it matters
It puts molecular subtyping into the preoperative window, where it can still change the operation.
Don't overread it
A retrospective classification study with no prospective use and no outcome data.
The statistics, in plain English
AUC measures how well a model ranks cases, not how often it is right about one patient: 0.92 is strong discrimination, 0.69 is barely better than a coin weighted slightly in the right direction. Accuracy figures of 0.81 to 0.95 look better than the AUCs because the categories are unbalanced — a rare tumour can be classified accurately by being predicted rarely. External validation is the study's strength; single-country data remain a limit.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for radiology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free