- Design
- Prognostic model development with external validation; eight accelerated failure time models with Weibull distribution
- Population
- 1,180 patients with a first demyelinating attack in Barcelona (median follow-up 10.8 years) and 108 in Melbourne (11.1 years)
- Primary outcome
- Eight outcomes including McDonald 2017 diagnosis, second attack, disability worsening, progression independent of relapse activity and EDSS 3.0
- Effect
- Harrell's C 0.653-0.823 in derivation and 0.623-0.766 externally, except relapse-associated worsening models which did not validate. 67.5% met McDonald criteria, 24.6% developed progression independent of relapse activity
A first demyelinating attack is where treatment decisions in multiple sclerosis are hardest, because the range of possible futures is enormous and nothing about the presentation obviously separates them. Spider-MS is an attempt to give each of those futures a number at that first visit.
The model was built from 1,180 patients in the Barcelona first-attack cohort, all under 50 and seen within three months of onset, and externally validated in 108 patients from Melbourne, with median follow-up around 11 years in both. Rather than one prediction it produces eight — McDonald 2017 diagnosis, a second attack, more than two new T2 lesions a year, relapse-associated worsening at the first and subsequent attacks, confirmed and sustained disability worsening, progression independent of relapse activity, and reaching EDSS 3.0. The predictors are all available at first presentation: age, sex, attack topography, brain and spinal cord lesion counts, CSF oligoclonal bands, and the proportion of time on high or moderate efficacy treatment.
Accuracy was moderate to high — Harrell's C between 0.653 and 0.823 internally, and 0.623 to 0.766 externally for all outcomes except relapse-associated worsening, where the external validation failed. The direction of the predictors is consistent and unsurprising: older age, cord involvement at the first attack, more brain and cord lesions, oligoclonal bands, and less time on effective treatment all predicted worse outcomes.
The practical value is not the score but the structure. A patient asking 'what will happen to me' is asking eight different questions, and they have different answers.
- Record cord involvement at the first attack explicitly — it predicted worse outcomes across several endpoints
- Send CSF oligoclonal bands; they carry prognostic weight, not only diagnostic weight
- Count brain and cord lesions rather than describing them qualitatively
- Note that treatment exposure is a model input, so predictions are not fixed at baseline
- Do not quote a single 'prognosis' — the outcomes separate, and so should the conversation
Why it matters
It separates 'what is my prognosis' into the eight distinct questions patients are actually asking at the first attack.
The statistics, in plain English
Harrell's C is the survival equivalent of a C-statistic: 0.5 is chance, 0.8 is good. A range of 0.653 to 0.823 means some of these eight models work well and others barely beat guessing, and the external figures are lower, as they nearly always are. The external cohort was 108 patients, which is small for validating eight models, and the relapse-associated worsening models failed there — a specific, disclosed failure rather than a general caveat.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for neurology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free