- Design
- Interobserver reproducibility study using whole-slide images
- Population
- 20 appendiceal goblet cell adenocarcinoma cases graded by 7 pathologists
- Primary outcome
- Agreement on WHO 5th-edition three-tier grade
- Effect
- Full agreement in 4/20; Fleiss kappa 0.29 (0.14–0.44); worst for necrosis (0.07) and architectural disarray (0.05)
The WHO 5th-edition classification grades appendiceal goblet cell adenocarcinoma in three tiers, but its reproducibility had not been tested. Seven pathologists with an interest in appendiceal pathology independently graded whole-slide images from 20 cases, scoring 15 defined histological features.
All seven agreed on the grade in only four of 20 cases. Overall agreement was fair — a Fleiss kappa of 0.29 — and pairwise agreement ranged widely, with a median around 0.13. Agreement was best for extracellular mucin and tumour sheets, and worst for mild architectural disarray and necrosis, two features meant to drive the grade.
For sign-out, this is a caution about how much weight a three-tier grade can bear when it rests on features experts cannot apply consistently. The authors suggest a two-tier low-versus-high scheme would be more reproducible; until the classification changes, flag borderline cases and consider a second read rather than reporting a precise tier as if it were robust.
- All seven pathologists agreed on grade in only 4 of 20 cases (20%).
- Overall agreement was fair (Fleiss kappa 0.29, 95% CI 0.14–0.44); median pairwise agreement about 0.13.
- Agreement was worst for the grade-defining features of architectural disarray (kappa 0.05) and necrosis (0.07).
- Flag borderline goblet cell adenocarcinoma grades and consider a second read; a two-tier scheme may be more reproducible.
Why it matters
It warns against over-trusting a grade that determines prognosis when experts cannot reproduce it.
Don't overread it
This is 20 cases and seven readers — it shows poor reproducibility, not that the grade is prognostically meaningless.
The statistics, in plain English
A kappa of 0.29 is 'fair' agreement — well below the 0.6–0.8 you would want before a grade drives management. That the grade-defining features (necrosis, architectural disarray) had the worst agreement shows the problem is in the criteria, not just the observers.
Read the rest in the app
You have read your two free briefings this month. The app carries all 27 specialties, every morning, free — and this finding is waiting in it.

Scan to keep reading on your phone. No account needed to start.
Tomorrow morning, before your first patient
One edition a day for pathology, written by the desk, every claim tied to its paper. Six minutes.
Get the app — free