Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
9.3 KiB
Compare Screen Redesign — Expert Data Review
Date: 2026-07-11
Reviewer: subagent briefed as an English education-standards / DfE-Ofsted data expert
Subject: desktop + mobile compare mockups and the redesign spec
(2026-07-11-compare-screen-redesign-design.md)
Status: first-pass must-fixes applied 2026-07-12; second-pass
findings (below) applied 2026-07-12 — mockups + spec §4/§8 updated
Must-fix
- COVID gap is wrong and drops a real results year. KS2 tests were cancelled 2019/20 and 2020/21 only; they resumed in 2021/22 with published school-level results (England RWM ≈ 59%). The mockup charts omit 2021/22 entirely and the tooltip claims no tests were held 2019/20–2021/22. Fix: add 2021/22 to axis and all series; shrink the gap band; optionally annotate 2021/22 with DfE's post-pandemic comparability caution.
- Report-card at-a-glance summary miscounts areas. Detail list has 4 Strong / 2 Expected / 1 Attention needed + Safeguarding met, but the summary says "3 areas Expected standard" — it counts safeguarding as a graded area. Safeguarding is a separate binary judgement and must be excluded from rating counts.
- "Where the offers went" derivation is unsound. Places − 1st-pref
offers ≠ "second or third choices": the residual can include 4th–6th
preference offers (pan-London scheme) and LA-allocated children who
didn't choose the school; and offers don't necessarily equal PAN.
Use the real 2nd/3rd-preference fields being promoted from
raw.ees_admissions; until then drop the row. - Ofsted timeline in the copy is wrong. Overall grades were abolished September 2024, not November 2025; Sept 2024–Nov 2025 inspections kept the four key judgements without an overall grade (ungraded inspections carried grades forward). Neither mockup shows the interim regime, which will dominate real comparisons. Fix copy and add an interim example.
- Barclay's "published an overall grade only — no area-by-area detail" misdescribes inspections. No inspection type does that; a 2021 graded inspection necessarily had subgrades — the gap is in our dataset. If it was an ungraded (s8) inspection, "Outstanding" is a carried-forward grade and should say so. Fix: "We don't hold area-by-area detail for this inspection", and distinguish graded vs ungraded in the data model.
Should-fix
- Writing is teacher assessment, not a test — "national tests and teacher assessments"; note TA caveat on the Writing strip.
- Verify renewed-framework wording against Ofsted's final toolkit: likely "Needs attention" (not "Attention needed") and "Personal development and well-being" (which otherwise collides with the identically-named legacy judgement). Pin every label to the published toolkit.
- "Expected standard" now means two things on one page (Ofsted area rating vs KS2 measure) — disambiguate in tooltips.
- Disadvantaged row: DfE definition includes looked-after / previously looked-after children, not just FSM6; benchmark labels inconsistent across desktop/mobile; subgroup percentages need cohort sizes or a volatility threshold before chips are attached.
- "Trend, last 7 years" spans ten years; sparklines render the COVID gap as equal spacing (the exact defect the audit criticises) and "Improved: 52% → 87%" endpoint-cherry-picks a volatile series.
- At-a-glance "Getting a place" uses different metrics per school (Barclay is also oversubscribed on total preferences but shows a green chip). Standardise on first-preference success %. Explain the equal-preference rule; condition "living close by matters" on the school's actual oversubscription criteria.
- "457 applications for 180 places" = total preferences at any rank, not head-to-head applicants; lead with first preferences vs places. Add offers-vs-final-intake (waiting lists/appeals) caveat.
- Elmhurst's subgrade list is likely missing Early years provision (school has a nursery) — possible pipeline gap.
- "Ofsted rating" label is obsolete post-Sept-2024 — use "Latest Ofsted inspection"; check whether Oct 2021 is the latest inspection or merely the latest graded one.
- SEN: "EHCP plans" is redundant; 28% SEN support often indicates resourced provision — add a note; England SEN-support ≈ 14%, not 13%.
Nice-to-have
- Consistent labelling of official DfE vs dataset-computed benchmarks (and medians shouldn't be called averages inconsistently).
- England 2015/16 RWM (53%) exists in DfE publications — the null is a dataset gap; source it or the England line looks broken.
- "1 in 4 first choices missed out" — actually more than 1 in 4.
- "1,273 of 1,260 places (full)" is over capacity; capacity figures are often stale — say "at or above capacity".
- State the actual suppression rule (DfE: ≤5 pupils suppressed, small numbers rounded) instead of "a handful".
- Spec §4.3 progress chips can't exist for displayed years: KS2 progress ended with 2022/23 (no KS1 baseline) and returns ~2027/28 with the reception baseline. Make explicit in the spec. IDACI (spec §4.5) is absent from mockups; if shipped, caveat it describes pupils' neighbourhoods, not the school.
- Tooltips should give the official term "first preference" alongside the plain-English "first choice".
Overall assessment (verbatim gist)
The bones are genuinely good by education-data standards — England-average anchoring, explicit non-comparability messaging across Ofsted regimes, refusal to synthesise an overall grade, time-true x-axis, neutral FSM/EAL framing — better than most commercial school-comparison sites. But items 1–5 are outright factual errors or misdescriptions that a well-informed parent or Ofsted would catch; the admissions section needs the most conceptual work (equal preference, preferences-vs-applicants, offers-vs-intake). Fix 1–5 before user testing; the rest fold into the planned PRs.
Second-pass review (2026-07-12)
Same reviewer, after the must-fixes and the new three-tier metric exposure model were applied.
Verification of first-pass must-fixes
- 1 (COVID/2021/22): resolved. Time-true axis, band covers only the cancelled years, England 58.7% consistent with official figures, dataset gaps break lines honestly; reading/maths England series all match published figures; RWM ≤ min(subject) checks pass.
- 2 (report-card count): resolved — safeguarding excluded, spec §8.2.
- 3 (offers derivation): resolved — row removed, spec §8.3 bans it.
- 4 (Ofsted timeline): resolved on desktop; mobile omits the interim regime clause (see finding 6).
- 5 (Barclay explanation): resolved.
New findings
- Should-fix — scaled-score strip domain contradicts caption. Caption says "scaled scores run 80–120", strips render 100–120; truncated domain exaggerates small gaps and below-100 averages would fall off the edge. Render 80–120, or caption the 100–120 window honestly and define below-100 behaviour.
- Should-fix — scaled-score England ticks (106/105/105) unsourced. Plausible but hand-entered; verify against DfE 2024/25 tables and add loading official England scaled scores to the pipeline list (absent from §8.1/§8.6).
- Should-fix — "Writing" listed under "Higher standard" in the picker. Writing TA outcome is "greater depth" (GDS), never "higher standard". Label "Writing — greater depth (teacher assessment)"; tooltip the combined higher-standard composition.
- Nice — "grammar & punctuation" summary line drops "spelling" (GPS).
- Nice — science is teacher-assessed (no KS2 test since 2009) and coarse; tooltip it like writing; reconsider its tier-2 slot.
- Should-fix — mobile Ofsted copy skips the interim regime (Sept 2024–Nov 2025) that desktop explains. One clause fixes it.
- Should-fix — benchmark provenance still inconsistent (EAL tooltip unsourced; FSM/disadvantaged chips vs tooltips use three vocabularies; header note says all England averages are official). Adopt one house style: official = "England average", computed = "benchmark / typical state school (our dataset)". Also tighten EAL definition to census wording ("first language known or believed to be other than English").
- Nice — "community primaries" distance note attached to an academy (Elmhurst); say "non-faith primaries" or condition on policy field.
- Nice — "Improving since 2022" → "since 2022/23".
- Nice — England chart tooltips show decimals; §7 mandates whole percents.
Residual gaps not covered by spec §8
- Spec promises IDACI-in-words, Attendance section, and tier-2 gender/absence that the mockups never show — mark post-v1 or demonstrate, so implementation scope is unambiguous.
- Add official England scaled-score averages to the pipeline task list.
- Add the writing/greater-depth terminology rule to §8.7.
Verdict
All must-fixes genuinely resolved; the tier model is conceptually sound ("no measure is lost", honest dataset-gap breaks, grouped picker). Remaining issues are contained: one internal contradiction (80–120 vs 100–120), one provenance inconsistency, one terminology error (writing/GDS). With findings 1–3 and 6–7 addressed, the data framing is fit to put in front of parents.