Files
school_compare/docs/superpowers/specs/2026-07-11-compare-screen-expert-review.md
T

178 lines
9.3 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Compare Screen Redesign — Expert Data Review
**Date:** 2026-07-11
**Reviewer:** subagent briefed as an English education-standards / DfE-Ofsted data expert
**Subject:** desktop + mobile compare mockups and the redesign spec
(`2026-07-11-compare-screen-redesign-design.md`)
**Status:** first-pass must-fixes applied 2026-07-12; second-pass
findings (below) applied 2026-07-12 — mockups + spec §4/§8 updated
## Must-fix
1. **COVID gap is wrong and drops a real results year.** KS2 tests were
cancelled 2019/20 and 2020/21 only; they resumed in 2021/22 with
published school-level results (England RWM ≈ 59%). The mockup charts
omit 2021/22 entirely and the tooltip claims no tests were held
2019/202021/22. Fix: add 2021/22 to axis and all series; shrink the
gap band; optionally annotate 2021/22 with DfE's post-pandemic
comparability caution.
2. **Report-card at-a-glance summary miscounts areas.** Detail list has
4 Strong / 2 Expected / 1 Attention needed + Safeguarding met, but
the summary says "3 areas Expected standard" — it counts safeguarding
as a graded area. Safeguarding is a separate binary judgement and
must be excluded from rating counts.
3. **"Where the offers went" derivation is unsound.** Places 1st-pref
offers ≠ "second or third choices": the residual can include 4th6th
preference offers (pan-London scheme) and LA-allocated children who
didn't choose the school; and offers don't necessarily equal PAN.
Use the real 2nd/3rd-preference fields being promoted from
`raw.ees_admissions`; until then drop the row.
4. **Ofsted timeline in the copy is wrong.** Overall grades were
abolished September 2024, not November 2025; Sept 2024Nov 2025
inspections kept the four key judgements without an overall grade
(ungraded inspections carried grades forward). Neither mockup shows
the interim regime, which will dominate real comparisons. Fix copy
and add an interim example.
5. **Barclay's "published an overall grade only — no area-by-area
detail" misdescribes inspections.** No inspection type does that; a
2021 graded inspection necessarily had subgrades — the gap is in our
dataset. If it was an ungraded (s8) inspection, "Outstanding" is a
carried-forward grade and should say so. Fix: "We don't hold
area-by-area detail for this inspection", and distinguish graded vs
ungraded in the data model.
## Should-fix
6. Writing is teacher assessment, not a test — "national tests and
teacher assessments"; note TA caveat on the Writing strip.
7. Verify renewed-framework wording against Ofsted's final toolkit:
likely "Needs attention" (not "Attention needed") and "Personal
development and well-being" (which otherwise collides with the
identically-named legacy judgement). Pin every label to the
published toolkit.
8. "Expected standard" now means two things on one page (Ofsted area
rating vs KS2 measure) — disambiguate in tooltips.
9. Disadvantaged row: DfE definition includes looked-after / previously
looked-after children, not just FSM6; benchmark labels inconsistent
across desktop/mobile; subgroup percentages need cohort sizes or a
volatility threshold before chips are attached.
10. "Trend, last 7 years" spans ten years; sparklines render the COVID
gap as equal spacing (the exact defect the audit criticises) and
"Improved: 52% → 87%" endpoint-cherry-picks a volatile series.
11. At-a-glance "Getting a place" uses different metrics per school
(Barclay is also oversubscribed on total preferences but shows a
green chip). Standardise on first-preference success %. Explain the
equal-preference rule; condition "living close by matters" on the
school's actual oversubscription criteria.
12. "457 applications for 180 places" = total preferences at any rank,
not head-to-head applicants; lead with first preferences vs places.
Add offers-vs-final-intake (waiting lists/appeals) caveat.
13. Elmhurst's subgrade list is likely missing Early years provision
(school has a nursery) — possible pipeline gap.
14. "Ofsted rating" label is obsolete post-Sept-2024 — use "Latest
Ofsted inspection"; check whether Oct 2021 is the latest inspection
or merely the latest graded one.
15. SEN: "EHCP plans" is redundant; 28% SEN support often indicates
resourced provision — add a note; England SEN-support ≈ 14%, not 13%.
## Nice-to-have
16. Consistent labelling of official DfE vs dataset-computed benchmarks
(and medians shouldn't be called averages inconsistently).
17. England 2015/16 RWM (53%) exists in DfE publications — the null is
a dataset gap; source it or the England line looks broken.
18. "1 in 4 first choices missed out" — actually more than 1 in 4.
19. "1,273 of 1,260 places (full)" is over capacity; capacity figures
are often stale — say "at or above capacity".
20. State the actual suppression rule (DfE: ≤5 pupils suppressed,
small numbers rounded) instead of "a handful".
21. Spec §4.3 progress chips can't exist for displayed years: KS2
progress ended with 2022/23 (no KS1 baseline) and returns
~2027/28 with the reception baseline. Make explicit in the spec.
IDACI (spec §4.5) is absent from mockups; if shipped, caveat it
describes pupils' neighbourhoods, not the school.
22. Tooltips should give the official term "first preference" alongside
the plain-English "first choice".
## Overall assessment (verbatim gist)
The bones are genuinely good by education-data standards —
England-average anchoring, explicit non-comparability messaging across
Ofsted regimes, refusal to synthesise an overall grade, time-true
x-axis, neutral FSM/EAL framing — better than most commercial
school-comparison sites. But items 15 are outright factual errors or
misdescriptions that a well-informed parent or Ofsted would catch;
the admissions section needs the most conceptual work (equal
preference, preferences-vs-applicants, offers-vs-intake). Fix 15
before user testing; the rest fold into the planned PRs.
---
# Second-pass review (2026-07-12)
Same reviewer, after the must-fixes and the new three-tier metric
exposure model were applied.
## Verification of first-pass must-fixes
- **1 (COVID/2021/22): resolved.** Time-true axis, band covers only the
cancelled years, England 58.7% consistent with official figures,
dataset gaps break lines honestly; reading/maths England series all
match published figures; RWM ≤ min(subject) checks pass.
- **2 (report-card count): resolved** — safeguarding excluded, spec §8.2.
- **3 (offers derivation): resolved** — row removed, spec §8.3 bans it.
- **4 (Ofsted timeline): resolved on desktop; mobile omits the interim
regime clause** (see finding 6).
- **5 (Barclay explanation): resolved.**
## New findings
1. **Should-fix — scaled-score strip domain contradicts caption.**
Caption says "scaled scores run 80120", strips render 100120;
truncated domain exaggerates small gaps and below-100 averages
would fall off the edge. Render 80120, or caption the 100120
window honestly and define below-100 behaviour.
2. **Should-fix — scaled-score England ticks (106/105/105) unsourced.**
Plausible but hand-entered; verify against DfE 2024/25 tables and
add loading official England scaled scores to the pipeline list
(absent from §8.1/§8.6).
3. **Should-fix — "Writing" listed under "Higher standard" in the
picker.** Writing TA outcome is "greater depth" (GDS), never
"higher standard". Label "Writing — greater depth (teacher
assessment)"; tooltip the combined higher-standard composition.
4. Nice — "grammar & punctuation" summary line drops "spelling" (GPS).
5. Nice — science is teacher-assessed (no KS2 test since 2009) and
coarse; tooltip it like writing; reconsider its tier-2 slot.
6. **Should-fix — mobile Ofsted copy skips the interim regime**
(Sept 2024Nov 2025) that desktop explains. One clause fixes it.
7. **Should-fix — benchmark provenance still inconsistent** (EAL
tooltip unsourced; FSM/disadvantaged chips vs tooltips use three
vocabularies; header note says all England averages are official).
Adopt one house style: official = "England average", computed =
"benchmark / typical state school (our dataset)". Also tighten EAL
definition to census wording ("first language known or believed to
be other than English").
8. Nice — "community primaries" distance note attached to an academy
(Elmhurst); say "non-faith primaries" or condition on policy field.
9. Nice — "Improving since 2022" → "since 2022/23".
10. Nice — England chart tooltips show decimals; §7 mandates whole
percents.
## Residual gaps not covered by spec §8
11. Spec promises IDACI-in-words, Attendance section, and tier-2
gender/absence that the mockups never show — mark post-v1 or
demonstrate, so implementation scope is unambiguous.
12. Add official England scaled-score averages to the pipeline task
list.
13. Add the writing/greater-depth terminology rule to §8.7.
## Verdict
All must-fixes genuinely resolved; the tier model is conceptually
sound ("no measure is lost", honest dataset-gap breaks, grouped
picker). Remaining issues are contained: one internal contradiction
(80120 vs 100120), one provenance inconsistency, one terminology
error (writing/GDS). With findings 13 and 67 addressed, the data
framing is fit to put in front of parents.