Files
school_compare/docs/superpowers/specs/2026-08-28-destination-measures-design.md
T
TudorandClaude Opus 5 2e9b5c83c5
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 7m23s
fix(destinations): the masking pass can no longer exit unsafely
Review found _mask_for_disclosure could return with its invariant broken
and say nothing. add_companion only ever withheld a *published* cell, so a
group with one suppressed category and every other one not_applicable —
routine in special schools and AP, where few categories apply — left the
loop with the lone suppressed cell still solvable. Reproduced on a
nine-pupil cohort: one hidden cell, cohort served, residual intact.

A disclosure-control pass that fails silently is worse than none, because
everything downstream trusts it. The loop now runs until the invariant
holds and escalates when no companion exists: the pupil group is dropped
from the payload, and an empty block serialises as None so the section is
absent rather than an empty shell. disclosure_invariant_holds() is exported
so tests assert it directly instead of re-deriving it, and an exhaustive
test sweeps all 81 suppression patterns of a four-category group.

Also fixes a test that set up six measures and checked one: the loop was
`for measure in ["school_sixth_form"]`. It now checks every measure, and
against the real invariant — none hidden, or at least two, rather than
"at least two", which the five published measures would have failed.

No regression on real data: 262 mainstream secondaries, all-pupils bar
still drawable on 94%, zero invariant violations, one disadvantaged group
dropped by the new escalation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-30 21:35:03 +01:00

18 KiB
Raw Blame History

Destination Measures — Design

Date: 2026-08-28 Status: awaiting review Scope: secondary school detail pages only

Goal

Say what happened to a school's leavers after they left. Two sections on the secondary template:

  • After Year 11 — every secondary, from the KS4 destination measures
  • After the sixth form — sixth-form schools only, from the 16-18 measures

This replaces the "Post-16 destination data coming soon" placeholder standing in nextjs-app/components/school/SecondaryAdmissionsSection.tsx:117 since the exam phase taxonomy work, and fills the ks5_destinations_pct slot specified but never built in 2026-07-07-exam-phase-taxonomy-design.md:201.

Mockup, with all three data states live: https://claude.ai/code/artifact/5be149d6-252f-473c-9a4f-4c36b05161b0

The finding that shapes everything

Suppression is per cell, and the cells sum to the cohort.

DfE withholds a figure it considers disclosive by writing c. It does this at the level of an individual destination category, not the whole school, and it publishes the cohort total alongside. The categories form a clean partition. So where exactly one category is suppressed, subtracting the published ones from the cohort recovers it exactly.

Verified against three real schools in the 2022/23 file:

School URN Withheld Recovers to
North East Futures UTC 145900 School sixth form 3 pupils
Whitley Bay High School 108638 Further education 18 pupils
St Matthew's RC High School 148389 School sixth form 4 pupils

Those are the precise numbers the c exists to hide, and in a random 400-school sample 22% of mainstream secondaries have exactly one suppressed category in their disadvantaged group. This is the normal case, not an edge case.

Three rules follow, and everything else in this document is downstream of them.

R1 — Never publish enough to derive a remainder.

An earlier draft of this rule said "never render a derived remainder", and that was the defect code review caught in PR #137. Not drawing a number does nothing to stop it being computed: GET /api/schools/{urn} is public and unauthenticated, so anything in the payload is published whatever the UI chooses to draw. The rendering guards shipped; the payload still carried the cohort and every published category, and cohort - sum(published) returned Whitley Bay's withheld figure exactly.

The rule is therefore about the serialiser, and the UI guards are a second line of defence behind it. Two identities have to be closed:

  • within a pupil group the categories sum to the cohort, so a group with exactly one suppressed category gives it away;
  • across groups, disadvantaged + other = all for every category, so a category suppressed in exactly one of the three gives itself away.

_mask_for_disclosure applies DfE's own answer — secondary suppression — withholding a companion cell until every row and every column hides either none or at least two. It iterates, because each new suppression can break the other identity, and terminates because cells are only ever added.

The companion must carry pupils. Suppressing a zero looks like secondary suppression and protects nothing: the residual still equals the original withheld figure.

Where no companion can do the job — a sparse cohort whose every other category is not_applicable, routine in special schools and alternative provision — the pupil group is dropped from the payload entirely. A first version simply returned at that point with the violation intact and no signal, which review caught: a disclosure-control pass that fails silently is worse than none, because everything downstream trusts it. The function now cannot terminate except in a state where disclosure_invariant_holds() is true, and an exhaustive test sweeps all 81 suppression patterns of a four-category group to prove it.

Measured cost on the 400-school sample: the all-pupils bar survives on 94% of mainstream secondaries rather than 100%. That is the price of not republishing what DfE withheld.

R2 — Never aggregate across a suppression boundary. Summing published components to fill a gap is R1 with extra steps.

DfE's own aggregates (Sustained education destination, Sustained education, employment & apprenticeships) are ingested but not served. An aggregate spanning exactly one suppressed component names it, and nothing renders them today — an unused field that leaks is not a trade-off worth carrying. They can be re-added with their own guard if the fallback ladder is ever built.

R3 — The three pupil groups are one disclosure surface, not three. Disadvantaged and Not-known-to-be-disadvantaged partition All pupils, so rendering any two of them recovers the third. Where a category is suppressed in the disadvantaged group, it must therefore also be withheld from all other pupils — the all-pupils view is the primary one and keeps it.

This costs almost nothing, because DfE already applies the same masking: across the sample, 493 of 498 suppressed disadvantaged cells were suppressed in the other group too. The mart enforces the remaining 5, which fell on 2 schools of 262. The all-pupils bar is unaffected — masking the whole page wherever the disadvantaged group is thin would remove the bar from 80% of schools, and is not what this rule says.

R1 and R2 both hold within a group and still leak across the switch, which is why R3 is stated separately.

The convention that would break this quietly

macros/safe_numeric.sql coerces every EES sentinel — z, c, x, q, u — to NULL, deliberately and correctly for attainment, where "suppressed" and "no data" are equally unrenderable. Here they are not the same thing: one must print withheld, the other must print nothing at all, and the difference is what keeps R1 enforceable.

safe_numeric must not be used on destination counts. The staging model keeps the sentinel in a companion status column. This is the single most likely way for this feature to regress into a disclosure, so it gets its own dbt test.

What is actually available

Measured against the EES public API (open, no key). Both datasets carry geographicLevel: School with urn on every location option, so the join to dim_school is direct.

KS4 16-18
Dataset id 019d4f41-22d1-71b2-a1a7-f3b91026815b 019d4e73-6440-7523-b60c-bfab1ad4a30d
Rows 1,871,739 3,862,658
Institutions 4,946 3,065
Time periods 2009/10–2022/23 2016/17–2022/23

Destination categories (KS4). School sixth form · Sixth form college · Further education · Other education destination · Sustained apprenticeships (with level breakdown) · Sustained employment destination · Not recorded as a sustained destination · Activity not captured. Plus the aggregates Sustained education destination and Sustained education, employment & apprenticeships.

16-18 adds UK higher education institution and FE split by level, which is what makes the post-16 section worth having.

Breakdowns. Disadvantage Status gives Disadvantaged / Not known to be disadvantaged / Total — exactly the three-way switch. Sex, ethnicity, FSM status, prior attainment and SEN provision also travel in the same table; we ingest none of them.

Indicators. Both counts and percentages, plus the cohort size. Bar widths use the counts — the published percentages do not sum to 100.

Coverage, and what degrades

Random 400-school sample, 2022/23, mainstream secondaries (n=262):

View As published by DfE After R1–R3 masking Consequence
All pupils, all categories 100% 94% Bar works nearly everywhere
Disadvantaged, headline rate 95% 95% Gap panel works
Disadvantaged, three grouped cards 68% 68% Degrades card by card
Disadvantaged, all six categories 20% 20% Bar unusable for this group

The middle column is what the site actually serves. Masking costs the all-pupils bar on 6% of mainstream secondaries — those are schools where a category was suppressed in exactly one pupil group and no non-zero companion existed below the all-pupils row.

Special schools and alternative provision are far worse: 13% and 41% respectively have the whole cohort suppressed even for all pupils. The empty state is load-bearing, not defensive.

The display

Question-led. Three cards over one bar, with the cards acting as a lens on the bar rather than a summary beside it — hovering a card dims the bar, table and England reference to the categories that card is built from. The full mockup is linked above; what matters for implementation:

The headline is not the sustained rate. That figure sits between 92% and 97% for nearly every school in England. The mix is what varies, so the mix leads.

The grouping is ours, not DfE's. "Academic route" = school sixth form + sixth-form college; "College" = FE and other colleges; "Work" = apprenticeship + employment. This is the most arguable thing on the page, so it lives in one place in lib/destinations.ts, is explained in a tooltip, and is reversible in one edit.

The absence is hatched neutral, never a colour. "Activity not captured" means no record in the sources DfE holds — it includes independent schools, moving abroad and private training. Colouring it as a bad outcome would be a factual error rendered in CSS. The hatch also fixes a real contrast problem: neutral against the employment blue failed CVD separation at ΔE 7.6, and texture is the secondary encoding that rescues it. Every other adjacent pair clears ΔE 10.9 under protanopia.

Colour tokens. Education is one hue in three steps (school-like to college-like); apprenticeship and employment are separate hues. Six new tokens in globals.css, defined in both themes, per the existing token discipline.

The disadvantage split rides the same control. One visualisation serving three cohorts, with the England reference repointing to the matching national group. The gap statement stays visible below the bar whatever is selected, because a gap nobody clicks on is a gap nobody sees.

Data model

Extraction

A new tap-uk-ees-destinations extractor, separate from tap-uk-ees. The existing tap downloads a release ZIP and reads a CSV inside it; the destinations files are far larger than we need and the query API filters server-side, so this one POSTs to /v1/data-sets/{id}/query and pages through results.

With every dimension pinned — destination measures, disadvantage status, sex Total, characteristic topic Total — one year returns 252,610 rows across all geographic levels. Three school-level years is comfortably tractable.

Pinning is mandatory, not an optimisation: leaving the characteristic dimensions unconstrained returned 45 rows where 9 were wanted, because every breakdown shares one table.

The tap emits the raw value as text. It does not coerce c.

Staging

stg_ees_ks4_destinations / stg_ees_ks5_destinations. Each raw value becomes two columns:

case when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then raw::numeric end as pupils,
case
    when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then 'published'
    when lower(trim(raw)) = 'c'        then 'suppressed'
    else 'not_applicable'
end as status

Marts

fact_ks4_destinations and fact_ks5_destinations, long format:

urn, year, pupil_group, destination_category, cohort_pupils, pupils, percentage, status

This departs from the wide house pattern (fact_ks4_performance and friends) on purpose. pupil_group is a genuine third dimension; going wide would need three sets of every column, and R2 is far easier to test on rows than on columns.

Roughly 8 categories × 3 groups × 4,946 schools × 3 years ≈ 356k rows.

fact_destination_national carries the same grain for England, so the page's England reference repoints with the switch.

dbt tests

  • assert_destinations_no_derived_remainder — for every (urn, year, pupil_group) with exactly one suppressed category, assert no aggregate row exists that would let the residual be recovered. This is the R1 guard.
  • assert_destinations_group_masking — for every (urn, year, category), if the disadvantaged group carries suppressed, so does the other-pupils group. This is the R3 guard, applied in the mart so no consumer can reach an unmasked combination.
  • assert_destination_status_null_agreement — pupils is null wherever status != 'published', and never null where it is.
  • assert_destinations_join_dim_school — no orphaned URNs, matching the existing assert_no_orphaned_facts pattern.

API

GET /api/schools/{urn} gains a destinations block:

{
  "ks4": {
    "cohort_year": "2022/23",
    "published": "2026-04",
    "groups": {
      "all":           { "cohort": 180, "categories": [ … ], "aggregates": { … } },
      "disadvantaged": { … },
      "other":         { … }
    }
  },
  "ks5": { … }
}

Each category carries pupils, percentage and status. The serialiser never emits a computed remainder, and a backend test asserts that a group containing a suppressed category serialises no total that closes the gap.

null for the whole block where nothing is published — the frontend renders the empty state from its absence, not from a sentinel.

Frontend

File Kind Job
lib/destinations.ts pure Category list, the academic/college/work grouping, canAggregate() enforcing R2, percentage derivation from counts
components/school/DestinationsSection.tsx server Section shell, renders all pupils into the HTML
components/school/DestinationsView.tsx client Cohort switch, card↔bar linkage
components/school/Post16DestinationsSection.tsx server Year 13 section, sixth-form schools only
app/globals.css tokens Six destination colours, both themes

Server-first matches the directory's existing discipline — every component in components/school/ is a server component except AdmissionsViewToggle, which is the precedent this follows. All-pupils figures are in the HTML for crawlers and for no-JS; only the switch and the hover linkage need the client.

lib/schoolSections.ts gains hasKs4Destinations / hasKs5Destinations flags and the nav items, following the existing computeSchoolFlags pattern.

Placement on the secondary template: GCSE results → After Year 11 → After the sixth form → admissions. Destinations follow attainment because they answer "and then what happened".

Dating. The latest destination year is 2022/23, published April 2026, while the site's newest KS4 year is 2024/25. The section header states its own cohort year, or it reads as stale data next to the GCSE section above it.

Edge states

State Frequency Behaviour
Whole cohort suppressed 13% of special, 41% of AP Section renders the explanation, no chart
Some categories withheld 80% of disadvantaged views Cards degrade individually; no bar; table marks withheld rows
Disadvantaged group suppressed entirely 5% Switch drops to two options, gap panel not rendered
No sixth form — Post-16 section not rendered at all — absence is correct, a "no data" placeholder would imply something is missing
School too new — "First figures expected in 2026", not a bare no

Testing

Per CLAUDE.md, user-facing behaviour extends e2e/ in the same PR.

Unit — lib/destinations.ts is where R1 and R2 live, so it carries the heaviest tests: canAggregate() refuses a group containing one suppressed cell, allows one spanning two, and the bar builder refuses to emit segments for any group with suppression. These are the tests that must fail loudly if someone later "fixes" a gap in the chart.

dbt — the three tests above.

Backend — the serialiser emits no closing total for a partially suppressed group.

E2E — a school with full data renders three cards and a bar; a school with a partially suppressed disadvantaged group renders the withheld state and no bar element; a suppressed school renders the explanation; a school with no sixth form renders no post-16 section.

Note the staging caveat: mart changes are inert until the Airflow pipeline runs, and the staging E2E gate runs post-merge.

Out of scope

  • Compare view and rankings. The long mart shape supports both; neither is built here. Flagged because "% to a school sixth form" is a plausible rankings metric and the mart shape should not have to change to allow it.
  • Ethnicity, sex, SEN and prior-attainment breakdowns. Available in the same file, ingested deliberately not at all — each is a separate editorial decision about what a school page should assert.
  • Longer term destinations (3 and 5 years out) and Progression to higher education — separate publications, worth a later look for sixth forms.
  • Primary schools. No KS2 destination measures publication exists; DfE tracking starts at KS4. Naming the secondaries a primary's leavers go to needs the National Pupil Database, which is not publishable at that grain.

Risks

A later change reintroduces the disclosure. The likeliest routes are applying safe_numeric to a destination column for consistency, adding a coalesce in a mart, or — as happened in review — enforcing a disclosure rule at the rendering layer instead of the publishing layer. Mitigation is the dbt tests plus backend/tests/test_destinations_api.py, which reconstructs the residual the way an attacker would and asserts it no longer resolves.

The two-year lag reads as staleness. Mitigated by dating the cohort in the section header rather than only in a tooltip.

Sixth-form retention will be misread. "41% went to a school sixth form" says nothing about which school. The published file reports destination type, never destination institution. Copy must never imply "stayed on here", and the tooltip should say so.

Section length. The secondary template is already long and this adds two sections. If it becomes a problem the post-16 section is the one to collapse behind a disclosure, not the Year 11 one.

Open questions

  1. Is the disadvantage split its own section or a sub-block inside the destinations section? Modelled as a sub-block; it is the most differentiating figure on the page and the most easily misread on a small cohort.
  2. Do we ingest the apprenticeship level breakdown (intermediate / advanced / higher) now, or collapse to one apprenticeship figure and revisit? Collapsed in this design.