From 576013d627e88b7f299f100741c651431707eddb Mon Sep 17 00:00:00 2001 From: Tudor Date: Fri, 28 Aug 2026 15:28:09 +0100 Subject: [PATCH] docs(destinations): design for KS4 and post-16 destination measures MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The published files suppress individual cells, not whole cohorts, and the categories sum to the cohort — so on 22% of mainstream secondaries the withheld figure can be recovered by subtraction. Three disclosure rules fall out of that, and the rest of the design is downstream of them. Verified against the EES API rather than assumed: both datasets carry school-level rows keyed by URN, with the disadvantage split and every destination category the display needs. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob --- .../2026-08-28-destination-measures-design.md | 352 ++++++++++++++++++ 1 file changed, 352 insertions(+) create mode 100644 docs/superpowers/specs/2026-08-28-destination-measures-design.md diff --git a/docs/superpowers/specs/2026-08-28-destination-measures-design.md b/docs/superpowers/specs/2026-08-28-destination-measures-design.md new file mode 100644 index 0000000..22ab448 --- /dev/null +++ b/docs/superpowers/specs/2026-08-28-destination-measures-design.md @@ -0,0 +1,352 @@ +# Destination Measures — Design + +**Date:** 2026-08-28 +**Status:** awaiting review +**Scope:** secondary school detail pages only + +## Goal + +Say what happened to a school's leavers after they left. Two sections on the +secondary template: + +- **After Year 11** — every secondary, from the KS4 destination measures +- **After the sixth form** — sixth-form schools only, from the 16-18 measures + +This replaces the "Post-16 destination data coming soon" placeholder standing in +`nextjs-app/components/school/SecondaryAdmissionsSection.tsx:117` since the exam +phase taxonomy work, and fills the `ks5_destinations_pct` slot specified but +never built in `2026-07-07-exam-phase-taxonomy-design.md:201`. + +Mockup, with all three data states live: + + +## The finding that shapes everything + +**Suppression is per cell, and the cells sum to the cohort.** + +DfE withholds a figure it considers disclosive by writing `c`. It does this at +the level of an individual destination category, not the whole school, and it +publishes the cohort total alongside. The categories form a clean partition. So +where exactly one category is suppressed, subtracting the published ones from the +cohort recovers it exactly. + +Verified against three real schools in the 2022/23 file: + +| School | URN | Withheld | Recovers to | +|---|---|---|---| +| North East Futures UTC | 145900 | School sixth form | **3 pupils** | +| Whitley Bay High School | 108638 | Further education | **18 pupils** | +| St Matthew's RC High School | 148389 | School sixth form | **4 pupils** | + +Those are the precise numbers the `c` exists to hide, and in a random 400-school +sample **22% of mainstream secondaries** have exactly one suppressed category in +their disadvantaged group. This is the normal case, not an edge case. + +Two rules follow, and everything else in this document is downstream of them. + +**R1 — Never render a derived remainder.** Not as a number, and not as a bar +segment: a segment sized by the residual can be read straight off the axis. Where +any category in a pupil group is suppressed, the page shows the published +categories, says the rest are withheld, and stops. + +**R2 — Never aggregate across a suppression boundary.** A group total is +publishable only when DfE published that total itself, or when the aggregate +spans **two or more** suppressed cells. Summing published components to fill a +gap is R1 with extra steps. + +**R3 — The three pupil groups are one disclosure surface, not three.** +Disadvantaged and Not-known-to-be-disadvantaged partition All pupils, so +rendering any *two* of them recovers the third. Where a category is suppressed in +the disadvantaged group, it must therefore also be withheld from **all other +pupils** — the all-pupils view is the primary one and keeps it. + +This costs almost nothing, because DfE already applies the same masking: across +the sample, 493 of 498 suppressed disadvantaged cells were suppressed in the +other group too. The mart enforces the remaining 5, which fell on 2 schools of +262. **The all-pupils bar is unaffected** — masking the whole page wherever the +disadvantaged group is thin would remove the bar from 80% of schools, and is not +what this rule says. + +R1 and R2 both hold within a group and still leak across the switch, which is why +R3 is stated separately. + +### The convention that would break this quietly + +`macros/safe_numeric.sql` coerces every EES sentinel — `z`, `c`, `x`, `q`, `u` — +to `NULL`, deliberately and correctly for attainment, where "suppressed" and "no +data" are equally unrenderable. Here they are not the same thing: one must print +*withheld*, the other must print nothing at all, and the difference is what keeps +R1 enforceable. + +**`safe_numeric` must not be used on destination counts.** The staging model +keeps the sentinel in a companion status column. This is the single most likely +way for this feature to regress into a disclosure, so it gets its own dbt test. + +## What is actually available + +Measured against the EES public API (open, no key). Both datasets carry +`geographicLevel: School` with `urn` on every location option, so the join to +`dim_school` is direct. + +| | KS4 | 16-18 | +|---|---|---| +| Dataset id | `019d4f41-22d1-71b2-a1a7-f3b91026815b` | `019d4e73-6440-7523-b60c-bfab1ad4a30d` | +| Rows | 1,871,739 | 3,862,658 | +| Institutions | 4,946 | 3,065 | +| Time periods | 2009/10–2022/23 | 2016/17–2022/23 | + +**Destination categories (KS4).** School sixth form · Sixth form college · +Further education · Other education destination · Sustained apprenticeships (with +level breakdown) · Sustained employment destination · Not recorded as a sustained +destination · Activity not captured. Plus the aggregates `Sustained education +destination` and `Sustained education, employment & apprenticeships`. + +**16-18 adds** UK higher education institution and FE split by level, which is +what makes the post-16 section worth having. + +**Breakdowns.** `Disadvantage Status` gives Disadvantaged / Not known to be +disadvantaged / Total — exactly the three-way switch. Sex, ethnicity, FSM status, +prior attainment and SEN provision also travel in the same table; we ingest none +of them. + +**Indicators.** Both counts and percentages, plus the cohort size. Bar widths use +the counts — the published percentages do not sum to 100. + +### Coverage, and what degrades + +Random 400-school sample, 2022/23, mainstream secondaries (n=262): + +| View | Published | Consequence | +|---|---|---| +| All pupils, all categories | **100%** | Full bar works everywhere | +| Disadvantaged, headline rate | 95% | Gap panel works | +| Disadvantaged, three grouped cards | 68% | Degrades card by card | +| Disadvantaged, all six categories | **20%** | Bar unusable for this group | + +Special schools and alternative provision are far worse: 13% and 41% respectively +have the whole cohort suppressed even for all pupils. The empty state is +load-bearing, not defensive. + +## The display + +Question-led. Three cards over one bar, with the cards acting as a lens on the +bar rather than a summary beside it — hovering a card dims the bar, table and +England reference to the categories that card is built from. The full mockup is +linked above; what matters for implementation: + +**The headline is not the sustained rate.** That figure sits between 92% and 97% +for nearly every school in England. The mix is what varies, so the mix leads. + +**The grouping is ours, not DfE's.** "Academic route" = school sixth form + +sixth-form college; "College" = FE and other colleges; "Work" = apprenticeship + +employment. This is the most arguable thing on the page, so it lives in one place +in `lib/destinations.ts`, is explained in a tooltip, and is reversible in one +edit. + +**The absence is hatched neutral, never a colour.** "Activity not captured" means +no record in the sources DfE holds — it includes independent schools, moving +abroad and private training. Colouring it as a bad outcome would be a factual +error rendered in CSS. The hatch also fixes a real contrast problem: neutral +against the employment blue failed CVD separation at ΔE 7.6, and texture is the +secondary encoding that rescues it. Every other adjacent pair clears ΔE 10.9 +under protanopia. + +**Colour tokens.** Education is one hue in three steps (school-like to +college-like); apprenticeship and employment are separate hues. Six new tokens in +`globals.css`, defined in both themes, per the existing token discipline. + +**The disadvantage split rides the same control.** One visualisation serving +three cohorts, with the England reference repointing to the matching national +group. The gap statement stays visible below the bar whatever is selected, +because a gap nobody clicks on is a gap nobody sees. + +## Data model + +### Extraction + +A new `tap-uk-ees-destinations` extractor, separate from `tap-uk-ees`. The +existing tap downloads a release ZIP and reads a CSV inside it; the destinations +files are far larger than we need and the query API filters server-side, so this +one POSTs to `/v1/data-sets/{id}/query` and pages through results. + +With every dimension pinned — destination measures, disadvantage status, sex +Total, characteristic topic Total — one year returns **252,610 rows** across all +geographic levels. Three school-level years is comfortably tractable. + +Pinning is mandatory, not an optimisation: leaving the characteristic dimensions +unconstrained returned 45 rows where 9 were wanted, because every breakdown +shares one table. + +The tap emits the raw value as text. **It does not coerce `c`.** + +### Staging + +`stg_ees_ks4_destinations` / `stg_ees_ks5_destinations`. Each raw value becomes +two columns: + +```sql +case when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then raw::numeric end as pupils, +case + when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then 'published' + when lower(trim(raw)) = 'c' then 'suppressed' + else 'not_applicable' +end as status +``` + +### Marts + +`fact_ks4_destinations` and `fact_ks5_destinations`, **long format**: + +``` +urn, year, pupil_group, destination_category, cohort_pupils, pupils, percentage, status +``` + +This departs from the wide house pattern (`fact_ks4_performance` and friends) on +purpose. `pupil_group` is a genuine third dimension; going wide would need three +sets of every column, and R2 is far easier to test on rows than on columns. + +Roughly 8 categories × 3 groups × 4,946 schools × 3 years ≈ 356k rows. + +`fact_destination_national` carries the same grain for England, so the page's +England reference repoints with the switch. + +### dbt tests + +- `assert_destinations_no_derived_remainder` — for every (urn, year, + pupil_group) with exactly one suppressed category, assert no aggregate row + exists that would let the residual be recovered. **This is the R1 guard.** +- `assert_destinations_group_masking` — for every (urn, year, category), if the + disadvantaged group carries `suppressed`, so does the other-pupils group. + **This is the R3 guard**, applied in the mart so no consumer can reach an + unmasked combination. +- `assert_destination_status_null_agreement` — `pupils is null` wherever + `status != 'published'`, and never null where it is. +- `assert_destinations_join_dim_school` — no orphaned URNs, matching the + existing `assert_no_orphaned_facts` pattern. + +## API + +`GET /api/schools/{urn}` gains a `destinations` block: + +```json +{ + "ks4": { + "cohort_year": "2022/23", + "published": "2026-04", + "groups": { + "all": { "cohort": 180, "categories": [ … ], "aggregates": { … } }, + "disadvantaged": { … }, + "other": { … } + } + }, + "ks5": { … } +} +``` + +Each category carries `pupils`, `percentage` and `status`. **The serialiser never +emits a computed remainder**, and a backend test asserts that a group containing a +suppressed category serialises no total that closes the gap. + +`null` for the whole block where nothing is published — the frontend renders the +empty state from its absence, not from a sentinel. + +## Frontend + +| File | Kind | Job | +|---|---|---| +| `lib/destinations.ts` | pure | Category list, the academic/college/work grouping, `canAggregate()` enforcing R2, percentage derivation from counts | +| `components/school/DestinationsSection.tsx` | server | Section shell, renders **all pupils** into the HTML | +| `components/school/DestinationsView.tsx` | client | Cohort switch, card↔bar linkage | +| `components/school/Post16DestinationsSection.tsx` | server | Year 13 section, sixth-form schools only | +| `app/globals.css` | tokens | Six destination colours, both themes | + +Server-first matches the directory's existing discipline — every component in +`components/school/` is a server component except `AdmissionsViewToggle`, which +is the precedent this follows. All-pupils figures are in the HTML for crawlers +and for no-JS; only the switch and the hover linkage need the client. + +`lib/schoolSections.ts` gains `hasKs4Destinations` / `hasKs5Destinations` flags +and the nav items, following the existing `computeSchoolFlags` pattern. + +**Placement** on the secondary template: GCSE results → After Year 11 → After the +sixth form → admissions. Destinations follow attainment because they answer "and +then what happened". + +**Dating.** The latest destination year is 2022/23, published April 2026, while +the site's newest KS4 year is 2024/25. The section header states its own cohort +year, or it reads as stale data next to the GCSE section above it. + +## Edge states + +| State | Frequency | Behaviour | +|---|---|---| +| Whole cohort suppressed | 13% of special, 41% of AP | Section renders the explanation, no chart | +| Some categories withheld | 80% of disadvantaged views | Cards degrade individually; **no bar**; table marks withheld rows | +| Disadvantaged group suppressed entirely | 5% | Switch drops to two options, gap panel not rendered | +| No sixth form | — | Post-16 section not rendered at all — absence is correct, a "no data" placeholder would imply something is missing | +| School too new | — | "First figures expected in 2026", not a bare no | + +## Testing + +Per CLAUDE.md, user-facing behaviour extends `e2e/` in the same PR. + +**Unit** — `lib/destinations.ts` is where R1 and R2 live, so it carries the +heaviest tests: `canAggregate()` refuses a group containing one suppressed cell, +allows one spanning two, and the bar builder refuses to emit segments for any +group with suppression. These are the tests that must fail loudly if someone +later "fixes" a gap in the chart. + +**dbt** — the three tests above. + +**Backend** — the serialiser emits no closing total for a partially suppressed +group. + +**E2E** — a school with full data renders three cards and a bar; a school with a +partially suppressed disadvantaged group renders the withheld state and **no bar +element**; a suppressed school renders the explanation; a school with no sixth +form renders no post-16 section. + +Note the staging caveat: mart changes are inert until the Airflow pipeline runs, +and the staging E2E gate runs post-merge. + +## Out of scope + +- **Compare view and rankings.** The long mart shape supports both; neither is + built here. Flagged because "% to a school sixth form" is a plausible rankings + metric and the mart shape should not have to change to allow it. +- **Ethnicity, sex, SEN and prior-attainment breakdowns.** Available in the same + file, ingested deliberately not at all — each is a separate editorial decision + about what a school page should assert. +- **Longer term destinations** (3 and 5 years out) and **Progression to higher + education** — separate publications, worth a later look for sixth forms. +- **Primary schools.** No KS2 destination measures publication exists; DfE + tracking starts at KS4. Naming the secondaries a primary's leavers go to needs + the National Pupil Database, which is not publishable at that grain. + +## Risks + +**A later change reintroduces the disclosure.** The likeliest route is someone +applying `safe_numeric` to a destination column for consistency, or adding a +`coalesce` in a mart. Mitigation is the dbt test plus the unit tests on +`canAggregate()` — the rule has to be executable, not documentary. + +**The two-year lag reads as staleness.** Mitigated by dating the cohort in the +section header rather than only in a tooltip. + +**Sixth-form retention will be misread.** "41% went to a school sixth form" says +nothing about *which* school. The published file reports destination type, never +destination institution. Copy must never imply "stayed on here", and the tooltip +should say so. + +**Section length.** The secondary template is already long and this adds two +sections. If it becomes a problem the post-16 section is the one to collapse +behind a disclosure, not the Year 11 one. + +## Open questions + +1. Is the disadvantage split its own section or a sub-block inside the + destinations section? Modelled as a sub-block; it is the most differentiating + figure on the page and the most easily misread on a small cohort. +2. Do we ingest the apprenticeship level breakdown (intermediate / advanced / + higher) now, or collapse to one apprenticeship figure and revisit? Collapsed + in this design.