PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 7m23s
Review found _mask_for_disclosure could return with its invariant broken and say nothing. add_companion only ever withheld a *published* cell, so a group with one suppressed category and every other one not_applicable — routine in special schools and AP, where few categories apply — left the loop with the lone suppressed cell still solvable. Reproduced on a nine-pupil cohort: one hidden cell, cohort served, residual intact. A disclosure-control pass that fails silently is worse than none, because everything downstream trusts it. The loop now runs until the invariant holds and escalates when no companion exists: the pupil group is dropped from the payload, and an empty block serialises as None so the section is absent rather than an empty shell. disclosure_invariant_holds() is exported so tests assert it directly instead of re-deriving it, and an exhaustive test sweeps all 81 suppression patterns of a four-category group. Also fixes a test that set up six measures and checked one: the loop was `for measure in ["school_sixth_form"]`. It now checks every measure, and against the real invariant — none hidden, or at least two, rather than "at least two", which the five published measures would have failed. No regression on real data: 262 mainstream secondaries, all-pupils bar still drawable on 94%, zero invariant violations, one disadvantaged group dropped by the new escalation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
400 lines
18 KiB
Markdown
400 lines
18 KiB
Markdown
# Destination Measures — Design
|
||
|
||
**Date:** 2026-08-28
|
||
**Status:** awaiting review
|
||
**Scope:** secondary school detail pages only
|
||
|
||
## Goal
|
||
|
||
Say what happened to a school's leavers after they left. Two sections on the
|
||
secondary template:
|
||
|
||
- **After Year 11** — every secondary, from the KS4 destination measures
|
||
- **After the sixth form** — sixth-form schools only, from the 16-18 measures
|
||
|
||
This replaces the "Post-16 destination data coming soon" placeholder standing in
|
||
`nextjs-app/components/school/SecondaryAdmissionsSection.tsx:117` since the exam
|
||
phase taxonomy work, and fills the `ks5_destinations_pct` slot specified but
|
||
never built in `2026-07-07-exam-phase-taxonomy-design.md:201`.
|
||
|
||
Mockup, with all three data states live:
|
||
<https://claude.ai/code/artifact/5be149d6-252f-473c-9a4f-4c36b05161b0>
|
||
|
||
## The finding that shapes everything
|
||
|
||
**Suppression is per cell, and the cells sum to the cohort.**
|
||
|
||
DfE withholds a figure it considers disclosive by writing `c`. It does this at
|
||
the level of an individual destination category, not the whole school, and it
|
||
publishes the cohort total alongside. The categories form a clean partition. So
|
||
where exactly one category is suppressed, subtracting the published ones from the
|
||
cohort recovers it exactly.
|
||
|
||
Verified against three real schools in the 2022/23 file:
|
||
|
||
| School | URN | Withheld | Recovers to |
|
||
|---|---|---|---|
|
||
| North East Futures UTC | 145900 | School sixth form | **3 pupils** |
|
||
| Whitley Bay High School | 108638 | Further education | **18 pupils** |
|
||
| St Matthew's RC High School | 148389 | School sixth form | **4 pupils** |
|
||
|
||
Those are the precise numbers the `c` exists to hide, and in a random 400-school
|
||
sample **22% of mainstream secondaries** have exactly one suppressed category in
|
||
their disadvantaged group. This is the normal case, not an edge case.
|
||
|
||
Three rules follow, and everything else in this document is downstream of them.
|
||
|
||
**R1 — Never *publish* enough to derive a remainder.**
|
||
|
||
An earlier draft of this rule said "never *render* a derived remainder", and
|
||
that was the defect code review caught in PR #137. Not drawing a number does
|
||
nothing to stop it being computed: `GET /api/schools/{urn}` is public and
|
||
unauthenticated, so anything in the payload is published whatever the UI
|
||
chooses to draw. The rendering guards shipped; the payload still carried the
|
||
cohort and every published category, and `cohort - sum(published)` returned
|
||
Whitley Bay's withheld figure exactly.
|
||
|
||
The rule is therefore about the serialiser, and the UI guards are a second line
|
||
of defence behind it. Two identities have to be closed:
|
||
|
||
- within a pupil group the categories sum to the cohort, so a group with
|
||
exactly **one** suppressed category gives it away;
|
||
- across groups, disadvantaged + other = all for every category, so a category
|
||
suppressed in exactly **one** of the three gives itself away.
|
||
|
||
`_mask_for_disclosure` applies DfE's own answer — secondary suppression —
|
||
withholding a companion cell until every row and every column hides either none
|
||
or at least two. It iterates, because each new suppression can break the other
|
||
identity, and terminates because cells are only ever added.
|
||
|
||
The companion must carry pupils. Suppressing a zero looks like secondary
|
||
suppression and protects nothing: the residual still equals the original
|
||
withheld figure.
|
||
|
||
Where no companion can do the job — a sparse cohort whose every other category
|
||
is `not_applicable`, routine in special schools and alternative provision — the
|
||
pupil group is **dropped from the payload entirely**. A first version simply
|
||
returned at that point with the violation intact and no signal, which review
|
||
caught: a disclosure-control pass that fails silently is worse than none,
|
||
because everything downstream trusts it. The function now cannot terminate
|
||
except in a state where `disclosure_invariant_holds()` is true, and an
|
||
exhaustive test sweeps all 81 suppression patterns of a four-category group to
|
||
prove it.
|
||
|
||
Measured cost on the 400-school sample: the all-pupils bar survives on **94%**
|
||
of mainstream secondaries rather than 100%. That is the price of not
|
||
republishing what DfE withheld.
|
||
|
||
**R2 — Never aggregate across a suppression boundary.** Summing published
|
||
components to fill a gap is R1 with extra steps.
|
||
|
||
DfE's own aggregates (`Sustained education destination`, `Sustained education,
|
||
employment & apprenticeships`) are ingested but **not served**. An aggregate
|
||
spanning exactly one suppressed component names it, and nothing renders them
|
||
today — an unused field that leaks is not a trade-off worth carrying. They can
|
||
be re-added with their own guard if the fallback ladder is ever built.
|
||
|
||
**R3 — The three pupil groups are one disclosure surface, not three.**
|
||
Disadvantaged and Not-known-to-be-disadvantaged partition All pupils, so
|
||
rendering any *two* of them recovers the third. Where a category is suppressed in
|
||
the disadvantaged group, it must therefore also be withheld from **all other
|
||
pupils** — the all-pupils view is the primary one and keeps it.
|
||
|
||
This costs almost nothing, because DfE already applies the same masking: across
|
||
the sample, 493 of 498 suppressed disadvantaged cells were suppressed in the
|
||
other group too. The mart enforces the remaining 5, which fell on 2 schools of
|
||
262. **The all-pupils bar is unaffected** — masking the whole page wherever the
|
||
disadvantaged group is thin would remove the bar from 80% of schools, and is not
|
||
what this rule says.
|
||
|
||
R1 and R2 both hold within a group and still leak across the switch, which is why
|
||
R3 is stated separately.
|
||
|
||
### The convention that would break this quietly
|
||
|
||
`macros/safe_numeric.sql` coerces every EES sentinel — `z`, `c`, `x`, `q`, `u` —
|
||
to `NULL`, deliberately and correctly for attainment, where "suppressed" and "no
|
||
data" are equally unrenderable. Here they are not the same thing: one must print
|
||
*withheld*, the other must print nothing at all, and the difference is what keeps
|
||
R1 enforceable.
|
||
|
||
**`safe_numeric` must not be used on destination counts.** The staging model
|
||
keeps the sentinel in a companion status column. This is the single most likely
|
||
way for this feature to regress into a disclosure, so it gets its own dbt test.
|
||
|
||
## What is actually available
|
||
|
||
Measured against the EES public API (open, no key). Both datasets carry
|
||
`geographicLevel: School` with `urn` on every location option, so the join to
|
||
`dim_school` is direct.
|
||
|
||
| | KS4 | 16-18 |
|
||
|---|---|---|
|
||
| Dataset id | `019d4f41-22d1-71b2-a1a7-f3b91026815b` | `019d4e73-6440-7523-b60c-bfab1ad4a30d` |
|
||
| Rows | 1,871,739 | 3,862,658 |
|
||
| Institutions | 4,946 | 3,065 |
|
||
| Time periods | 2009/10–2022/23 | 2016/17–2022/23 |
|
||
|
||
**Destination categories (KS4).** School sixth form · Sixth form college ·
|
||
Further education · Other education destination · Sustained apprenticeships (with
|
||
level breakdown) · Sustained employment destination · Not recorded as a sustained
|
||
destination · Activity not captured. Plus the aggregates `Sustained education
|
||
destination` and `Sustained education, employment & apprenticeships`.
|
||
|
||
**16-18 adds** UK higher education institution and FE split by level, which is
|
||
what makes the post-16 section worth having.
|
||
|
||
**Breakdowns.** `Disadvantage Status` gives Disadvantaged / Not known to be
|
||
disadvantaged / Total — exactly the three-way switch. Sex, ethnicity, FSM status,
|
||
prior attainment and SEN provision also travel in the same table; we ingest none
|
||
of them.
|
||
|
||
**Indicators.** Both counts and percentages, plus the cohort size. Bar widths use
|
||
the counts — the published percentages do not sum to 100.
|
||
|
||
### Coverage, and what degrades
|
||
|
||
Random 400-school sample, 2022/23, mainstream secondaries (n=262):
|
||
|
||
| View | As published by DfE | After R1–R3 masking | Consequence |
|
||
|---|---|---|---|
|
||
| All pupils, all categories | 100% | **94%** | Bar works nearly everywhere |
|
||
| Disadvantaged, headline rate | 95% | 95% | Gap panel works |
|
||
| Disadvantaged, three grouped cards | 68% | 68% | Degrades card by card |
|
||
| Disadvantaged, all six categories | 20% | **20%** | Bar unusable for this group |
|
||
|
||
The middle column is what the site actually serves. Masking costs the
|
||
all-pupils bar on 6% of mainstream secondaries — those are schools where a
|
||
category was suppressed in exactly one pupil group and no non-zero companion
|
||
existed below the all-pupils row.
|
||
|
||
Special schools and alternative provision are far worse: 13% and 41% respectively
|
||
have the whole cohort suppressed even for all pupils. The empty state is
|
||
load-bearing, not defensive.
|
||
|
||
## The display
|
||
|
||
Question-led. Three cards over one bar, with the cards acting as a lens on the
|
||
bar rather than a summary beside it — hovering a card dims the bar, table and
|
||
England reference to the categories that card is built from. The full mockup is
|
||
linked above; what matters for implementation:
|
||
|
||
**The headline is not the sustained rate.** That figure sits between 92% and 97%
|
||
for nearly every school in England. The mix is what varies, so the mix leads.
|
||
|
||
**The grouping is ours, not DfE's.** "Academic route" = school sixth form +
|
||
sixth-form college; "College" = FE and other colleges; "Work" = apprenticeship +
|
||
employment. This is the most arguable thing on the page, so it lives in one place
|
||
in `lib/destinations.ts`, is explained in a tooltip, and is reversible in one
|
||
edit.
|
||
|
||
**The absence is hatched neutral, never a colour.** "Activity not captured" means
|
||
no record in the sources DfE holds — it includes independent schools, moving
|
||
abroad and private training. Colouring it as a bad outcome would be a factual
|
||
error rendered in CSS. The hatch also fixes a real contrast problem: neutral
|
||
against the employment blue failed CVD separation at ΔE 7.6, and texture is the
|
||
secondary encoding that rescues it. Every other adjacent pair clears ΔE 10.9
|
||
under protanopia.
|
||
|
||
**Colour tokens.** Education is one hue in three steps (school-like to
|
||
college-like); apprenticeship and employment are separate hues. Six new tokens in
|
||
`globals.css`, defined in both themes, per the existing token discipline.
|
||
|
||
**The disadvantage split rides the same control.** One visualisation serving
|
||
three cohorts, with the England reference repointing to the matching national
|
||
group. The gap statement stays visible below the bar whatever is selected,
|
||
because a gap nobody clicks on is a gap nobody sees.
|
||
|
||
## Data model
|
||
|
||
### Extraction
|
||
|
||
A new `tap-uk-ees-destinations` extractor, separate from `tap-uk-ees`. The
|
||
existing tap downloads a release ZIP and reads a CSV inside it; the destinations
|
||
files are far larger than we need and the query API filters server-side, so this
|
||
one POSTs to `/v1/data-sets/{id}/query` and pages through results.
|
||
|
||
With every dimension pinned — destination measures, disadvantage status, sex
|
||
Total, characteristic topic Total — one year returns **252,610 rows** across all
|
||
geographic levels. Three school-level years is comfortably tractable.
|
||
|
||
Pinning is mandatory, not an optimisation: leaving the characteristic dimensions
|
||
unconstrained returned 45 rows where 9 were wanted, because every breakdown
|
||
shares one table.
|
||
|
||
The tap emits the raw value as text. **It does not coerce `c`.**
|
||
|
||
### Staging
|
||
|
||
`stg_ees_ks4_destinations` / `stg_ees_ks5_destinations`. Each raw value becomes
|
||
two columns:
|
||
|
||
```sql
|
||
case when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then raw::numeric end as pupils,
|
||
case
|
||
when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then 'published'
|
||
when lower(trim(raw)) = 'c' then 'suppressed'
|
||
else 'not_applicable'
|
||
end as status
|
||
```
|
||
|
||
### Marts
|
||
|
||
`fact_ks4_destinations` and `fact_ks5_destinations`, **long format**:
|
||
|
||
```
|
||
urn, year, pupil_group, destination_category, cohort_pupils, pupils, percentage, status
|
||
```
|
||
|
||
This departs from the wide house pattern (`fact_ks4_performance` and friends) on
|
||
purpose. `pupil_group` is a genuine third dimension; going wide would need three
|
||
sets of every column, and R2 is far easier to test on rows than on columns.
|
||
|
||
Roughly 8 categories × 3 groups × 4,946 schools × 3 years ≈ 356k rows.
|
||
|
||
`fact_destination_national` carries the same grain for England, so the page's
|
||
England reference repoints with the switch.
|
||
|
||
### dbt tests
|
||
|
||
- `assert_destinations_no_derived_remainder` — for every (urn, year,
|
||
pupil_group) with exactly one suppressed category, assert no aggregate row
|
||
exists that would let the residual be recovered. **This is the R1 guard.**
|
||
- `assert_destinations_group_masking` — for every (urn, year, category), if the
|
||
disadvantaged group carries `suppressed`, so does the other-pupils group.
|
||
**This is the R3 guard**, applied in the mart so no consumer can reach an
|
||
unmasked combination.
|
||
- `assert_destination_status_null_agreement` — `pupils is null` wherever
|
||
`status != 'published'`, and never null where it is.
|
||
- `assert_destinations_join_dim_school` — no orphaned URNs, matching the
|
||
existing `assert_no_orphaned_facts` pattern.
|
||
|
||
## API
|
||
|
||
`GET /api/schools/{urn}` gains a `destinations` block:
|
||
|
||
```json
|
||
{
|
||
"ks4": {
|
||
"cohort_year": "2022/23",
|
||
"published": "2026-04",
|
||
"groups": {
|
||
"all": { "cohort": 180, "categories": [ … ], "aggregates": { … } },
|
||
"disadvantaged": { … },
|
||
"other": { … }
|
||
}
|
||
},
|
||
"ks5": { … }
|
||
}
|
||
```
|
||
|
||
Each category carries `pupils`, `percentage` and `status`. **The serialiser never
|
||
emits a computed remainder**, and a backend test asserts that a group containing a
|
||
suppressed category serialises no total that closes the gap.
|
||
|
||
`null` for the whole block where nothing is published — the frontend renders the
|
||
empty state from its absence, not from a sentinel.
|
||
|
||
## Frontend
|
||
|
||
| File | Kind | Job |
|
||
|---|---|---|
|
||
| `lib/destinations.ts` | pure | Category list, the academic/college/work grouping, `canAggregate()` enforcing R2, percentage derivation from counts |
|
||
| `components/school/DestinationsSection.tsx` | server | Section shell, renders **all pupils** into the HTML |
|
||
| `components/school/DestinationsView.tsx` | client | Cohort switch, card↔bar linkage |
|
||
| `components/school/Post16DestinationsSection.tsx` | server | Year 13 section, sixth-form schools only |
|
||
| `app/globals.css` | tokens | Six destination colours, both themes |
|
||
|
||
Server-first matches the directory's existing discipline — every component in
|
||
`components/school/` is a server component except `AdmissionsViewToggle`, which
|
||
is the precedent this follows. All-pupils figures are in the HTML for crawlers
|
||
and for no-JS; only the switch and the hover linkage need the client.
|
||
|
||
`lib/schoolSections.ts` gains `hasKs4Destinations` / `hasKs5Destinations` flags
|
||
and the nav items, following the existing `computeSchoolFlags` pattern.
|
||
|
||
**Placement** on the secondary template: GCSE results → After Year 11 → After the
|
||
sixth form → admissions. Destinations follow attainment because they answer "and
|
||
then what happened".
|
||
|
||
**Dating.** The latest destination year is 2022/23, published April 2026, while
|
||
the site's newest KS4 year is 2024/25. The section header states its own cohort
|
||
year, or it reads as stale data next to the GCSE section above it.
|
||
|
||
## Edge states
|
||
|
||
| State | Frequency | Behaviour |
|
||
|---|---|---|
|
||
| Whole cohort suppressed | 13% of special, 41% of AP | Section renders the explanation, no chart |
|
||
| Some categories withheld | 80% of disadvantaged views | Cards degrade individually; **no bar**; table marks withheld rows |
|
||
| Disadvantaged group suppressed entirely | 5% | Switch drops to two options, gap panel not rendered |
|
||
| No sixth form | — | Post-16 section not rendered at all — absence is correct, a "no data" placeholder would imply something is missing |
|
||
| School too new | — | "First figures expected in 2026", not a bare no |
|
||
|
||
## Testing
|
||
|
||
Per CLAUDE.md, user-facing behaviour extends `e2e/` in the same PR.
|
||
|
||
**Unit** — `lib/destinations.ts` is where R1 and R2 live, so it carries the
|
||
heaviest tests: `canAggregate()` refuses a group containing one suppressed cell,
|
||
allows one spanning two, and the bar builder refuses to emit segments for any
|
||
group with suppression. These are the tests that must fail loudly if someone
|
||
later "fixes" a gap in the chart.
|
||
|
||
**dbt** — the three tests above.
|
||
|
||
**Backend** — the serialiser emits no closing total for a partially suppressed
|
||
group.
|
||
|
||
**E2E** — a school with full data renders three cards and a bar; a school with a
|
||
partially suppressed disadvantaged group renders the withheld state and **no bar
|
||
element**; a suppressed school renders the explanation; a school with no sixth
|
||
form renders no post-16 section.
|
||
|
||
Note the staging caveat: mart changes are inert until the Airflow pipeline runs,
|
||
and the staging E2E gate runs post-merge.
|
||
|
||
## Out of scope
|
||
|
||
- **Compare view and rankings.** The long mart shape supports both; neither is
|
||
built here. Flagged because "% to a school sixth form" is a plausible rankings
|
||
metric and the mart shape should not have to change to allow it.
|
||
- **Ethnicity, sex, SEN and prior-attainment breakdowns.** Available in the same
|
||
file, ingested deliberately not at all — each is a separate editorial decision
|
||
about what a school page should assert.
|
||
- **Longer term destinations** (3 and 5 years out) and **Progression to higher
|
||
education** — separate publications, worth a later look for sixth forms.
|
||
- **Primary schools.** No KS2 destination measures publication exists; DfE
|
||
tracking starts at KS4. Naming the secondaries a primary's leavers go to needs
|
||
the National Pupil Database, which is not publishable at that grain.
|
||
|
||
## Risks
|
||
|
||
**A later change reintroduces the disclosure.** The likeliest routes are
|
||
applying `safe_numeric` to a destination column for consistency, adding a
|
||
`coalesce` in a mart, or — as happened in review — enforcing a disclosure rule
|
||
at the rendering layer instead of the publishing layer. Mitigation is the dbt
|
||
tests plus `backend/tests/test_destinations_api.py`, which reconstructs the
|
||
residual the way an attacker would and asserts it no longer resolves.
|
||
|
||
**The two-year lag reads as staleness.** Mitigated by dating the cohort in the
|
||
section header rather than only in a tooltip.
|
||
|
||
**Sixth-form retention will be misread.** "41% went to a school sixth form" says
|
||
nothing about *which* school. The published file reports destination type, never
|
||
destination institution. Copy must never imply "stayed on here", and the tooltip
|
||
should say so.
|
||
|
||
**Section length.** The secondary template is already long and this adds two
|
||
sections. If it becomes a problem the post-16 section is the one to collapse
|
||
behind a disclosure, not the Year 11 one.
|
||
|
||
## Open questions
|
||
|
||
1. Is the disadvantage split its own section or a sub-block inside the
|
||
destinations section? Modelled as a sub-block; it is the most differentiating
|
||
figure on the page and the most easily misread on a small cohort.
|
||
2. Do we ingest the apprenticeship level breakdown (intermediate / advanced /
|
||
higher) now, or collapse to one apprenticeship figure and revisit? Collapsed
|
||
in this design.
|