diff --git a/docs/superpowers/plans/2026-08-28-destination-measures.md b/docs/superpowers/plans/2026-08-28-destination-measures.md new file mode 100644 index 0000000..b57dabb --- /dev/null +++ b/docs/superpowers/plans/2026-08-28-destination-measures.md @@ -0,0 +1,1754 @@ +# Destination Measures Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Show what happened to a school's leavers after Year 11 and after the sixth form, on secondary school detail pages, without ever republishing a figure DfE withheld. + +**Architecture:** A new Meltano tap pulls the EES destinations query API into `raw`; dbt staging preserves the `c` suppression sentinel as a status column rather than nulling it; long-format marts carry one row per school × year × pupil group × destination category; the API serialises a `destinations` block; a server component renders all-pupils into the HTML with one client component for the cohort switch. All three disclosure rules live as executable guards in `lib/destinations.ts`. + +**Tech Stack:** Python 3.12 / Singer SDK / Meltano · dbt + PostgreSQL · FastAPI + SQLAlchemy · Next.js (App Router) + TypeScript + CSS Modules · Jest · Playwright + +**Spec:** `docs/superpowers/specs/2026-08-28-destination-measures-design.md` + +## Global Constraints + +- **R1 — Never render a derived remainder.** Not as a number, not as a bar segment. Where any category in a pupil group is suppressed, no bar is drawn for that group. +- **R2 — Never aggregate across a suppression boundary.** Compute an aggregate from components only when every component is published. Render a DfE-published aggregate only when the count of suppressed components within it is 0 or ≥ 2. +- **R3 — Where a category is suppressed for the disadvantaged group, it is also withheld for the other-pupils group.** The all-pupils view keeps it. Enforced in the mart. +- **`safe_numeric` must never be applied to a destination count or percentage.** It coerces `c` to `NULL`, destroying the distinction between *withheld* and *no data*. +- Destination categories, verbatim from the EES filter: `School sixth form`, `Sixth form college`, `Further education`, `Other education destination`, `Sustained apprenticeships`, `Sustained employment destination`, `Not recorded as a sustained destination`, `Activity not captured`. Aggregates: `Sustained education destination`, `Sustained education, employment & apprenticeships`. +- Card grouping (ours, not DfE's): academic = school sixth form + sixth-form college; college = further education + other education; work = apprenticeship + employment. +- Copy must never imply a pupil "stayed on here" — the file reports destination *type*, never destination *institution*. +- Every new colour is a token in `nextjs-app/app/globals.css`, defined in `:root` and in both dark blocks. Never style a component from inside a theme block. +- Percentages for display are rounded; bar widths derive from unrounded pupil counts. +- EES API: `https://api.education.gov.uk/statistics/v1`. KS4 dataset `019d4f41-22d1-71b2-a1a7-f3b91026815b`; 16-18 dataset `019d4e73-6440-7523-b60c-bfab1ad4a30d`. Time periods use the `2022/2023` form, not `2022/23`. + +**Pipeline reality:** `dbt` and `meltano` do not run locally. Tasks 3–5 are verified by unit tests and by SQL review; the models only produce data once Tudor triggers the Airflow DAG on staging. Do not claim mart data exists until that has run. + +--- + +### Task 1: Destination domain logic + +The disclosure rules are here, in pure functions, so they can be tested without a database, a network, or a browser. Every later task consumes this module. + +**Files:** +- Create: `nextjs-app/lib/destinations.ts` +- Test: `nextjs-app/__tests__/lib/destinations.test.ts` + +**Interfaces:** +- Consumes: nothing +- Produces: + - `type DestinationCategory` — the eight category slugs + - `type PupilGroup = 'all' | 'disadvantaged' | 'other'` + - `type DestinationStatus = 'published' | 'suppressed' | 'not_applicable'` + - `interface DestinationCell { category; pupils: number | null; percentage: number | null; status }` + - `interface DestinationGroup { cohort: number; cells: DestinationCell[]; aggregates: Record }` + - `CARD_GROUPS: Record` + - `canAggregate(cells: DestinationCell[]): boolean` + - `aggregateCells(cells: DestinationCell[], cohort: number): { pupils: number; percentage: number } | null` + - `canRenderPublishedAggregate(components: DestinationCell[]): boolean` + - `canRenderBar(group: DestinationGroup): boolean` + - `toBarSegments(group: DestinationGroup): { category; pupils; widthPct; labelPct }[]` + - `suppressedCount(cells: DestinationCell[]): number` + +- [ ] **Step 1: Write the failing test** + +Create `nextjs-app/__tests__/lib/destinations.test.ts`: + +```ts +import { + canAggregate, aggregateCells, canRenderPublishedAggregate, + canRenderBar, toBarSegments, CARD_GROUPS, + type DestinationCell, type DestinationGroup, +} from '@/lib/destinations'; + +const pub = (category: any, pupils: number, cohort: number): DestinationCell => ({ + category, pupils, percentage: (pupils / cohort) * 100, status: 'published', +}); +const sup = (category: any): DestinationCell => ({ + category, pupils: null, percentage: null, status: 'suppressed', +}); + +const fullGroup = (): DestinationGroup => ({ + cohort: 180, + cells: [ + pub('school_sixth_form', 75, 180), pub('sixth_form_college', 21, 180), + pub('further_education', 55, 180), pub('other_education', 6, 180), + pub('apprenticeship', 8, 180), pub('employment', 6, 180), + pub('not_sustained', 5, 180), pub('not_captured', 4, 180), + ], + aggregates: {}, +}); + +describe('canAggregate — R2, computing from components', () => { + it('allows a sum when every component is published', () => { + expect(canAggregate([pub('apprenticeship', 8, 180), pub('employment', 6, 180)])).toBe(true); + }); + + it('refuses a sum when any component is suppressed', () => { + expect(canAggregate([pub('apprenticeship', 8, 180), sup('employment')])).toBe(false); + }); + + it('refuses a sum when every component is suppressed', () => { + expect(canAggregate([sup('apprenticeship'), sup('employment')])).toBe(false); + }); +}); + +describe('aggregateCells', () => { + it('sums published cells and derives a percentage from the cohort', () => { + expect(aggregateCells([pub('apprenticeship', 8, 180), pub('employment', 6, 180)], 180)) + .toEqual({ pupils: 14, percentage: (14 / 180) * 100 }); + }); + + it('returns null rather than a partial sum when a component is suppressed', () => { + expect(aggregateCells([pub('apprenticeship', 8, 180), sup('employment')], 180)).toBeNull(); + }); +}); + +describe('canRenderPublishedAggregate — R2, a total DfE published itself', () => { + it('allows it when no component is suppressed', () => { + expect(canRenderPublishedAggregate([pub('school_sixth_form', 75, 180), pub('sixth_form_college', 21, 180)])).toBe(true); + }); + + it('REFUSES it when exactly one component is suppressed — the aggregate identifies it', () => { + expect(canRenderPublishedAggregate([pub('school_sixth_form', 75, 180), sup('sixth_form_college')])).toBe(false); + }); + + it('allows it when two or more components are suppressed', () => { + expect(canRenderPublishedAggregate([sup('school_sixth_form'), sup('sixth_form_college')])).toBe(true); + }); +}); + +describe('canRenderBar — R1', () => { + it('allows a bar when the whole group is published', () => { + expect(canRenderBar(fullGroup())).toBe(true); + }); + + it('refuses a bar when a single category is suppressed', () => { + const g = fullGroup(); + g.cells[1] = sup('sixth_form_college'); + expect(canRenderBar(g)).toBe(false); + }); +}); + +describe('toBarSegments', () => { + it('derives widths from counts, not from rounded percentages', () => { + const segs = toBarSegments(fullGroup()); + expect(segs).toHaveLength(8); + expect(segs[0].widthPct).toBeCloseTo((75 / 180) * 100, 10); + expect(segs.reduce((a, s) => a + s.widthPct, 0)).toBeCloseTo(100, 6); + }); + + it('throws rather than silently leaving a gap when the group is suppressed', () => { + const g = fullGroup(); + g.cells[1] = sup('sixth_form_college'); + expect(() => toBarSegments(g)).toThrow(/suppressed/i); + }); +}); + +describe('CARD_GROUPS', () => { + it('partitions every destination category exactly once, plus the absence', () => { + const grouped = Object.values(CARD_GROUPS).flat(); + expect(new Set(grouped).size).toBe(grouped.length); + expect(grouped).toEqual(expect.arrayContaining([ + 'school_sixth_form', 'sixth_form_college', 'further_education', + 'other_education', 'apprenticeship', 'employment', + ])); + expect(grouped).not.toContain('not_sustained'); + expect(grouped).not.toContain('not_captured'); + }); +}); +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `cd nextjs-app && npx jest __tests__/lib/destinations.test.ts` +Expected: FAIL — `Cannot find module '@/lib/destinations'` + +- [ ] **Step 3: Write the implementation** + +Create `nextjs-app/lib/destinations.ts`: + +```ts +/** + * Destination measures — categories, the card grouping, and the disclosure + * guards. + * + * DfE suppresses individual cells with `c`, and the categories sum to the + * cohort. So subtracting the published cells from the cohort total recovers a + * lone suppressed cell exactly — on 22% of mainstream secondaries. The guards + * below are what stop this module's consumers doing that by accident, and they + * are the reason percentages are never reconstructed from a partial sum. + * + * See docs/superpowers/specs/2026-08-28-destination-measures-design.md. + */ + +export type DestinationCategory = + | 'school_sixth_form' + | 'sixth_form_college' + | 'further_education' + | 'other_education' + | 'apprenticeship' + | 'employment' + | 'not_sustained' + | 'not_captured'; + +export type PupilGroup = 'all' | 'disadvantaged' | 'other'; + +export type DestinationStatus = 'published' | 'suppressed' | 'not_applicable'; + +export type CardGroup = 'academic' | 'college' | 'work'; + +export interface DestinationCell { + category: DestinationCategory; + pupils: number | null; + percentage: number | null; + status: DestinationStatus; +} + +export interface DestinationGroup { + cohort: number; + cells: DestinationCell[]; + /** Aggregates DfE published itself, keyed by slug. */ + aggregates: Partial>; +} + +/** Display order, which is also bar order: education, then work, then absence. */ +export const CATEGORY_ORDER: DestinationCategory[] = [ + 'school_sixth_form', 'sixth_form_college', 'further_education', 'other_education', + 'apprenticeship', 'employment', 'not_sustained', 'not_captured', +]; + +/** + * Our grouping, not DfE's — the single most arguable thing on the page, which + * is why it lives in exactly one place. `not_sustained` and `not_captured` are + * deliberately absent: they are an absence of destination, not a route. + */ +export const CARD_GROUPS: Record = { + academic: ['school_sixth_form', 'sixth_form_college'], + college: ['further_education', 'other_education'], + work: ['apprenticeship', 'employment'], +}; + +export function suppressedCount(cells: DestinationCell[]): number { + return cells.filter(c => c.status === 'suppressed').length; +} + +/** R2: a sum computed from components is safe only if every component is published. */ +export function canAggregate(cells: DestinationCell[]): boolean { + return cells.length > 0 && cells.every(c => c.status === 'published'); +} + +export function aggregateCells( + cells: DestinationCell[], cohort: number, +): { pupils: number; percentage: number } | null { + if (!canAggregate(cells) || cohort <= 0) return null; + const pupils = cells.reduce((sum, c) => sum + (c.pupils ?? 0), 0); + return { pupils, percentage: (pupils / cohort) * 100 }; +} + +/** + * R2, the other direction: DfE published this total itself. Showing it beside + * the components is safe only when it spans no suppressed component, or two or + * more. Exactly one and the total names the withheld figure. + */ +export function canRenderPublishedAggregate(components: DestinationCell[]): boolean { + return suppressedCount(components) !== 1; +} + +/** R1: a bar is drawable only when nothing in the group is withheld. */ +export function canRenderBar(group: DestinationGroup): boolean { + return group.cohort > 0 && group.cells.every(c => c.status === 'published'); +} + +export interface BarSegment { + category: DestinationCategory; + pupils: number; + /** Exact width from the count — never the rounded percentage. */ + widthPct: number; + /** Rounded value for the segment label. */ + labelPct: number; +} + +export function toBarSegments(group: DestinationGroup): BarSegment[] { + if (!canRenderBar(group)) { + throw new Error( + 'toBarSegments: refusing to draw a bar for a group with suppressed categories — ' + + 'the gap would disclose the withheld figure (R1).', + ); + } + const byCategory = new Map(group.cells.map(c => [c.category, c])); + return CATEGORY_ORDER.flatMap(category => { + const cell = byCategory.get(category); + if (!cell || cell.pupils === null) return []; + const widthPct = (cell.pupils / group.cohort) * 100; + return [{ category, pupils: cell.pupils, widthPct, labelPct: Math.round(widthPct) }]; + }); +} + +export const CATEGORY_LABELS: Record = { + school_sixth_form: 'State-funded school sixth form', + sixth_form_college: 'Sixth-form college', + further_education: 'FE and other colleges', + other_education: 'Other education destination', + apprenticeship: 'Apprenticeship', + employment: 'Employment', + not_sustained: 'Not recorded as a sustained destination', + not_captured: 'Activity not captured', +}; + +export const CARD_QUESTIONS: Record = { + academic: { question: 'Do leavers stay on an academic route?', hint: 'a school sixth form or a sixth-form college' }, + college: { question: 'Or move to a college?', hint: 'an FE or other college' }, + work: { question: 'Or straight into work?', hint: 'an apprenticeship or a job' }, +}; +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `cd nextjs-app && npx jest __tests__/lib/destinations.test.ts` +Expected: PASS, 12 tests + +- [ ] **Step 5: Typecheck and commit** + +```bash +cd nextjs-app && npm run typecheck +cd .. && git add nextjs-app/lib/destinations.ts nextjs-app/__tests__/lib/destinations.test.ts +git commit -m "feat(destinations): the disclosure rules, as executable guards" +``` + +--- + +### Task 2: Destination colour tokens + +Six tokens, both themes. The absence is neutral plus a hatch rather than a colour, which is both a factual point (it isn't a bad outcome) and the secondary encoding that rescues a failing CVD pair. + +**Files:** +- Modify: `nextjs-app/app/globals.css` (`:root`, the `prefers-color-scheme` block, and the `[data-theme="dark"]` block if one exists) +- Test: `nextjs-app/__tests__/components/darkThemeSafety.test.ts` + +**Interfaces:** +- Consumes: nothing +- Produces: CSS custom properties `--dest-sixthform`, `--dest-sfcollege`, `--dest-fecollege`, `--dest-apprentice`, `--dest-employment`, `--dest-none`, `--dest-none-hatch` + +- [ ] **Step 1: Read the existing token blocks** + +Run: `grep -n "\-\-series-1\|prefers-color-scheme" nextjs-app/app/globals.css` + +Add the new tokens immediately after the `--series-*` group in each block so the palette stays in one place. + +- [ ] **Step 2: Write the failing test** + +Append to `nextjs-app/__tests__/components/darkThemeSafety.test.ts`: + +```ts +describe('destination tokens', () => { + const css = readFileSync(join(process.cwd(), 'app/globals.css'), 'utf8'); + const tokens = [ + '--dest-sixthform', '--dest-sfcollege', '--dest-fecollege', + '--dest-apprentice', '--dest-employment', '--dest-none', '--dest-none-hatch', + ]; + + it('defines every destination token in the light palette', () => { + const root = css.slice(css.indexOf(':root {'), css.indexOf('@media (prefers-color-scheme: dark)')); + tokens.forEach(t => expect(root).toContain(t + ':')); + }); + + it('redefines every destination token for dark', () => { + const dark = css.slice(css.indexOf('@media (prefers-color-scheme: dark)')); + tokens.forEach(t => expect(dark).toContain(t + ':')); + }); +}); +``` + +- [ ] **Step 3: Run test to verify it fails** + +Run: `cd nextjs-app && npx jest __tests__/components/darkThemeSafety.test.ts` +Expected: FAIL — light palette missing `--dest-sixthform:` + +- [ ] **Step 4: Add the tokens** + +In the `:root` block: + +```css + /* ── Destination measures ─────────────────────────────────────────── + Education is one hue in three steps (school-like → college-like) so the + three education destinations read as one family; apprenticeship and + employment are separate hues. The absence is neutral and hatched, never + a colour — "activity not captured" includes independent schools and + moving abroad, so rendering it as a bad outcome would be wrong. The + hatch is also what rescues the neutral/blue pair, which fails CVD + separation at ΔE 7.6 as flat fills. Every other adjacent pair clears + ΔE 10.9 under protanopia. */ + --dest-sixthform: #0F766E; + --dest-sfcollege: #4A9E96; + --dest-fecollege: #7CBFB8; + --dest-apprentice: #806200; + --dest-employment: #2F6F8F; + --dest-none: #6B7580; + --dest-none-hatch: rgba(107, 117, 128, 0.34); +``` + +In the `@media (prefers-color-scheme: dark)` block (and the `[data-theme="dark"]` block if present): + +```css + --dest-sixthform: #5FC7BB; + --dest-sfcollege: #3E9B92; + --dest-fecollege: #2A716B; + --dest-apprentice: #EFC658; + --dest-employment: #8FB4D9; + --dest-none: #8B9AA1; + --dest-none-hatch: rgba(139, 154, 161, 0.34); +``` + +- [ ] **Step 5: Run test to verify it passes** + +Run: `cd nextjs-app && npx jest __tests__/components/darkThemeSafety.test.ts` +Expected: PASS + +- [ ] **Step 6: Commit** + +```bash +git add nextjs-app/app/globals.css nextjs-app/__tests__/components/darkThemeSafety.test.ts +git commit -m "feat(destinations): colour tokens, with the absence hatched not coloured" +``` + +--- + +### Task 3: The destinations tap + +A separate extractor from `tap-uk-ees`. That tap downloads a release ZIP and reads a CSV inside it; the destinations files carry every breakdown we don't want, so this one POSTs to the query API and pages. + +**Files:** +- Create: `pipeline/plugins/extractors/tap-uk-ees-destinations/pyproject.toml` +- Create: `pipeline/plugins/extractors/tap-uk-ees-destinations/tap_uk_ees_destinations/__init__.py` +- Create: `pipeline/plugins/extractors/tap-uk-ees-destinations/tap_uk_ees_destinations/tap.py` +- Modify: `pipeline/meltano.yml` +- Modify: `pipeline/dags/school_data_pipeline.py:160-196` +- Test: `pipeline/plugins/extractors/tap-uk-ees-destinations/tests/test_tap.py` + +**Interfaces:** +- Consumes: nothing +- Produces: raw tables `ees_ks4_destinations` and `ees_ks5_destinations`, columns `urn`, `time_period`, `pupil_group`, `destination_measure`, `cohort_pupils`, `pupils_raw`, `percentage_raw` — the last two as **text**, sentinel preserved. + +- [ ] **Step 1: Write the failing test** + +Create `pipeline/plugins/extractors/tap-uk-ees-destinations/tests/test_tap.py`: + +```python +"""The tap's only job that can be tested without the network: mapping an API +row to a Singer record without destroying the suppression sentinel.""" +from tap_uk_ees_destinations.tap import row_to_record, DESTINATION_SLUGS, PUPIL_GROUP_SLUGS + + +def test_published_row_keeps_its_numbers_as_text(): + row = { + "timePeriod": {"period": "2022/2023"}, + "geographicLevel": "SCH", + "locations": {"SCH": "IXn5B"}, + "filters": {"wYXbx": "DCz1Q", "9ss4v": "p9WRS"}, + "values": {"Poghe": "264", "1roqi": "182", "dPjk0": "68.9"}, + } + rec = row_to_record(row, urn_by_location={"IXn5B": "137083"}) + assert rec["urn"] == "137083" + assert rec["time_period"] == "202223" + assert rec["destination_measure"] == "school_sixth_form" + assert rec["pupil_group"] == "all" + assert rec["cohort_pupils"] == "264" + assert rec["pupils_raw"] == "182" + assert rec["percentage_raw"] == "68.9" + + +def test_suppressed_row_preserves_the_c_sentinel(): + row = { + "timePeriod": {"period": "2022/2023"}, + "geographicLevel": "SCH", + "locations": {"SCH": "IXn5B"}, + "filters": {"wYXbx": "eLsdu", "9ss4v": "OvPnC"}, + "values": {"Poghe": "41", "1roqi": "c", "dPjk0": "c"}, + } + rec = row_to_record(row, urn_by_location={"IXn5B": "137083"}) + assert rec["pupils_raw"] == "c", "the sentinel must survive extraction" + assert rec["percentage_raw"] == "c" + assert rec["pupil_group"] == "disadvantaged" + + +def test_national_rows_are_kept_with_a_null_urn(): + """The England reference lives in the same response. It is kept, with urn + None, so fact_destination_national has something to read.""" + row = { + "timePeriod": {"period": "2022/2023"}, + "geographicLevel": "NAT", + "locations": {"NAT": "dP0Zw"}, + "filters": {"wYXbx": "DCz1Q", "9ss4v": "p9WRS"}, + "values": {"Poghe": "500000", "1roqi": "190000", "dPjk0": "38.0"}, + } + rec = row_to_record(row, urn_by_location={}) + assert rec is not None + assert rec["urn"] is None + assert rec["percentage_raw"] == "38.0" + + +def test_other_geographic_levels_are_dropped(): + """Local authority, district, region and constituency rows are noise here.""" + row = { + "timePeriod": {"period": "2022/2023"}, + "geographicLevel": "LA", + "locations": {"LA": "u9Oo4", "NAT": "dP0Zw"}, + "filters": {"wYXbx": "DCz1Q", "9ss4v": "p9WRS"}, + "values": {"Poghe": "1", "1roqi": "1", "dPjk0": "1"}, + } + assert row_to_record(row, urn_by_location={}) is None + + +def test_every_slug_maps_to_one_filter_id(): + assert len(set(DESTINATION_SLUGS.values())) == len(DESTINATION_SLUGS) + assert set(PUPIL_GROUP_SLUGS.values()) == {"all", "disadvantaged", "other"} +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `cd pipeline/plugins/extractors/tap-uk-ees-destinations && uv run --with singer-sdk --with requests pytest tests/ -v` +Expected: FAIL — `ModuleNotFoundError: tap_uk_ees_destinations` + +- [ ] **Step 3: Write the tap** + +Create `pipeline/plugins/extractors/tap-uk-ees-destinations/pyproject.toml`: + +```toml +[project] +name = "tap-uk-ees-destinations" +version = "0.1.0" +requires-python = ">=3.10" +dependencies = ["singer-sdk>=0.40", "requests>=2.31"] + +[project.scripts] +tap-uk-ees-destinations = "tap_uk_ees_destinations.tap:TapUKEESDestinations.cli" + +[build-system] +requires = ["hatchling"] +build-backend = "hatchling.build" +``` + +Create `pipeline/plugins/extractors/tap-uk-ees-destinations/tap_uk_ees_destinations/__init__.py` (empty file). + +Create `pipeline/plugins/extractors/tap-uk-ees-destinations/tap_uk_ees_destinations/tap.py`: + +```python +"""EES destinations tap — KS4 and 16-18 destination measures, school level. + +Separate from tap-uk-ees on purpose. That tap pulls a release ZIP and reads a +CSV inside it; the destinations files carry sex, ethnicity, FSM status, prior +attainment and SEN in the same table, so the whole-file route would download +millions of rows to keep a few hundred thousand. The query API filters server +side. + +The one thing this tap must not do is tidy up the data. EES writes `c` where a +figure is withheld, and the categories sum to the cohort — so turning `c` into +NULL here would let a downstream sum reconstruct exactly what DfE suppressed. +Counts and percentages are emitted as TEXT, sentinel intact. +""" + +from __future__ import annotations + +import requests +from singer_sdk import Stream, Tap +from singer_sdk import typing as th + +API_BASE = "https://api.education.gov.uk/statistics/v1" +TIMEOUT = 180 +PAGE_SIZE = 10000 + +KS4_DATASET = "019d4f41-22d1-71b2-a1a7-f3b91026815b" +KS5_DATASET = "019d4e73-6440-7523-b60c-bfab1ad4a30d" + +# Filter option ids, read from each dataset's /meta. KS4 and 16-18 use +# different ids for the same concepts, so they are declared separately. +KS4_DESTINATION_SLUGS = { + "DCz1Q": "school_sixth_form", + "eLsdu": "sixth_form_college", + "o2MJm": "further_education", + "b7v6t": "other_education", + "mlKo9": "apprenticeship", + "QIJEw": "employment", + "RZrek": "not_sustained", + "1j7Ui": "not_captured", + "WiEl2": "agg_sustained_education", + "EfSAq": "agg_sustained_all", +} +KS4_PUPIL_GROUP_SLUGS = {"p9WRS": "all", "OvPnC": "disadvantaged", "7VmdX": "other"} +# Sex=Total and characteristic topic=Total. Leaving these unpinned returns +# every breakdown crossed with every other — 45 rows where 9 are wanted. +KS4_PINNED = ["X542f", "jHdaA"] + +KS5_DESTINATION_SLUGS = { + "dkyu0": "higher_education", + "QIB6w": "school_sixth_form", + "9c3u4": "sixth_form_college", + "S4UgV": "further_education", + "EfYDq": "other_education", + "PaLTe": "apprenticeship", + "9c8k4": "employment", + "Wi6R2": "not_sustained", + "o2l7m": "not_captured", + "wBXtb": "agg_sustained_education", + "o2c4m": "agg_sustained_all", +} +KS5_PUPIL_GROUP_SLUGS = {"Y0PuH": "all", "HTeez": "disadvantaged", "CnSVI": "other"} +KS5_PINNED: list[str] = [] + +# Indicator ids are shared across both datasets. +IND_COHORT = "Poghe" +IND_PUPILS = "1roqi" +IND_PERCENT = "dPjk0" + +DESTINATION_SLUGS = KS4_DESTINATION_SLUGS +PUPIL_GROUP_SLUGS = KS4_PUPIL_GROUP_SLUGS + + +def _period_to_time_period(period: str) -> str: + """'2022/2023' -> '202223', matching the convention the other marts use.""" + start, end = period.split("/") + return start + end[-2:] + + +def row_to_record( + row: dict, + urn_by_location: dict[str, str], + destination_slugs: dict[str, str] | None = None, + pupil_group_slugs: dict[str, str] | None = None, +) -> dict | None: + """Map one API row to a Singer record, or None if it is not school level.""" + destination_slugs = destination_slugs or KS4_DESTINATION_SLUGS + pupil_group_slugs = pupil_group_slugs or KS4_PUPIL_GROUP_SLUGS + + # Read geographicLevel, not the locations keys: a school row also carries + # NAT, LA and REG entries for its parents, so "NAT in locations" is true + # for every row in the file and would let LA rows through as national ones. + level = row.get("geographicLevel") + if level == "SCH": + urn = urn_by_location.get(row.get("locations", {}).get("SCH")) + if not urn: + return None + elif level == "NAT": + urn = None + else: + return None + + filters = row.get("filters", {}) + destination = next( + (slug for fid, slug in destination_slugs.items() if fid in filters.values()), None + ) + group = next( + (slug for fid, slug in pupil_group_slugs.items() if fid in filters.values()), None + ) + if destination is None or group is None: + return None + + values = row.get("values", {}) + return { + "urn": urn, + "time_period": _period_to_time_period(row["timePeriod"]["period"]), + "pupil_group": group, + "destination_measure": destination, + "cohort_pupils": values.get(IND_COHORT), + "pupils_raw": values.get(IND_PUPILS), + "percentage_raw": values.get(IND_PERCENT), + } + + +def fetch_urn_by_location(dataset_id: str) -> dict[str, str]: + """Location id -> URN, from the dataset's meta.""" + resp = requests.get(f"{API_BASE}/data-sets/{dataset_id}/meta", timeout=TIMEOUT) + resp.raise_for_status() + for group in resp.json().get("locations", []): + if group.get("level", {}).get("code") == "SCH": + return {o["id"]: o["urn"] for o in group.get("options", []) if o.get("urn")} + return {} + + +def fetch_time_periods(dataset_id: str) -> list[str]: + resp = requests.get(f"{API_BASE}/data-sets/{dataset_id}/meta", timeout=TIMEOUT) + resp.raise_for_status() + return [t["period"] for t in resp.json().get("timePeriods", [])] + + +class DestinationsStream(Stream): + """One stream per dataset. Pages the query API, one time period at a time.""" + + _dataset_id: str + _destination_slugs: dict[str, str] + _pupil_group_slugs: dict[str, str] + _pinned: list[str] + + schema = th.PropertiesList( + th.Property("urn", th.StringType), + th.Property("time_period", th.StringType), + th.Property("pupil_group", th.StringType), + th.Property("destination_measure", th.StringType), + th.Property("cohort_pupils", th.StringType), + th.Property("pupils_raw", th.StringType), + th.Property("percentage_raw", th.StringType), + ).to_dict() + + primary_keys = ["urn", "time_period", "pupil_group", "destination_measure"] + replication_key = None + + def get_records(self, context): + urn_by_location = fetch_urn_by_location(self._dataset_id) + self.logger.info("%s: %d school locations", self.name, len(urn_by_location)) + + criteria = [{"filters": {"in": list(self._destination_slugs)}}, + {"filters": {"in": list(self._pupil_group_slugs)}}] + for pinned in self._pinned: + criteria.append({"filters": {"in": [pinned]}}) + + for period in fetch_time_periods(self._dataset_id): + page = 1 + while True: + body = { + "criteria": {"and": criteria + [ + {"timePeriods": {"in": [{"period": period, "code": "AY"}]}}, + ]}, + "indicators": [IND_COHORT, IND_PUPILS, IND_PERCENT], + "page": page, + "pageSize": PAGE_SIZE, + } + resp = requests.post( + f"{API_BASE}/data-sets/{self._dataset_id}/query", + json=body, timeout=TIMEOUT, + ) + resp.raise_for_status() + payload = resp.json() + + for row in payload.get("results", []): + record = row_to_record( + row, urn_by_location, + self._destination_slugs, self._pupil_group_slugs, + ) + if record is not None: + yield record + + paging = payload.get("paging", {}) + if page >= paging.get("totalPages", 1): + break + page += 1 + + +class KS4DestinationsStream(DestinationsStream): + name = "ees_ks4_destinations" + _dataset_id = KS4_DATASET + _destination_slugs = KS4_DESTINATION_SLUGS + _pupil_group_slugs = KS4_PUPIL_GROUP_SLUGS + _pinned = KS4_PINNED + + +class KS5DestinationsStream(DestinationsStream): + name = "ees_ks5_destinations" + _dataset_id = KS5_DATASET + _destination_slugs = KS5_DESTINATION_SLUGS + _pupil_group_slugs = KS5_PUPIL_GROUP_SLUGS + _pinned = KS5_PINNED + + +class TapUKEESDestinations(Tap): + name = "tap-uk-ees-destinations" + config_jsonschema = th.PropertiesList().to_dict() + + def discover_streams(self): + return [KS4DestinationsStream(self), KS5DestinationsStream(self)] + + +if __name__ == "__main__": + TapUKEESDestinations.cli() +``` + +- [ ] **Step 4: Run test to verify it passes** + +Run: `cd pipeline/plugins/extractors/tap-uk-ees-destinations && uv run --with singer-sdk --with requests --with pytest pytest tests/ -v` +Expected: PASS, 4 tests + +- [ ] **Step 5: Register the tap with Meltano** + +In `pipeline/meltano.yml`, after the `tap-uk-ees` block in `extractors:`: + +```yaml + - name: tap-uk-ees-destinations + namespace: uk_ees_destinations + pip_url: ./plugins/extractors/tap-uk-ees-destinations + executable: tap-uk-ees-destinations + settings: [] +``` + +- [ ] **Step 6: Wire it into the annual DAG** + +In `pipeline/dags/school_data_pipeline.py`, inside the `extract_ees` TaskGroup (around line 174), add a second operator and make the dbt selector cover the new models: + +```python + extract_ees_destinations = BashOperator( + task_id="extract_ees_destinations", + bash_command=f"cd {PIPELINE_DIR} && {MELTANO_BIN} run tap-uk-ees-destinations target-postgres", + ) + + extract_ees >> extract_ees_destinations +``` + +And extend the `dbt_build_ees` selector with `stg_ees_ks4_destinations+ stg_ees_ks5_destinations+`. + +- [ ] **Step 7: Commit** + +```bash +git add pipeline/plugins/extractors/tap-uk-ees-destinations pipeline/meltano.yml pipeline/dags/school_data_pipeline.py +git commit -m "feat(destinations): a tap that preserves the suppression sentinel" +``` + +--- + +### Task 4: Staging models + +**Files:** +- Create: `pipeline/transform/models/staging/stg_ees_ks4_destinations.sql` +- Create: `pipeline/transform/models/staging/stg_ees_ks5_destinations.sql` +- Modify: `pipeline/transform/models/staging/_stg_sources.yml` + +**Interfaces:** +- Consumes: raw tables from Task 3 +- Produces: `stg_ees_ks4_destinations` / `stg_ees_ks5_destinations` with columns `urn` (int), `year` (int), `pupil_group` (text), `destination_measure` (text), `cohort_pupils` (int), `pupils` (int, null when withheld), `percentage` (numeric, null when withheld), `status` (text: `published` | `suppressed` | `not_applicable`) + +- [ ] **Step 1: Declare the sources** + +In `pipeline/transform/models/staging/_stg_sources.yml`, under `tables:`: + +```yaml + - name: ees_ks4_destinations + description: > + KS4 leavers destinations, school level, long format — one row per + URN × year × pupil group × destination measure. pupils_raw and + percentage_raw are TEXT and may hold the 'c' suppression sentinel; + they must never be passed through safe_numeric. + + - name: ees_ks5_destinations + description: > + 16-18 study leavers destinations, same grain and same suppression + caveat as ees_ks4_destinations. +``` + +- [ ] **Step 2: Write the KS4 staging model** + +Create `pipeline/transform/models/staging/stg_ees_ks4_destinations.sql`: + +```sql +{{ config(materialized='table') }} + +-- Staging model: KS4 leavers destinations, school level. +-- +-- DELIBERATELY DOES NOT USE safe_numeric. That macro maps every EES sentinel +-- (z, c, x, q, u) to NULL, which is right for attainment — there, "suppressed" +-- and "not applicable" are equally unrenderable. Here they are different +-- claims: one prints "withheld", the other prints nothing. Collapsing them +-- would also let a downstream sum reconstruct a withheld figure, because the +-- destination categories add up to the cohort. +-- +-- See docs/superpowers/specs/2026-08-28-destination-measures-design.md. + +with source as ( + select * from {{ source('raw', 'ees_ks4_destinations') }} + -- National rows carry a null urn and feed fact_destination_national. + where (urn is null or urn ~ '^[0-9]+$') + and time_period ~ '^[0-9]+$' +) + +select + case when urn ~ '^[0-9]+$' then cast(trim(urn) as integer) end as urn, + cast(trim(time_period) as integer) as year, + trim(pupil_group) as pupil_group, + trim(destination_measure) as destination_measure, + + case when cohort_pupils ~ '^[0-9]+$' + then cast(cohort_pupils as integer) end as cohort_pupils, + + case when pupils_raw ~ '^[0-9]+$' + then cast(pupils_raw as integer) end as pupils, + + case when percentage_raw ~ '^-?[0-9]+(\.[0-9]+)?$' + then cast(percentage_raw as numeric) end as percentage, + + case + when pupils_raw ~ '^[0-9]+$' then 'published' + when lower(trim(pupils_raw)) = 'c' then 'suppressed' + else 'not_applicable' + end as status + +from source +``` + +- [ ] **Step 3: Write the 16-18 staging model** + +Create `pipeline/transform/models/staging/stg_ees_ks5_destinations.sql` — identical body, reading `source('raw', 'ees_ks5_destinations')`, with the header comment naming 16-18 study leavers. Repeat the full SQL rather than abstracting it; the two sources drift independently and a shared macro would couple their refresh cadences. + +- [ ] **Step 4: Verify the SQL compiles by eye against the sibling models** + +Run: `diff <(sed -n '1,12p' pipeline/transform/models/staging/stg_ees_ks4.sql) <(sed -n '1,12p' pipeline/transform/models/staging/stg_ees_ks4_destinations.sql)` + +Confirm the header comment style matches, and confirm by inspection that `safe_numeric` appears nowhere: + +Run: `grep -c safe_numeric pipeline/transform/models/staging/stg_ees_ks*_destinations.sql` +Expected: `0` for both files + +- [ ] **Step 5: Commit** + +```bash +git add pipeline/transform/models/staging/stg_ees_ks4_destinations.sql \ + pipeline/transform/models/staging/stg_ees_ks5_destinations.sql \ + pipeline/transform/models/staging/_stg_sources.yml +git commit -m "feat(destinations): staging models that keep 'withheld' distinct from 'absent'" +``` + +--- + +### Task 5: Marts and the disclosure tests + +**Files:** +- Create: `pipeline/transform/models/marts/fact_ks4_destinations.sql` +- Create: `pipeline/transform/models/marts/fact_ks5_destinations.sql` +- Create: `pipeline/transform/models/marts/fact_destination_national.sql` +- Create: `pipeline/transform/tests/assert_destinations_no_derived_remainder.sql` +- Create: `pipeline/transform/tests/assert_destinations_group_masking.sql` +- Create: `pipeline/transform/tests/assert_destination_status_null_agreement.sql` +- Modify: `pipeline/transform/models/marts/_marts_schema.yml` + +**Interfaces:** +- Consumes: `stg_ees_ks4_destinations`, `stg_ees_ks5_destinations`, `dim_school` +- Produces: `fact_ks4_destinations` / `fact_ks5_destinations` (`urn`, `year`, `pupil_group`, `destination_measure`, `cohort_pupils`, `pupils`, `percentage`, `status`) and `fact_destination_national` (same, without `urn`, plus `phase`) + +- [ ] **Step 1: Write the KS4 mart with R3 masking** + +Create `pipeline/transform/models/marts/fact_ks4_destinations.sql`: + +```sql +{{ config(materialized='table') }} + +-- Mart: KS4 leavers destinations — one row per URN × year × pupil group × +-- destination measure. +-- +-- Long format, unlike the wide fact_ks4_performance next door. pupil_group is +-- a real third dimension, so going wide would need three sets of every column, +-- and the disclosure tests below are far easier to write over rows. +-- +-- R3 is applied HERE rather than in the API: where a category is suppressed +-- for the disadvantaged group it is masked for the other-pupils group too, +-- because the two partition the whole and the all-pupils figure is published. +-- DfE already does this in 493 of 498 cases; this closes the remainder so no +-- consumer can reach an unmasked combination. + +with staged as ( + select s.* + from {{ ref('stg_ees_ks4_destinations') }} s + inner join {{ ref('dim_school') }} d on d.urn = s.urn +), + +-- Categories withheld for disadvantaged pupils at this school and year. +masked as ( + select distinct urn, year, destination_measure + from staged + where pupil_group = 'disadvantaged' and status = 'suppressed' +) + +select + s.urn, + s.year, + s.pupil_group, + s.destination_measure, + s.cohort_pupils, + case when m.urn is not null and s.pupil_group = 'other' + then null else s.pupils end as pupils, + case when m.urn is not null and s.pupil_group = 'other' + then null else s.percentage end as percentage, + case when m.urn is not null and s.pupil_group = 'other' + then 'suppressed' else s.status end as status +from staged s +left join masked m + on m.urn = s.urn + and m.year = s.year + and m.destination_measure = s.destination_measure +``` + +- [ ] **Step 2: Write the 16-18 mart** + +Create `pipeline/transform/models/marts/fact_ks5_destinations.sql` — the same body reading `stg_ees_ks5_destinations`, with a header naming 16-18 study leavers. + +- [ ] **Step 3: Write the national reference mart** + +Create `pipeline/transform/models/marts/fact_destination_national.sql`: + +```sql +{{ config(materialized='table') }} + +-- Mart: England destination measures by pupil group, for the page's national +-- reference. Kept separate from the school facts so the section's England bar +-- can repoint with the cohort switch — comparing a school's disadvantaged +-- pupils against the national all-pupils figure would flatter or damn the +-- school for its intake rather than its work. + +select 'ks4' as phase, year, pupil_group, destination_measure, + cohort_pupils, pupils, percentage, status +from {{ ref('stg_ees_ks4_destinations') }} +where urn is null + +union all + +select 'ks5' as phase, year, pupil_group, destination_measure, + cohort_pupils, pupils, percentage, status +from {{ ref('stg_ees_ks5_destinations') }} +where urn is null +``` + +- [ ] **Step 4: Write the R1 disclosure test** + +Create `pipeline/transform/tests/assert_destinations_no_derived_remainder.sql`: + +```sql +-- R1 GUARD. Fails if a school/year/group has exactly one suppressed category +-- while also publishing the cohort total — the combination that lets the +-- withheld figure be recovered by subtraction. +-- +-- This does not mean the mart is wrong: DfE publishes exactly this, and the +-- mart's job is to carry it faithfully. The test exists so that the condition +-- is visible and counted, and so that any consumer added later has to +-- acknowledge it. The API and the frontend are what must refuse to render the +-- remainder; this test is the tripwire that says how often the situation +-- arises. It is configured to warn, not error. +{{ config(severity='warn') }} + +select + urn, year, pupil_group, + count(*) filter (where status = 'suppressed') as suppressed_categories +from {{ ref('fact_ks4_destinations') }} +where destination_measure not like 'agg_%' +group by urn, year, pupil_group +having count(*) filter (where status = 'suppressed') = 1 +``` + +- [ ] **Step 5: Write the R3 masking test** + +Create `pipeline/transform/tests/assert_destinations_group_masking.sql`: + +```sql +-- R3 GUARD. Fails if a category is suppressed for disadvantaged pupils but +-- still published for the other-pupils group — the two partition the whole, so +-- publishing both alongside the all-pupils figure recovers the withheld cell. + +select d.urn, d.year, d.destination_measure +from {{ ref('fact_ks4_destinations') }} d +inner join {{ ref('fact_ks4_destinations') }} o + on o.urn = d.urn + and o.year = d.year + and o.destination_measure = d.destination_measure + and o.pupil_group = 'other' +where d.pupil_group = 'disadvantaged' + and d.status = 'suppressed' + and o.status = 'published' +``` + +- [ ] **Step 6: Write the status agreement test** + +Create `pipeline/transform/tests/assert_destination_status_null_agreement.sql`: + +```sql +-- pupils must be null wherever status is not 'published', and never null where +-- it is. This is what stops a later coalesce or a wide-format refactor turning +-- "withheld" into a zero. + +select urn, year, pupil_group, destination_measure, status, pupils +from {{ ref('fact_ks4_destinations') }} +where (status <> 'published' and pupils is not null) + or (status = 'published' and pupils is null) +``` + +- [ ] **Step 7: Document the marts** + +In `pipeline/transform/models/marts/_marts_schema.yml`, add entries for the three new models with a `description` for each and `tests: [not_null]` on `urn`, `year`, `pupil_group`, `destination_measure`, `status`, following the existing entries' shape. + +- [ ] **Step 8: Confirm no test references safe_numeric and commit** + +Run: `grep -rn safe_numeric pipeline/transform/models/marts/fact_ks*_destinations.sql pipeline/transform/models/marts/fact_destination_national.sql` +Expected: no output + +```bash +git add pipeline/transform/models/marts/fact_ks4_destinations.sql \ + pipeline/transform/models/marts/fact_ks5_destinations.sql \ + pipeline/transform/models/marts/fact_destination_national.sql \ + pipeline/transform/tests/assert_destinations_*.sql \ + pipeline/transform/tests/assert_destination_status_null_agreement.sql \ + pipeline/transform/models/marts/_marts_schema.yml +git commit -m "feat(destinations): marts, with R3 masking applied at the boundary" +``` + +--- + +### Task 6: Backend models, loader and API + +**Files:** +- Modify: `backend/models.py` (append after `FactFinance`, around line 265) +- Modify: `backend/data_loader.py:819-830` (`_empty_supplementary`) and `:833-958` (`get_supplementary_data_batch`) +- Modify: `backend/app.py:905-920` (the `get_school_details` return block) +- Test: `backend/tests/test_destinations_api.py` + +**Interfaces:** +- Consumes: `fact_ks4_destinations`, `fact_ks5_destinations`, `fact_destination_national` +- Produces: `destinations` key on `GET /api/schools/{urn}`, shaped `{ ks4: {...} | null, ks5: {...} | null }`; each phase `{ cohort_year, groups: { all, disadvantaged, other } }`; each group `{ cohort, categories: [{ category, pupils, percentage, status }], aggregates: {...} }` + +- [ ] **Step 1: Write the failing test** + +Create `backend/tests/test_destinations_api.py`: + +```python +"""The serialiser's contract: it carries suppression through, and never emits a +total that closes a gap left by a suppressed category.""" +import pytest + +from backend.data_loader import _destinations_block + + +def _row(group, measure, pupils, status, cohort=180, percentage=None): + return { + "pupil_group": group, "destination_measure": measure, + "pupils": pupils, "percentage": percentage, + "status": status, "cohort_pupils": cohort, "year": 202223, + } + + +def test_suppressed_category_serialises_as_suppressed_with_null_pupils(): + rows = [ + _row("all", "school_sixth_form", 75, "published", percentage=41.7), + _row("all", "sixth_form_college", None, "suppressed"), + ] + block = _destinations_block(rows) + cats = {c["category"]: c for c in block["groups"]["all"]["categories"]} + assert cats["sixth_form_college"]["status"] == "suppressed" + assert cats["sixth_form_college"]["pupils"] is None + assert cats["sixth_form_college"]["percentage"] is None + + +def test_no_closing_total_is_emitted_for_a_partially_suppressed_group(): + rows = [ + _row("all", "school_sixth_form", 75, "published", percentage=41.7), + _row("all", "sixth_form_college", None, "suppressed"), + _row("all", "further_education", 61, "published", percentage=33.9), + _row("all", "apprenticeship", 8, "published", percentage=4.4), + _row("all", "employment", 6, "published", percentage=3.3), + _row("all", "not_sustained", 5, "published", percentage=2.8), + _row("all", "not_captured", 4, "published", percentage=2.2), + ] + block = _destinations_block(rows) + group = block["groups"]["all"] + published = sum(c["pupils"] for c in group["categories"] if c["pupils"] is not None) + for value in group["aggregates"].values(): + if value is None or value.get("pupils") is None: + continue + assert value["pupils"] != group["cohort"] - published, ( + "an aggregate that equals the residual identifies the suppressed cell" + ) + + +def test_cohort_year_is_reported_so_the_page_can_date_itself(): + block = _destinations_block([_row("all", "school_sixth_form", 75, "published")]) + assert block["cohort_year"] == "2022/23" + + +def test_empty_rows_yield_none_not_an_empty_shell(): + assert _destinations_block([]) is None +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `cd /Users/tudor/projects/school_compare && uv run --with fastapi --with 'httpx==0.27.0' --with sqlalchemy --with pandas --with pytest --with pydantic-settings python -m pytest backend/tests/test_destinations_api.py -v` +Expected: FAIL — `ImportError: cannot import name '_destinations_block'` + +- [ ] **Step 3: Add the SQLAlchemy models** + +Append to `backend/models.py` after `FactFinance`: + +```python +class FactKs4Destinations(Base): + """KS4 leavers destinations — one row per URN, year, pupil group, measure.""" + __tablename__ = "fact_ks4_destinations" + __table_args__ = ( + Index("ix_ks4_dest_urn_year", "urn", "year"), + MARTS, + ) + + urn = Column(Integer, primary_key=True) + year = Column(Integer, primary_key=True) + pupil_group = Column(String(20), primary_key=True) + destination_measure = Column(String(40), primary_key=True) + cohort_pupils = Column(Integer) + pupils = Column(Integer) + percentage = Column(Float) + # 'published' | 'suppressed' | 'not_applicable'. Never collapse this to a + # null check: a suppressed cell prints "withheld", an absent one prints + # nothing, and the difference is what keeps the disclosure rules workable. + status = Column(String(20)) + + +class FactKs5Destinations(Base): + """16-18 study leavers destinations — same grain as FactKs4Destinations.""" + __tablename__ = "fact_ks5_destinations" + __table_args__ = ( + Index("ix_ks5_dest_urn_year", "urn", "year"), + MARTS, + ) + + urn = Column(Integer, primary_key=True) + year = Column(Integer, primary_key=True) + pupil_group = Column(String(20), primary_key=True) + destination_measure = Column(String(40), primary_key=True) + cohort_pupils = Column(Integer) + pupils = Column(Integer) + percentage = Column(Float) + status = Column(String(20)) +``` + +- [ ] **Step 4: Write the serialiser** + +Add to `backend/data_loader.py`, above `_empty_supplementary`: + +```python +_AGGREGATE_MEASURES = {"agg_sustained_education", "agg_sustained_all"} + + +def _format_cohort_year(year: int | None) -> str | None: + """202223 -> '2022/23'. The page must date its own cohort: destinations run + two GCSE years behind the results shown above them.""" + if not year: + return None + text = str(year) + return f"{text[:4]}/{text[6:8]}" if len(text) == 8 else f"{text[:4]}/{text[4:6]}" + + +def _destinations_block(rows: list[dict]) -> dict | None: + """Shape destination rows for one phase into the API's block. + + Carries `status` through untouched and emits no computed totals. The only + aggregates present are ones DfE published itself; the frontend decides + whether they are safe to show (see lib/destinations.ts, R2). + """ + if not rows: + return None + + latest_year = max(r["year"] for r in rows if r.get("year") is not None) + rows = [r for r in rows if r.get("year") == latest_year] + + groups: dict[str, dict] = {} + for row in rows: + group = groups.setdefault( + row["pupil_group"], + {"cohort": row.get("cohort_pupils"), "categories": [], "aggregates": {}}, + ) + measure = row["destination_measure"] + cell = { + "category": measure, + "pupils": row.get("pupils"), + "percentage": row.get("percentage"), + "status": row.get("status"), + } + if measure in _AGGREGATE_MEASURES: + group["aggregates"][measure.removeprefix("agg_")] = cell + else: + group["categories"].append(cell) + + if not groups: + return None + + return {"cohort_year": _format_cohort_year(latest_year), "groups": groups} +``` + +- [ ] **Step 5: Run test to verify it passes** + +Run: `cd /Users/tudor/projects/school_compare && uv run --with fastapi --with 'httpx==0.27.0' --with sqlalchemy --with pandas --with pytest --with pydantic-settings python -m pytest backend/tests/test_destinations_api.py -v` +Expected: PASS, 4 tests + +- [ ] **Step 6: Wire it into the batch loader** + +In `backend/data_loader.py`, add `"destinations": None` to the dict `_empty_supplementary` returns. Then add two query functions inside `get_supplementary_data_batch`, following the `_ofsted` / `_census` pattern exactly, each wrapped in `_safe`: + +```python + # Destinations — KS4 and 16-18, all years; _destinations_block picks the + # latest and shapes the groups. + def _destinations(): + from collections import defaultdict + per_urn_ks4 = defaultdict(list) + for r in (db.query(FactKs4Destinations) + .filter(FactKs4Destinations.urn.in_(urns)).all()): + per_urn_ks4[r.urn].append({ + "year": r.year, "pupil_group": r.pupil_group, + "destination_measure": r.destination_measure, + "cohort_pupils": r.cohort_pupils, "pupils": r.pupils, + "percentage": r.percentage, "status": r.status, + }) + per_urn_ks5 = defaultdict(list) + for r in (db.query(FactKs5Destinations) + .filter(FactKs5Destinations.urn.in_(urns)).all()): + per_urn_ks5[r.urn].append({ + "year": r.year, "pupil_group": r.pupil_group, + "destination_measure": r.destination_measure, + "cohort_pupils": r.cohort_pupils, "pupils": r.pupils, + "percentage": r.percentage, "status": r.status, + }) + for urn in urns: + ks4 = _destinations_block(per_urn_ks4.get(urn, [])) + ks5 = _destinations_block(per_urn_ks5.get(urn, [])) + result[urn]["destinations"] = ( + {"ks4": ks4, "ks5": ks5} if (ks4 or ks5) else None + ) + _safe(_destinations) +``` + +Import `FactKs4Destinations` and `FactKs5Destinations` alongside the other mart models at the top of the file. + +- [ ] **Step 7: Expose it on the endpoint** + +In `backend/app.py`, in the `get_school_details` return dict, after `"finance": supplementary.get("finance"),`: + +```python + "destinations": supplementary.get("destinations"), +``` + +- [ ] **Step 8: Run the full backend suite and commit** + +Run: `cd /Users/tudor/projects/school_compare && uv run --with fastapi --with 'httpx==0.27.0' --with sqlalchemy --with pandas --with pytest --with pydantic-settings python -m pytest backend/tests/ -q` +Expected: all pass + +```bash +git add backend/models.py backend/data_loader.py backend/app.py backend/tests/test_destinations_api.py +git commit -m "feat(destinations): serve destinations without closing the gaps" +``` + +--- + +### Task 7: Frontend types and section flags + +**Files:** +- Modify: `nextjs-app/lib/types.ts` +- Modify: `nextjs-app/lib/schoolSections.ts:12-46` (the interfaces) and `:48-95` (`computeSchoolFlags`) +- Test: `nextjs-app/__tests__/lib/schoolSections.destinations.test.ts` + +**Interfaces:** +- Consumes: `lib/destinations.ts` types from Task 1 +- Produces: `SchoolDestinations` type; `SchoolFlags.hasKs4Destinations` and `.hasKs5Destinations` + +- [ ] **Step 1: Write the failing test** + +Create `nextjs-app/__tests__/lib/schoolSections.destinations.test.ts`: + +```ts +import { computeSchoolFlags } from '@/lib/schoolSections'; + +const base = { + schoolInfo: { urn: 1, school_name: 'X', phase: 'Secondary', has_sixth_form: true } as any, + yearlyData: [], absenceData: null, census: null, deprivation: null, finance: null, +}; + +const ks4Only = { + ks4: { cohort_year: '2022/23', groups: { all: { cohort: 180, categories: [], aggregates: {} } } }, + ks5: null, +} as any; + +it('flags KS4 destinations when the block is present', () => { + const flags = computeSchoolFlags({ ...base, destinations: ks4Only }); + expect(flags.hasKs4Destinations).toBe(true); + expect(flags.hasKs5Destinations).toBe(false); +}); + +it('flags neither when the block is absent', () => { + const flags = computeSchoolFlags({ ...base, destinations: null }); + expect(flags.hasKs4Destinations).toBe(false); + expect(flags.hasKs5Destinations).toBe(false); +}); + +it('does not flag a group with no categories as renderable', () => { + const empty = { ks4: { cohort_year: '2022/23', groups: {} }, ks5: null } as any; + expect(computeSchoolFlags({ ...base, destinations: empty }).hasKs4Destinations).toBe(false); +}); +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `cd nextjs-app && npx jest __tests__/lib/schoolSections.destinations.test.ts` +Expected: FAIL — `hasKs4Destinations` is undefined + +- [ ] **Step 3: Add the types** + +In `nextjs-app/lib/types.ts`: + +```ts +import type { DestinationCell, PupilGroup } from './destinations'; + +export interface DestinationGroupPayload { + cohort: number | null; + categories: DestinationCell[]; + aggregates: Partial>; +} + +export interface DestinationPhase { + cohort_year: string | null; + groups: Partial>; +} + +export interface SchoolDestinations { + ks4: DestinationPhase | null; + ks5: DestinationPhase | null; +} +``` + +- [ ] **Step 4: Extend the flags** + +In `nextjs-app/lib/schoolSections.ts`, add `destinations: SchoolDestinations | null` to `SchoolFlagsInput`, add `hasKs4Destinations: boolean` and `hasKs5Destinations: boolean` to `SchoolFlags`, destructure `destinations` in `computeSchoolFlags`, and compute: + +```ts + // A phase counts as present only if some group actually carries categories — + // a block with an empty groups map is a pipeline artefact, not a section. + const phaseHasContent = (phase: DestinationPhase | null | undefined) => + !!phase && Object.values(phase.groups ?? {}).some(g => (g?.categories?.length ?? 0) > 0); + + const hasKs4Destinations = phaseHasContent(destinations?.ks4); + const hasKs5Destinations = phaseHasContent(destinations?.ks5); +``` + +Return both from `computeSchoolFlags`. Add nav items `{ id: 'destinations', label: 'After Year 11' }` and `{ id: 'post16-destinations', label: 'After the sixth form' }` in `buildNavItems`, gated on the two flags, positioned after the GCSE entry. + +- [ ] **Step 5: Run test and typecheck** + +Run: `cd nextjs-app && npx jest __tests__/lib/schoolSections.destinations.test.ts && npm run typecheck` +Expected: PASS + +- [ ] **Step 6: Commit** + +```bash +git add nextjs-app/lib/types.ts nextjs-app/lib/schoolSections.ts nextjs-app/__tests__/lib/schoolSections.destinations.test.ts +git commit -m "feat(destinations): types and section flags" +``` + +--- + +### Task 8: The After Year 11 section + +**Files:** +- Create: `nextjs-app/components/school/DestinationsSection.tsx` +- Create: `nextjs-app/components/school/DestinationsView.tsx` +- Create: `nextjs-app/components/school/destinations.module.css` +- Modify: `nextjs-app/components/school/SecondarySchoolSections.tsx` +- Test: `nextjs-app/__tests__/components/DestinationsSection.test.tsx` + +**Interfaces:** +- Consumes: `lib/destinations.ts` (Task 1), `SchoolDestinations` (Task 7), tokens (Task 2) +- Produces: `` + +- [ ] **Step 1: Write the failing test** + +Create `nextjs-app/__tests__/components/DestinationsSection.test.tsx`: + +```tsx +import { render, screen } from '@testing-library/react'; +import { DestinationsSection } from '@/components/school/DestinationsSection'; + +const cell = (category: string, pupils: number | null, status = 'published') => ({ + category, pupils, percentage: pupils === null ? null : (pupils / 180) * 100, status, +}); + +const fullPhase: any = { + cohort_year: '2022/23', + groups: { + all: { + cohort: 180, + categories: [ + cell('school_sixth_form', 75), cell('sixth_form_college', 21), + cell('further_education', 55), cell('other_education', 6), + cell('apprenticeship', 8), cell('employment', 6), + cell('not_sustained', 5), cell('not_captured', 4), + ], + aggregates: {}, + }, + }, +}; + +const suppressedPhase: any = { + cohort_year: '2022/23', + groups: { + all: { + cohort: 180, + categories: [ + cell('school_sixth_form', 75), cell('sixth_form_college', null, 'suppressed'), + cell('further_education', 55), cell('other_education', 6), + cell('apprenticeship', 8), cell('employment', 6), + cell('not_sustained', 5), cell('not_captured', 4), + ], + aggregates: {}, + }, + }, +}; + +it('dates its own cohort so it is not read as stale', () => { + render(); + expect(screen.getByText(/2022\/23/)).toBeInTheDocument(); +}); + +it('renders the bar when the group is fully published', () => { + const { container } = render(); + expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(8); +}); + +it('renders NO bar when a category is withheld', () => { + const { container } = render(); + expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(0); + expect(screen.getByText(/withheld/i)).toBeInTheDocument(); +}); + +it('never states a remainder for a partially suppressed group', () => { + const { container } = render(); + // 180 cohort - 155 published = 25, the withheld figure. It must appear nowhere. + expect(container.textContent).not.toMatch(/\b25\b/); +}); + +it('never claims a pupil stayed at this school', () => { + const { container } = render(); + expect(container.textContent).not.toMatch(/stayed on (here|at)/i); +}); +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `cd nextjs-app && npx jest __tests__/components/DestinationsSection.test.tsx` +Expected: FAIL — cannot find module + +- [ ] **Step 3: Write the server section** + +Create `nextjs-app/components/school/DestinationsSection.tsx`. It renders the `Section` shell, the title, a subtitle naming the cohort year and the publication lag, and delegates the interactive body to `DestinationsView` with `all` as the server-rendered default. Mark each bar segment with `data-destination-segment` so the tests and the E2E journeys can assert on its absence. + +Key structure: + +```tsx +/** + * DestinationsSection — where a school's Year 11 leavers went. Server component. + * + * The headline is deliberately NOT the sustained-destination rate: that figure + * sits between 92% and 97% for nearly every school in England, so leading with + * it would say nothing. The mix is what varies. + */ +import type { DestinationPhase } from '@/lib/types'; +import { Section, sectionStyles } from './sectionShared'; +import { DestinationsView } from './DestinationsView'; + +export function DestinationsSection({ + destinations, schoolName, +}: { destinations: DestinationPhase; schoolName: string }) { + const cohort = destinations.groups.all?.cohort ?? null; + return ( +
+

After Year 11

+

+ Where {cohort ? `the ${cohort} pupils` : 'the pupils'} who left Year 11 in{' '} + {destinations.cohort_year ?? 'the most recent year published'} went next. + Destination measures are published about two years after the exams above. +

+ +
+ ); +} +``` + +- [ ] **Step 4: Write the client view** + +Create `nextjs-app/components/school/DestinationsView.tsx` with `'use client'`. It owns the cohort switch (`role="radiogroup"`, arrow-key navigation), the card↔bar hover linkage, and the render decisions: + +- Cards from `CARD_GROUPS` — `aggregateCells` for the value, or a "Not published" card when it returns `null` +- Bar only when `canRenderBar(group)`; otherwise a panel explaining that the categories add up to the cohort so the rest cannot be drawn +- A published aggregate is shown only when `canRenderPublishedAggregate(components)` is true +- The full table always, with withheld rows marked +- Segment widths from `toBarSegments`, each carrying `data-destination-segment` + +Wrap the `toBarSegments` call in the `canRenderBar` guard rather than a try/catch — the throw is a backstop for programmer error, not control flow. + +- [ ] **Step 5: Write the stylesheet** + +Create `nextjs-app/components/school/destinations.module.css` using only the tokens from Task 2 plus the existing section tokens. The absence segment is `background-image: repeating-linear-gradient(45deg, var(--dest-none-hatch) 0 3px, transparent 3px 7px)` over `var(--bg-card)` with a `1px` inset ring in `var(--dest-none)`. Segments sit in a flex row with `gap: 2px`. + +- [ ] **Step 6: Mount it on the secondary template** + +In `nextjs-app/components/school/SecondarySchoolSections.tsx`, render `` after the GCSE section and before admissions, gated on `flags.hasKs4Destinations`. + +- [ ] **Step 7: Run tests, typecheck, commit** + +Run: `cd nextjs-app && npx jest __tests__/components/DestinationsSection.test.tsx && npm run typecheck` +Expected: PASS, 5 tests + +```bash +git add nextjs-app/components/school/DestinationsSection.tsx \ + nextjs-app/components/school/DestinationsView.tsx \ + nextjs-app/components/school/destinations.module.css \ + nextjs-app/components/school/SecondarySchoolSections.tsx \ + nextjs-app/__tests__/components/DestinationsSection.test.tsx +git commit -m "feat(destinations): the After Year 11 section" +``` + +--- + +### Task 9: The post-16 section, and removing the placeholder + +**Files:** +- Create: `nextjs-app/components/school/Post16DestinationsSection.tsx` +- Modify: `nextjs-app/components/school/SecondaryAdmissionsSection.tsx:110-120` +- Modify: `nextjs-app/components/school/SecondarySchoolSections.tsx` +- Test: `nextjs-app/__tests__/components/Post16DestinationsSection.test.tsx` + +**Interfaces:** +- Consumes: everything from Task 8; reuses `DestinationsView` +- Produces: `` + +- [ ] **Step 1: Write the failing test** + +Create `nextjs-app/__tests__/components/Post16DestinationsSection.test.tsx`: + +```tsx +import { render, screen } from '@testing-library/react'; +import { Post16DestinationsSection } from '@/components/school/Post16DestinationsSection'; + +const phase: any = { + cohort_year: '2022/23', + groups: { + all: { + cohort: 96, + categories: [ + { category: 'higher_education', pupils: 56, percentage: 58.3, status: 'published' }, + { category: 'further_education', pupils: 12, percentage: 12.5, status: 'published' }, + { category: 'apprenticeship', pupils: 9, percentage: 9.4, status: 'published' }, + { category: 'employment', pupils: 13, percentage: 13.5, status: 'published' }, + { category: 'not_sustained', pupils: 6, percentage: 6.3, status: 'published' }, + ], + aggregates: {}, + }, + }, +}; + +it('names the Year 13 cohort, not Year 11', () => { + render(); + expect(screen.getByText(/Year 13/)).toBeInTheDocument(); +}); + +it('reports university destinations', () => { + render(); + expect(screen.getByText(/higher education|university/i)).toBeInTheDocument(); +}); +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `cd nextjs-app && npx jest __tests__/components/Post16DestinationsSection.test.tsx` +Expected: FAIL — cannot find module + +- [ ] **Step 3: Write the section** + +Create `nextjs-app/components/school/Post16DestinationsSection.tsx` mirroring `DestinationsSection` with `id="post16-destinations"`, the heading "After the sixth form", copy naming Year 13, and the same `DestinationsView` body. + +- [ ] **Step 4: Remove the placeholder** + +In `nextjs-app/components/school/SecondaryAdmissionsSection.tsx`, delete the "Post-16 destination data coming soon" paragraph at line ~117 and its surrounding conditional. The sixth-form badge in the header stays. + +- [ ] **Step 5: Mount it** + +In `SecondarySchoolSections.tsx`, render `` after ``, gated on `flags.hasKs5Destinations`. Where the school has no sixth form the section is simply not rendered — no placeholder, because absence is the correct statement. + +- [ ] **Step 6: Confirm the placeholder is gone, run tests, commit** + +Run: `grep -rn "coming soon" nextjs-app/components/` +Expected: no output + +Run: `cd nextjs-app && npx jest && npm run typecheck` +Expected: PASS + +```bash +git add nextjs-app/components/school/Post16DestinationsSection.tsx \ + nextjs-app/components/school/SecondaryAdmissionsSection.tsx \ + nextjs-app/components/school/SecondarySchoolSections.tsx \ + nextjs-app/__tests__/components/Post16DestinationsSection.test.tsx +git commit -m "feat(destinations): the post-16 section, replacing the placeholder" +``` + +--- + +### Task 10: E2E journeys + +Per CLAUDE.md, user-facing behaviour extends `e2e/` in the same PR. Note the staging E2E gate runs post-merge — these journeys cannot pass in PR checks until the Airflow DAG has populated the marts on staging. + +**Files:** +- Modify: `e2e/tests/journeys.spec.ts` + +**Interfaces:** +- Consumes: the rendered pages from Tasks 8 and 9 +- Produces: nothing + +- [ ] **Step 1: Write the journeys** + +Append to `e2e/tests/journeys.spec.ts`: + +```ts +test('a secondary school page says where its Year 11 leavers went', async ({ page }) => { + await page.goto('/school/abbey-grange-church-of-england-academy-137083'); + const section = page.locator('#destinations'); + await expect(section).toBeVisible(); + // The section must date its own cohort — destinations run two GCSE years + // behind the results above them, and an undated figure reads as stale. + await expect(section).toContainText(/20\d{2}\/\d{2}/); +}); + +test('the destinations section draws no bar for a group with withheld figures', async ({ page }) => { + await page.goto('/school/abbey-grange-church-of-england-academy-137083'); + const section = page.locator('#destinations'); + await section.getByRole('radio', { name: /disadvantaged/i }).click(); + + const withheld = section.getByText(/withheld/i); + if (await withheld.count() > 0) { + // R1: where anything is withheld, the bar must be absent entirely — a bar + // with a gap in it publishes the withheld figure by its width. + await expect(section.locator('[data-destination-segment]')).toHaveCount(0); + } +}); + +test('a school with no sixth form has no post-16 destinations section', async ({ page }) => { + await page.goto('/school/abbey-grange-church-of-england-academy-137083'); + const hasSixthForm = await page.getByText(/sixth form/i).count() > 0; + if (!hasSixthForm) { + await expect(page.locator('#post16-destinations')).toHaveCount(0); + } +}); + +test('the destinations section never claims a pupil stayed at this school', async ({ page }) => { + await page.goto('/school/abbey-grange-church-of-england-academy-137083'); + const text = await page.locator('#destinations').textContent(); + // The published file reports destination TYPE, never destination institution. + expect(text ?? '').not.toMatch(/stayed on (here|at this school)/i); +}); +``` + +- [ ] **Step 2: Verify the slug resolves** + +Run: `grep -n "school/" e2e/tests/journeys.spec.ts | head -5` + +Match the slug format the existing school-page journeys use. If they build slugs from an API call rather than hardcoding, follow that pattern instead of the literal above. + +- [ ] **Step 3: Commit** + +```bash +git add e2e/tests/journeys.spec.ts +git commit -m "test(e2e): destination journeys, including the no-bar rule" +``` + +--- + +## Self-Review + +**Spec coverage.** Every section of the design maps to a task: disclosure rules → Task 1 (guards) and Task 5 (mart tests); availability/extraction → Task 3; staging and the `safe_numeric` prohibition → Task 4; marts and R3 → Task 5; API → Task 6; display → Tasks 2, 7, 8, 9; edge states → Tasks 8 and 9; testing → every task plus Task 10. + +**Gap found and closed.** A first draft had Task 3 drop every non-school row, which would have left `fact_destination_national` (Task 5) reading an empty table — and a note telling the executor to go back and amend an earlier task. Task 3 now keeps national rows with `urn = None` from the start, and its tests cover both that and the LA rows that must still be dropped. + +**Type consistency.** `row_to_record` reads `geographicLevel`, not the `locations` keys, because a school row also carries `NAT`, `LA` and `REG` entries for its parents — keying off `"NAT" in locations` would admit every LA row as national. `status` takes the same three values in the tap, the staging models, the marts, the SQLAlchemy models, the API and `lib/destinations.ts`. `pupil_group` is `all` / `disadvantaged` / `other` throughout.