Files
school_compare/docs/superpowers/plans/2026-08-28-destination-measures.md
T

1754 lines
70 KiB
Markdown
Raw Normal View History

# Destination Measures Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Show what happened to a school's leavers after Year 11 and after the sixth form, on secondary school detail pages, without ever republishing a figure DfE withheld.
**Architecture:** A new Meltano tap pulls the EES destinations query API into `raw`; dbt staging preserves the `c` suppression sentinel as a status column rather than nulling it; long-format marts carry one row per school × year × pupil group × destination category; the API serialises a `destinations` block; a server component renders all-pupils into the HTML with one client component for the cohort switch. All three disclosure rules live as executable guards in `lib/destinations.ts`.
**Tech Stack:** Python 3.12 / Singer SDK / Meltano · dbt + PostgreSQL · FastAPI + SQLAlchemy · Next.js (App Router) + TypeScript + CSS Modules · Jest · Playwright
**Spec:** `docs/superpowers/specs/2026-08-28-destination-measures-design.md`
## Global Constraints
- **R1 — Never render a derived remainder.** Not as a number, not as a bar segment. Where any category in a pupil group is suppressed, no bar is drawn for that group.
- **R2 — Never aggregate across a suppression boundary.** Compute an aggregate from components only when every component is published. Render a DfE-published aggregate only when the count of suppressed components within it is 0 or ≥ 2.
- **R3 — Where a category is suppressed for the disadvantaged group, it is also withheld for the other-pupils group.** The all-pupils view keeps it. Enforced in the mart.
- **`safe_numeric` must never be applied to a destination count or percentage.** It coerces `c` to `NULL`, destroying the distinction between *withheld* and *no data*.
- Destination categories, verbatim from the EES filter: `School sixth form`, `Sixth form college`, `Further education`, `Other education destination`, `Sustained apprenticeships`, `Sustained employment destination`, `Not recorded as a sustained destination`, `Activity not captured`. Aggregates: `Sustained education destination`, `Sustained education, employment & apprenticeships`.
- Card grouping (ours, not DfE's): academic = school sixth form + sixth-form college; college = further education + other education; work = apprenticeship + employment.
- Copy must never imply a pupil "stayed on here" — the file reports destination *type*, never destination *institution*.
- Every new colour is a token in `nextjs-app/app/globals.css`, defined in `:root` and in both dark blocks. Never style a component from inside a theme block.
- Percentages for display are rounded; bar widths derive from unrounded pupil counts.
- EES API: `https://api.education.gov.uk/statistics/v1`. KS4 dataset `019d4f41-22d1-71b2-a1a7-f3b91026815b`; 16-18 dataset `019d4e73-6440-7523-b60c-bfab1ad4a30d`. Time periods use the `2022/2023` form, not `2022/23`.
**Pipeline reality:** `dbt` and `meltano` do not run locally. Tasks 3–5 are verified by unit tests and by SQL review; the models only produce data once Tudor triggers the Airflow DAG on staging. Do not claim mart data exists until that has run.
---
### Task 1: Destination domain logic
The disclosure rules are here, in pure functions, so they can be tested without a database, a network, or a browser. Every later task consumes this module.
**Files:**
- Create: `nextjs-app/lib/destinations.ts`
- Test: `nextjs-app/__tests__/lib/destinations.test.ts`
**Interfaces:**
- Consumes: nothing
- Produces:
- `type DestinationCategory` — the eight category slugs
- `type PupilGroup = 'all' | 'disadvantaged' | 'other'`
- `type DestinationStatus = 'published' | 'suppressed' | 'not_applicable'`
- `interface DestinationCell { category; pupils: number | null; percentage: number | null; status }`
- `interface DestinationGroup { cohort: number; cells: DestinationCell[]; aggregates: Record<string, DestinationCell> }`
- `CARD_GROUPS: Record<CardGroup, DestinationCategory[]>`
- `canAggregate(cells: DestinationCell[]): boolean`
- `aggregateCells(cells: DestinationCell[], cohort: number): { pupils: number; percentage: number } | null`
- `canRenderPublishedAggregate(components: DestinationCell[]): boolean`
- `canRenderBar(group: DestinationGroup): boolean`
- `toBarSegments(group: DestinationGroup): { category; pupils; widthPct; labelPct }[]`
- `suppressedCount(cells: DestinationCell[]): number`
- [ ] **Step 1: Write the failing test**
Create `nextjs-app/__tests__/lib/destinations.test.ts`:
```ts
import {
canAggregate, aggregateCells, canRenderPublishedAggregate,
canRenderBar, toBarSegments, CARD_GROUPS,
type DestinationCell, type DestinationGroup,
} from '@/lib/destinations';
const pub = (category: any, pupils: number, cohort: number): DestinationCell => ({
category, pupils, percentage: (pupils / cohort) * 100, status: 'published',
});
const sup = (category: any): DestinationCell => ({
category, pupils: null, percentage: null, status: 'suppressed',
});
const fullGroup = (): DestinationGroup => ({
cohort: 180,
cells: [
pub('school_sixth_form', 75, 180), pub('sixth_form_college', 21, 180),
pub('further_education', 55, 180), pub('other_education', 6, 180),
pub('apprenticeship', 8, 180), pub('employment', 6, 180),
pub('not_sustained', 5, 180), pub('not_captured', 4, 180),
],
aggregates: {},
});
describe('canAggregate — R2, computing from components', () => {
it('allows a sum when every component is published', () => {
expect(canAggregate([pub('apprenticeship', 8, 180), pub('employment', 6, 180)])).toBe(true);
});
it('refuses a sum when any component is suppressed', () => {
expect(canAggregate([pub('apprenticeship', 8, 180), sup('employment')])).toBe(false);
});
it('refuses a sum when every component is suppressed', () => {
expect(canAggregate([sup('apprenticeship'), sup('employment')])).toBe(false);
});
});
describe('aggregateCells', () => {
it('sums published cells and derives a percentage from the cohort', () => {
expect(aggregateCells([pub('apprenticeship', 8, 180), pub('employment', 6, 180)], 180))
.toEqual({ pupils: 14, percentage: (14 / 180) * 100 });
});
it('returns null rather than a partial sum when a component is suppressed', () => {
expect(aggregateCells([pub('apprenticeship', 8, 180), sup('employment')], 180)).toBeNull();
});
});
describe('canRenderPublishedAggregate — R2, a total DfE published itself', () => {
it('allows it when no component is suppressed', () => {
expect(canRenderPublishedAggregate([pub('school_sixth_form', 75, 180), pub('sixth_form_college', 21, 180)])).toBe(true);
});
it('REFUSES it when exactly one component is suppressed — the aggregate identifies it', () => {
expect(canRenderPublishedAggregate([pub('school_sixth_form', 75, 180), sup('sixth_form_college')])).toBe(false);
});
it('allows it when two or more components are suppressed', () => {
expect(canRenderPublishedAggregate([sup('school_sixth_form'), sup('sixth_form_college')])).toBe(true);
});
});
describe('canRenderBar — R1', () => {
it('allows a bar when the whole group is published', () => {
expect(canRenderBar(fullGroup())).toBe(true);
});
it('refuses a bar when a single category is suppressed', () => {
const g = fullGroup();
g.cells[1] = sup('sixth_form_college');
expect(canRenderBar(g)).toBe(false);
});
});
describe('toBarSegments', () => {
it('derives widths from counts, not from rounded percentages', () => {
const segs = toBarSegments(fullGroup());
expect(segs).toHaveLength(8);
expect(segs[0].widthPct).toBeCloseTo((75 / 180) * 100, 10);
expect(segs.reduce((a, s) => a + s.widthPct, 0)).toBeCloseTo(100, 6);
});
it('throws rather than silently leaving a gap when the group is suppressed', () => {
const g = fullGroup();
g.cells[1] = sup('sixth_form_college');
expect(() => toBarSegments(g)).toThrow(/suppressed/i);
});
});
describe('CARD_GROUPS', () => {
it('partitions every destination category exactly once, plus the absence', () => {
const grouped = Object.values(CARD_GROUPS).flat();
expect(new Set(grouped).size).toBe(grouped.length);
expect(grouped).toEqual(expect.arrayContaining([
'school_sixth_form', 'sixth_form_college', 'further_education',
'other_education', 'apprenticeship', 'employment',
]));
expect(grouped).not.toContain('not_sustained');
expect(grouped).not.toContain('not_captured');
});
});
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd nextjs-app && npx jest __tests__/lib/destinations.test.ts`
Expected: FAIL — `Cannot find module '@/lib/destinations'`
- [ ] **Step 3: Write the implementation**
Create `nextjs-app/lib/destinations.ts`:
```ts
/**
* Destination measures — categories, the card grouping, and the disclosure
* guards.
*
* DfE suppresses individual cells with `c`, and the categories sum to the
* cohort. So subtracting the published cells from the cohort total recovers a
* lone suppressed cell exactly — on 22% of mainstream secondaries. The guards
* below are what stop this module's consumers doing that by accident, and they
* are the reason percentages are never reconstructed from a partial sum.
*
* See docs/superpowers/specs/2026-08-28-destination-measures-design.md.
*/
export type DestinationCategory =
| 'school_sixth_form'
| 'sixth_form_college'
| 'further_education'
| 'other_education'
| 'apprenticeship'
| 'employment'
| 'not_sustained'
| 'not_captured';
export type PupilGroup = 'all' | 'disadvantaged' | 'other';
export type DestinationStatus = 'published' | 'suppressed' | 'not_applicable';
export type CardGroup = 'academic' | 'college' | 'work';
export interface DestinationCell {
category: DestinationCategory;
pupils: number | null;
percentage: number | null;
status: DestinationStatus;
}
export interface DestinationGroup {
cohort: number;
cells: DestinationCell[];
/** Aggregates DfE published itself, keyed by slug. */
aggregates: Partial<Record<'sustained_education' | 'sustained_all', DestinationCell>>;
}
/** Display order, which is also bar order: education, then work, then absence. */
export const CATEGORY_ORDER: DestinationCategory[] = [
'school_sixth_form', 'sixth_form_college', 'further_education', 'other_education',
'apprenticeship', 'employment', 'not_sustained', 'not_captured',
];
/**
* Our grouping, not DfE's — the single most arguable thing on the page, which
* is why it lives in exactly one place. `not_sustained` and `not_captured` are
* deliberately absent: they are an absence of destination, not a route.
*/
export const CARD_GROUPS: Record<CardGroup, DestinationCategory[]> = {
academic: ['school_sixth_form', 'sixth_form_college'],
college: ['further_education', 'other_education'],
work: ['apprenticeship', 'employment'],
};
export function suppressedCount(cells: DestinationCell[]): number {
return cells.filter(c => c.status === 'suppressed').length;
}
/** R2: a sum computed from components is safe only if every component is published. */
export function canAggregate(cells: DestinationCell[]): boolean {
return cells.length > 0 && cells.every(c => c.status === 'published');
}
export function aggregateCells(
cells: DestinationCell[], cohort: number,
): { pupils: number; percentage: number } | null {
if (!canAggregate(cells) || cohort <= 0) return null;
const pupils = cells.reduce((sum, c) => sum + (c.pupils ?? 0), 0);
return { pupils, percentage: (pupils / cohort) * 100 };
}
/**
* R2, the other direction: DfE published this total itself. Showing it beside
* the components is safe only when it spans no suppressed component, or two or
* more. Exactly one and the total names the withheld figure.
*/
export function canRenderPublishedAggregate(components: DestinationCell[]): boolean {
return suppressedCount(components) !== 1;
}
/** R1: a bar is drawable only when nothing in the group is withheld. */
export function canRenderBar(group: DestinationGroup): boolean {
return group.cohort > 0 && group.cells.every(c => c.status === 'published');
}
export interface BarSegment {
category: DestinationCategory;
pupils: number;
/** Exact width from the count — never the rounded percentage. */
widthPct: number;
/** Rounded value for the segment label. */
labelPct: number;
}
export function toBarSegments(group: DestinationGroup): BarSegment[] {
if (!canRenderBar(group)) {
throw new Error(
'toBarSegments: refusing to draw a bar for a group with suppressed categories — '
+ 'the gap would disclose the withheld figure (R1).',
);
}
const byCategory = new Map(group.cells.map(c => [c.category, c]));
return CATEGORY_ORDER.flatMap(category => {
const cell = byCategory.get(category);
if (!cell || cell.pupils === null) return [];
const widthPct = (cell.pupils / group.cohort) * 100;
return [{ category, pupils: cell.pupils, widthPct, labelPct: Math.round(widthPct) }];
});
}
export const CATEGORY_LABELS: Record<DestinationCategory, string> = {
school_sixth_form: 'State-funded school sixth form',
sixth_form_college: 'Sixth-form college',
further_education: 'FE and other colleges',
other_education: 'Other education destination',
apprenticeship: 'Apprenticeship',
employment: 'Employment',
not_sustained: 'Not recorded as a sustained destination',
not_captured: 'Activity not captured',
};
export const CARD_QUESTIONS: Record<CardGroup, { question: string; hint: string }> = {
academic: { question: 'Do leavers stay on an academic route?', hint: 'a school sixth form or a sixth-form college' },
college: { question: 'Or move to a college?', hint: 'an FE or other college' },
work: { question: 'Or straight into work?', hint: 'an apprenticeship or a job' },
};
```
- [ ] **Step 4: Run test to verify it passes**
Run: `cd nextjs-app && npx jest __tests__/lib/destinations.test.ts`
Expected: PASS, 12 tests
- [ ] **Step 5: Typecheck and commit**
```bash
cd nextjs-app && npm run typecheck
cd .. && git add nextjs-app/lib/destinations.ts nextjs-app/__tests__/lib/destinations.test.ts
git commit -m "feat(destinations): the disclosure rules, as executable guards"
```
---
### Task 2: Destination colour tokens
Six tokens, both themes. The absence is neutral plus a hatch rather than a colour, which is both a factual point (it isn't a bad outcome) and the secondary encoding that rescues a failing CVD pair.
**Files:**
- Modify: `nextjs-app/app/globals.css` (`:root`, the `prefers-color-scheme` block, and the `[data-theme="dark"]` block if one exists)
- Test: `nextjs-app/__tests__/components/darkThemeSafety.test.ts`
**Interfaces:**
- Consumes: nothing
- Produces: CSS custom properties `--dest-sixthform`, `--dest-sfcollege`, `--dest-fecollege`, `--dest-apprentice`, `--dest-employment`, `--dest-none`, `--dest-none-hatch`
- [ ] **Step 1: Read the existing token blocks**
Run: `grep -n "\-\-series-1\|prefers-color-scheme" nextjs-app/app/globals.css`
Add the new tokens immediately after the `--series-*` group in each block so the palette stays in one place.
- [ ] **Step 2: Write the failing test**
Append to `nextjs-app/__tests__/components/darkThemeSafety.test.ts`:
```ts
describe('destination tokens', () => {
const css = readFileSync(join(process.cwd(), 'app/globals.css'), 'utf8');
const tokens = [
'--dest-sixthform', '--dest-sfcollege', '--dest-fecollege',
'--dest-apprentice', '--dest-employment', '--dest-none', '--dest-none-hatch',
];
it('defines every destination token in the light palette', () => {
const root = css.slice(css.indexOf(':root {'), css.indexOf('@media (prefers-color-scheme: dark)'));
tokens.forEach(t => expect(root).toContain(t + ':'));
});
it('redefines every destination token for dark', () => {
const dark = css.slice(css.indexOf('@media (prefers-color-scheme: dark)'));
tokens.forEach(t => expect(dark).toContain(t + ':'));
});
});
```
- [ ] **Step 3: Run test to verify it fails**
Run: `cd nextjs-app && npx jest __tests__/components/darkThemeSafety.test.ts`
Expected: FAIL — light palette missing `--dest-sixthform:`
- [ ] **Step 4: Add the tokens**
In the `:root` block:
```css
/* ── Destination measures ───────────────────────────────────────────
Education is one hue in three steps (school-like → college-like) so the
three education destinations read as one family; apprenticeship and
employment are separate hues. The absence is neutral and hatched, never
a colour — "activity not captured" includes independent schools and
moving abroad, so rendering it as a bad outcome would be wrong. The
hatch is also what rescues the neutral/blue pair, which fails CVD
separation at ΔE 7.6 as flat fills. Every other adjacent pair clears
ΔE 10.9 under protanopia. */
--dest-sixthform: #0F766E;
--dest-sfcollege: #4A9E96;
--dest-fecollege: #7CBFB8;
--dest-apprentice: #806200;
--dest-employment: #2F6F8F;
--dest-none: #6B7580;
--dest-none-hatch: rgba(107, 117, 128, 0.34);
```
In the `@media (prefers-color-scheme: dark)` block (and the `[data-theme="dark"]` block if present):
```css
--dest-sixthform: #5FC7BB;
--dest-sfcollege: #3E9B92;
--dest-fecollege: #2A716B;
--dest-apprentice: #EFC658;
--dest-employment: #8FB4D9;
--dest-none: #8B9AA1;
--dest-none-hatch: rgba(139, 154, 161, 0.34);
```
- [ ] **Step 5: Run test to verify it passes**
Run: `cd nextjs-app && npx jest __tests__/components/darkThemeSafety.test.ts`
Expected: PASS
- [ ] **Step 6: Commit**
```bash
git add nextjs-app/app/globals.css nextjs-app/__tests__/components/darkThemeSafety.test.ts
git commit -m "feat(destinations): colour tokens, with the absence hatched not coloured"
```
---
### Task 3: The destinations tap
A separate extractor from `tap-uk-ees`. That tap downloads a release ZIP and reads a CSV inside it; the destinations files carry every breakdown we don't want, so this one POSTs to the query API and pages.
**Files:**
- Create: `pipeline/plugins/extractors/tap-uk-ees-destinations/pyproject.toml`
- Create: `pipeline/plugins/extractors/tap-uk-ees-destinations/tap_uk_ees_destinations/__init__.py`
- Create: `pipeline/plugins/extractors/tap-uk-ees-destinations/tap_uk_ees_destinations/tap.py`
- Modify: `pipeline/meltano.yml`
- Modify: `pipeline/dags/school_data_pipeline.py:160-196`
- Test: `pipeline/plugins/extractors/tap-uk-ees-destinations/tests/test_tap.py`
**Interfaces:**
- Consumes: nothing
- Produces: raw tables `ees_ks4_destinations` and `ees_ks5_destinations`, columns `urn`, `time_period`, `pupil_group`, `destination_measure`, `cohort_pupils`, `pupils_raw`, `percentage_raw` — the last two as **text**, sentinel preserved.
- [ ] **Step 1: Write the failing test**
Create `pipeline/plugins/extractors/tap-uk-ees-destinations/tests/test_tap.py`:
```python
"""The tap's only job that can be tested without the network: mapping an API
row to a Singer record without destroying the suppression sentinel."""
from tap_uk_ees_destinations.tap import row_to_record, DESTINATION_SLUGS, PUPIL_GROUP_SLUGS
def test_published_row_keeps_its_numbers_as_text():
row = {
"timePeriod": {"period": "2022/2023"},
"geographicLevel": "SCH",
"locations": {"SCH": "IXn5B"},
"filters": {"wYXbx": "DCz1Q", "9ss4v": "p9WRS"},
"values": {"Poghe": "264", "1roqi": "182", "dPjk0": "68.9"},
}
rec = row_to_record(row, urn_by_location={"IXn5B": "137083"})
assert rec["urn"] == "137083"
assert rec["time_period"] == "202223"
assert rec["destination_measure"] == "school_sixth_form"
assert rec["pupil_group"] == "all"
assert rec["cohort_pupils"] == "264"
assert rec["pupils_raw"] == "182"
assert rec["percentage_raw"] == "68.9"
def test_suppressed_row_preserves_the_c_sentinel():
row = {
"timePeriod": {"period": "2022/2023"},
"geographicLevel": "SCH",
"locations": {"SCH": "IXn5B"},
"filters": {"wYXbx": "eLsdu", "9ss4v": "OvPnC"},
"values": {"Poghe": "41", "1roqi": "c", "dPjk0": "c"},
}
rec = row_to_record(row, urn_by_location={"IXn5B": "137083"})
assert rec["pupils_raw"] == "c", "the sentinel must survive extraction"
assert rec["percentage_raw"] == "c"
assert rec["pupil_group"] == "disadvantaged"
def test_national_rows_are_kept_with_a_null_urn():
"""The England reference lives in the same response. It is kept, with urn
None, so fact_destination_national has something to read."""
row = {
"timePeriod": {"period": "2022/2023"},
"geographicLevel": "NAT",
"locations": {"NAT": "dP0Zw"},
"filters": {"wYXbx": "DCz1Q", "9ss4v": "p9WRS"},
"values": {"Poghe": "500000", "1roqi": "190000", "dPjk0": "38.0"},
}
rec = row_to_record(row, urn_by_location={})
assert rec is not None
assert rec["urn"] is None
assert rec["percentage_raw"] == "38.0"
def test_other_geographic_levels_are_dropped():
"""Local authority, district, region and constituency rows are noise here."""
row = {
"timePeriod": {"period": "2022/2023"},
"geographicLevel": "LA",
"locations": {"LA": "u9Oo4", "NAT": "dP0Zw"},
"filters": {"wYXbx": "DCz1Q", "9ss4v": "p9WRS"},
"values": {"Poghe": "1", "1roqi": "1", "dPjk0": "1"},
}
assert row_to_record(row, urn_by_location={}) is None
def test_every_slug_maps_to_one_filter_id():
assert len(set(DESTINATION_SLUGS.values())) == len(DESTINATION_SLUGS)
assert set(PUPIL_GROUP_SLUGS.values()) == {"all", "disadvantaged", "other"}
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd pipeline/plugins/extractors/tap-uk-ees-destinations && uv run --with singer-sdk --with requests pytest tests/ -v`
Expected: FAIL — `ModuleNotFoundError: tap_uk_ees_destinations`
- [ ] **Step 3: Write the tap**
Create `pipeline/plugins/extractors/tap-uk-ees-destinations/pyproject.toml`:
```toml
[project]
name = "tap-uk-ees-destinations"
version = "0.1.0"
requires-python = ">=3.10"
dependencies = ["singer-sdk>=0.40", "requests>=2.31"]
[project.scripts]
tap-uk-ees-destinations = "tap_uk_ees_destinations.tap:TapUKEESDestinations.cli"
[build-system]
requires = ["hatchling"]
build-backend = "hatchling.build"
```
Create `pipeline/plugins/extractors/tap-uk-ees-destinations/tap_uk_ees_destinations/__init__.py` (empty file).
Create `pipeline/plugins/extractors/tap-uk-ees-destinations/tap_uk_ees_destinations/tap.py`:
```python
"""EES destinations tap — KS4 and 16-18 destination measures, school level.
Separate from tap-uk-ees on purpose. That tap pulls a release ZIP and reads a
CSV inside it; the destinations files carry sex, ethnicity, FSM status, prior
attainment and SEN in the same table, so the whole-file route would download
millions of rows to keep a few hundred thousand. The query API filters server
side.
The one thing this tap must not do is tidy up the data. EES writes `c` where a
figure is withheld, and the categories sum to the cohort — so turning `c` into
NULL here would let a downstream sum reconstruct exactly what DfE suppressed.
Counts and percentages are emitted as TEXT, sentinel intact.
"""
from __future__ import annotations
import requests
from singer_sdk import Stream, Tap
from singer_sdk import typing as th
API_BASE = "https://api.education.gov.uk/statistics/v1"
TIMEOUT = 180
PAGE_SIZE = 10000
KS4_DATASET = "019d4f41-22d1-71b2-a1a7-f3b91026815b"
KS5_DATASET = "019d4e73-6440-7523-b60c-bfab1ad4a30d"
# Filter option ids, read from each dataset's /meta. KS4 and 16-18 use
# different ids for the same concepts, so they are declared separately.
KS4_DESTINATION_SLUGS = {
"DCz1Q": "school_sixth_form",
"eLsdu": "sixth_form_college",
"o2MJm": "further_education",
"b7v6t": "other_education",
"mlKo9": "apprenticeship",
"QIJEw": "employment",
"RZrek": "not_sustained",
"1j7Ui": "not_captured",
"WiEl2": "agg_sustained_education",
"EfSAq": "agg_sustained_all",
}
KS4_PUPIL_GROUP_SLUGS = {"p9WRS": "all", "OvPnC": "disadvantaged", "7VmdX": "other"}
# Sex=Total and characteristic topic=Total. Leaving these unpinned returns
# every breakdown crossed with every other — 45 rows where 9 are wanted.
KS4_PINNED = ["X542f", "jHdaA"]
KS5_DESTINATION_SLUGS = {
"dkyu0": "higher_education",
"QIB6w": "school_sixth_form",
"9c3u4": "sixth_form_college",
"S4UgV": "further_education",
"EfYDq": "other_education",
"PaLTe": "apprenticeship",
"9c8k4": "employment",
"Wi6R2": "not_sustained",
"o2l7m": "not_captured",
"wBXtb": "agg_sustained_education",
"o2c4m": "agg_sustained_all",
}
KS5_PUPIL_GROUP_SLUGS = {"Y0PuH": "all", "HTeez": "disadvantaged", "CnSVI": "other"}
KS5_PINNED: list[str] = []
# Indicator ids are shared across both datasets.
IND_COHORT = "Poghe"
IND_PUPILS = "1roqi"
IND_PERCENT = "dPjk0"
DESTINATION_SLUGS = KS4_DESTINATION_SLUGS
PUPIL_GROUP_SLUGS = KS4_PUPIL_GROUP_SLUGS
def _period_to_time_period(period: str) -> str:
"""'2022/2023' -> '202223', matching the convention the other marts use."""
start, end = period.split("/")
return start + end[-2:]
def row_to_record(
row: dict,
urn_by_location: dict[str, str],
destination_slugs: dict[str, str] | None = None,
pupil_group_slugs: dict[str, str] | None = None,
) -> dict | None:
"""Map one API row to a Singer record, or None if it is not school level."""
destination_slugs = destination_slugs or KS4_DESTINATION_SLUGS
pupil_group_slugs = pupil_group_slugs or KS4_PUPIL_GROUP_SLUGS
# Read geographicLevel, not the locations keys: a school row also carries
# NAT, LA and REG entries for its parents, so "NAT in locations" is true
# for every row in the file and would let LA rows through as national ones.
level = row.get("geographicLevel")
if level == "SCH":
urn = urn_by_location.get(row.get("locations", {}).get("SCH"))
if not urn:
return None
elif level == "NAT":
urn = None
else:
return None
filters = row.get("filters", {})
destination = next(
(slug for fid, slug in destination_slugs.items() if fid in filters.values()), None
)
group = next(
(slug for fid, slug in pupil_group_slugs.items() if fid in filters.values()), None
)
if destination is None or group is None:
return None
values = row.get("values", {})
return {
"urn": urn,
"time_period": _period_to_time_period(row["timePeriod"]["period"]),
"pupil_group": group,
"destination_measure": destination,
"cohort_pupils": values.get(IND_COHORT),
"pupils_raw": values.get(IND_PUPILS),
"percentage_raw": values.get(IND_PERCENT),
}
def fetch_urn_by_location(dataset_id: str) -> dict[str, str]:
"""Location id -> URN, from the dataset's meta."""
resp = requests.get(f"{API_BASE}/data-sets/{dataset_id}/meta", timeout=TIMEOUT)
resp.raise_for_status()
for group in resp.json().get("locations", []):
if group.get("level", {}).get("code") == "SCH":
return {o["id"]: o["urn"] for o in group.get("options", []) if o.get("urn")}
return {}
def fetch_time_periods(dataset_id: str) -> list[str]:
resp = requests.get(f"{API_BASE}/data-sets/{dataset_id}/meta", timeout=TIMEOUT)
resp.raise_for_status()
return [t["period"] for t in resp.json().get("timePeriods", [])]
class DestinationsStream(Stream):
"""One stream per dataset. Pages the query API, one time period at a time."""
_dataset_id: str
_destination_slugs: dict[str, str]
_pupil_group_slugs: dict[str, str]
_pinned: list[str]
schema = th.PropertiesList(
th.Property("urn", th.StringType),
th.Property("time_period", th.StringType),
th.Property("pupil_group", th.StringType),
th.Property("destination_measure", th.StringType),
th.Property("cohort_pupils", th.StringType),
th.Property("pupils_raw", th.StringType),
th.Property("percentage_raw", th.StringType),
).to_dict()
primary_keys = ["urn", "time_period", "pupil_group", "destination_measure"]
replication_key = None
def get_records(self, context):
urn_by_location = fetch_urn_by_location(self._dataset_id)
self.logger.info("%s: %d school locations", self.name, len(urn_by_location))
criteria = [{"filters": {"in": list(self._destination_slugs)}},
{"filters": {"in": list(self._pupil_group_slugs)}}]
for pinned in self._pinned:
criteria.append({"filters": {"in": [pinned]}})
for period in fetch_time_periods(self._dataset_id):
page = 1
while True:
body = {
"criteria": {"and": criteria + [
{"timePeriods": {"in": [{"period": period, "code": "AY"}]}},
]},
"indicators": [IND_COHORT, IND_PUPILS, IND_PERCENT],
"page": page,
"pageSize": PAGE_SIZE,
}
resp = requests.post(
f"{API_BASE}/data-sets/{self._dataset_id}/query",
json=body, timeout=TIMEOUT,
)
resp.raise_for_status()
payload = resp.json()
for row in payload.get("results", []):
record = row_to_record(
row, urn_by_location,
self._destination_slugs, self._pupil_group_slugs,
)
if record is not None:
yield record
paging = payload.get("paging", {})
if page >= paging.get("totalPages", 1):
break
page += 1
class KS4DestinationsStream(DestinationsStream):
name = "ees_ks4_destinations"
_dataset_id = KS4_DATASET
_destination_slugs = KS4_DESTINATION_SLUGS
_pupil_group_slugs = KS4_PUPIL_GROUP_SLUGS
_pinned = KS4_PINNED
class KS5DestinationsStream(DestinationsStream):
name = "ees_ks5_destinations"
_dataset_id = KS5_DATASET
_destination_slugs = KS5_DESTINATION_SLUGS
_pupil_group_slugs = KS5_PUPIL_GROUP_SLUGS
_pinned = KS5_PINNED
class TapUKEESDestinations(Tap):
name = "tap-uk-ees-destinations"
config_jsonschema = th.PropertiesList().to_dict()
def discover_streams(self):
return [KS4DestinationsStream(self), KS5DestinationsStream(self)]
if __name__ == "__main__":
TapUKEESDestinations.cli()
```
- [ ] **Step 4: Run test to verify it passes**
Run: `cd pipeline/plugins/extractors/tap-uk-ees-destinations && uv run --with singer-sdk --with requests --with pytest pytest tests/ -v`
Expected: PASS, 4 tests
- [ ] **Step 5: Register the tap with Meltano**
In `pipeline/meltano.yml`, after the `tap-uk-ees` block in `extractors:`:
```yaml
- name: tap-uk-ees-destinations
namespace: uk_ees_destinations
pip_url: ./plugins/extractors/tap-uk-ees-destinations
executable: tap-uk-ees-destinations
settings: []
```
- [ ] **Step 6: Wire it into the annual DAG**
In `pipeline/dags/school_data_pipeline.py`, inside the `extract_ees` TaskGroup (around line 174), add a second operator and make the dbt selector cover the new models:
```python
extract_ees_destinations = BashOperator(
task_id="extract_ees_destinations",
bash_command=f"cd {PIPELINE_DIR} && {MELTANO_BIN} run tap-uk-ees-destinations target-postgres",
)
extract_ees >> extract_ees_destinations
```
And extend the `dbt_build_ees` selector with `stg_ees_ks4_destinations+ stg_ees_ks5_destinations+`.
- [ ] **Step 7: Commit**
```bash
git add pipeline/plugins/extractors/tap-uk-ees-destinations pipeline/meltano.yml pipeline/dags/school_data_pipeline.py
git commit -m "feat(destinations): a tap that preserves the suppression sentinel"
```
---
### Task 4: Staging models
**Files:**
- Create: `pipeline/transform/models/staging/stg_ees_ks4_destinations.sql`
- Create: `pipeline/transform/models/staging/stg_ees_ks5_destinations.sql`
- Modify: `pipeline/transform/models/staging/_stg_sources.yml`
**Interfaces:**
- Consumes: raw tables from Task 3
- Produces: `stg_ees_ks4_destinations` / `stg_ees_ks5_destinations` with columns `urn` (int), `year` (int), `pupil_group` (text), `destination_measure` (text), `cohort_pupils` (int), `pupils` (int, null when withheld), `percentage` (numeric, null when withheld), `status` (text: `published` | `suppressed` | `not_applicable`)
- [ ] **Step 1: Declare the sources**
In `pipeline/transform/models/staging/_stg_sources.yml`, under `tables:`:
```yaml
- name: ees_ks4_destinations
description: >
KS4 leavers destinations, school level, long format — one row per
URN × year × pupil group × destination measure. pupils_raw and
percentage_raw are TEXT and may hold the 'c' suppression sentinel;
they must never be passed through safe_numeric.
- name: ees_ks5_destinations
description: >
16-18 study leavers destinations, same grain and same suppression
caveat as ees_ks4_destinations.
```
- [ ] **Step 2: Write the KS4 staging model**
Create `pipeline/transform/models/staging/stg_ees_ks4_destinations.sql`:
```sql
{{ config(materialized='table') }}
-- Staging model: KS4 leavers destinations, school level.
--
-- DELIBERATELY DOES NOT USE safe_numeric. That macro maps every EES sentinel
-- (z, c, x, q, u) to NULL, which is right for attainment — there, "suppressed"
-- and "not applicable" are equally unrenderable. Here they are different
-- claims: one prints "withheld", the other prints nothing. Collapsing them
-- would also let a downstream sum reconstruct a withheld figure, because the
-- destination categories add up to the cohort.
--
-- See docs/superpowers/specs/2026-08-28-destination-measures-design.md.
with source as (
select * from {{ source('raw', 'ees_ks4_destinations') }}
-- National rows carry a null urn and feed fact_destination_national.
where (urn is null or urn ~ '^[0-9]+$')
and time_period ~ '^[0-9]+$'
)
select
case when urn ~ '^[0-9]+$' then cast(trim(urn) as integer) end as urn,
cast(trim(time_period) as integer) as year,
trim(pupil_group) as pupil_group,
trim(destination_measure) as destination_measure,
case when cohort_pupils ~ '^[0-9]+$'
then cast(cohort_pupils as integer) end as cohort_pupils,
case when pupils_raw ~ '^[0-9]+$'
then cast(pupils_raw as integer) end as pupils,
case when percentage_raw ~ '^-?[0-9]+(\.[0-9]+)?$'
then cast(percentage_raw as numeric) end as percentage,
case
when pupils_raw ~ '^[0-9]+$' then 'published'
when lower(trim(pupils_raw)) = 'c' then 'suppressed'
else 'not_applicable'
end as status
from source
```
- [ ] **Step 3: Write the 16-18 staging model**
Create `pipeline/transform/models/staging/stg_ees_ks5_destinations.sql` — identical body, reading `source('raw', 'ees_ks5_destinations')`, with the header comment naming 16-18 study leavers. Repeat the full SQL rather than abstracting it; the two sources drift independently and a shared macro would couple their refresh cadences.
- [ ] **Step 4: Verify the SQL compiles by eye against the sibling models**
Run: `diff <(sed -n '1,12p' pipeline/transform/models/staging/stg_ees_ks4.sql) <(sed -n '1,12p' pipeline/transform/models/staging/stg_ees_ks4_destinations.sql)`
Confirm the header comment style matches, and confirm by inspection that `safe_numeric` appears nowhere:
Run: `grep -c safe_numeric pipeline/transform/models/staging/stg_ees_ks*_destinations.sql`
Expected: `0` for both files
- [ ] **Step 5: Commit**
```bash
git add pipeline/transform/models/staging/stg_ees_ks4_destinations.sql \
pipeline/transform/models/staging/stg_ees_ks5_destinations.sql \
pipeline/transform/models/staging/_stg_sources.yml
git commit -m "feat(destinations): staging models that keep 'withheld' distinct from 'absent'"
```
---
### Task 5: Marts and the disclosure tests
**Files:**
- Create: `pipeline/transform/models/marts/fact_ks4_destinations.sql`
- Create: `pipeline/transform/models/marts/fact_ks5_destinations.sql`
- Create: `pipeline/transform/models/marts/fact_destination_national.sql`
- Create: `pipeline/transform/tests/assert_destinations_no_derived_remainder.sql`
- Create: `pipeline/transform/tests/assert_destinations_group_masking.sql`
- Create: `pipeline/transform/tests/assert_destination_status_null_agreement.sql`
- Modify: `pipeline/transform/models/marts/_marts_schema.yml`
**Interfaces:**
- Consumes: `stg_ees_ks4_destinations`, `stg_ees_ks5_destinations`, `dim_school`
- Produces: `fact_ks4_destinations` / `fact_ks5_destinations` (`urn`, `year`, `pupil_group`, `destination_measure`, `cohort_pupils`, `pupils`, `percentage`, `status`) and `fact_destination_national` (same, without `urn`, plus `phase`)
- [ ] **Step 1: Write the KS4 mart with R3 masking**
Create `pipeline/transform/models/marts/fact_ks4_destinations.sql`:
```sql
{{ config(materialized='table') }}
-- Mart: KS4 leavers destinations — one row per URN × year × pupil group ×
-- destination measure.
--
-- Long format, unlike the wide fact_ks4_performance next door. pupil_group is
-- a real third dimension, so going wide would need three sets of every column,
-- and the disclosure tests below are far easier to write over rows.
--
-- R3 is applied HERE rather than in the API: where a category is suppressed
-- for the disadvantaged group it is masked for the other-pupils group too,
-- because the two partition the whole and the all-pupils figure is published.
-- DfE already does this in 493 of 498 cases; this closes the remainder so no
-- consumer can reach an unmasked combination.
with staged as (
select s.*
from {{ ref('stg_ees_ks4_destinations') }} s
inner join {{ ref('dim_school') }} d on d.urn = s.urn
),
-- Categories withheld for disadvantaged pupils at this school and year.
masked as (
select distinct urn, year, destination_measure
from staged
where pupil_group = 'disadvantaged' and status = 'suppressed'
)
select
s.urn,
s.year,
s.pupil_group,
s.destination_measure,
s.cohort_pupils,
case when m.urn is not null and s.pupil_group = 'other'
then null else s.pupils end as pupils,
case when m.urn is not null and s.pupil_group = 'other'
then null else s.percentage end as percentage,
case when m.urn is not null and s.pupil_group = 'other'
then 'suppressed' else s.status end as status
from staged s
left join masked m
on m.urn = s.urn
and m.year = s.year
and m.destination_measure = s.destination_measure
```
- [ ] **Step 2: Write the 16-18 mart**
Create `pipeline/transform/models/marts/fact_ks5_destinations.sql` — the same body reading `stg_ees_ks5_destinations`, with a header naming 16-18 study leavers.
- [ ] **Step 3: Write the national reference mart**
Create `pipeline/transform/models/marts/fact_destination_national.sql`:
```sql
{{ config(materialized='table') }}
-- Mart: England destination measures by pupil group, for the page's national
-- reference. Kept separate from the school facts so the section's England bar
-- can repoint with the cohort switch — comparing a school's disadvantaged
-- pupils against the national all-pupils figure would flatter or damn the
-- school for its intake rather than its work.
select 'ks4' as phase, year, pupil_group, destination_measure,
cohort_pupils, pupils, percentage, status
from {{ ref('stg_ees_ks4_destinations') }}
where urn is null
union all
select 'ks5' as phase, year, pupil_group, destination_measure,
cohort_pupils, pupils, percentage, status
from {{ ref('stg_ees_ks5_destinations') }}
where urn is null
```
- [ ] **Step 4: Write the R1 disclosure test**
Create `pipeline/transform/tests/assert_destinations_no_derived_remainder.sql`:
```sql
-- R1 GUARD. Fails if a school/year/group has exactly one suppressed category
-- while also publishing the cohort total — the combination that lets the
-- withheld figure be recovered by subtraction.
--
-- This does not mean the mart is wrong: DfE publishes exactly this, and the
-- mart's job is to carry it faithfully. The test exists so that the condition
-- is visible and counted, and so that any consumer added later has to
-- acknowledge it. The API and the frontend are what must refuse to render the
-- remainder; this test is the tripwire that says how often the situation
-- arises. It is configured to warn, not error.
{{ config(severity='warn') }}
select
urn, year, pupil_group,
count(*) filter (where status = 'suppressed') as suppressed_categories
from {{ ref('fact_ks4_destinations') }}
where destination_measure not like 'agg_%'
group by urn, year, pupil_group
having count(*) filter (where status = 'suppressed') = 1
```
- [ ] **Step 5: Write the R3 masking test**
Create `pipeline/transform/tests/assert_destinations_group_masking.sql`:
```sql
-- R3 GUARD. Fails if a category is suppressed for disadvantaged pupils but
-- still published for the other-pupils group — the two partition the whole, so
-- publishing both alongside the all-pupils figure recovers the withheld cell.
select d.urn, d.year, d.destination_measure
from {{ ref('fact_ks4_destinations') }} d
inner join {{ ref('fact_ks4_destinations') }} o
on o.urn = d.urn
and o.year = d.year
and o.destination_measure = d.destination_measure
and o.pupil_group = 'other'
where d.pupil_group = 'disadvantaged'
and d.status = 'suppressed'
and o.status = 'published'
```
- [ ] **Step 6: Write the status agreement test**
Create `pipeline/transform/tests/assert_destination_status_null_agreement.sql`:
```sql
-- pupils must be null wherever status is not 'published', and never null where
-- it is. This is what stops a later coalesce or a wide-format refactor turning
-- "withheld" into a zero.
select urn, year, pupil_group, destination_measure, status, pupils
from {{ ref('fact_ks4_destinations') }}
where (status <> 'published' and pupils is not null)
or (status = 'published' and pupils is null)
```
- [ ] **Step 7: Document the marts**
In `pipeline/transform/models/marts/_marts_schema.yml`, add entries for the three new models with a `description` for each and `tests: [not_null]` on `urn`, `year`, `pupil_group`, `destination_measure`, `status`, following the existing entries' shape.
- [ ] **Step 8: Confirm no test references safe_numeric and commit**
Run: `grep -rn safe_numeric pipeline/transform/models/marts/fact_ks*_destinations.sql pipeline/transform/models/marts/fact_destination_national.sql`
Expected: no output
```bash
git add pipeline/transform/models/marts/fact_ks4_destinations.sql \
pipeline/transform/models/marts/fact_ks5_destinations.sql \
pipeline/transform/models/marts/fact_destination_national.sql \
pipeline/transform/tests/assert_destinations_*.sql \
pipeline/transform/tests/assert_destination_status_null_agreement.sql \
pipeline/transform/models/marts/_marts_schema.yml
git commit -m "feat(destinations): marts, with R3 masking applied at the boundary"
```
---
### Task 6: Backend models, loader and API
**Files:**
- Modify: `backend/models.py` (append after `FactFinance`, around line 265)
- Modify: `backend/data_loader.py:819-830` (`_empty_supplementary`) and `:833-958` (`get_supplementary_data_batch`)
- Modify: `backend/app.py:905-920` (the `get_school_details` return block)
- Test: `backend/tests/test_destinations_api.py`
**Interfaces:**
- Consumes: `fact_ks4_destinations`, `fact_ks5_destinations`, `fact_destination_national`
- Produces: `destinations` key on `GET /api/schools/{urn}`, shaped `{ ks4: {...} | null, ks5: {...} | null }`; each phase `{ cohort_year, groups: { all, disadvantaged, other } }`; each group `{ cohort, categories: [{ category, pupils, percentage, status }], aggregates: {...} }`
- [ ] **Step 1: Write the failing test**
Create `backend/tests/test_destinations_api.py`:
```python
"""The serialiser's contract: it carries suppression through, and never emits a
total that closes a gap left by a suppressed category."""
import pytest
from backend.data_loader import _destinations_block
def _row(group, measure, pupils, status, cohort=180, percentage=None):
return {
"pupil_group": group, "destination_measure": measure,
"pupils": pupils, "percentage": percentage,
"status": status, "cohort_pupils": cohort, "year": 202223,
}
def test_suppressed_category_serialises_as_suppressed_with_null_pupils():
rows = [
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
_row("all", "sixth_form_college", None, "suppressed"),
]
block = _destinations_block(rows)
cats = {c["category"]: c for c in block["groups"]["all"]["categories"]}
assert cats["sixth_form_college"]["status"] == "suppressed"
assert cats["sixth_form_college"]["pupils"] is None
assert cats["sixth_form_college"]["percentage"] is None
def test_no_closing_total_is_emitted_for_a_partially_suppressed_group():
rows = [
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
_row("all", "sixth_form_college", None, "suppressed"),
_row("all", "further_education", 61, "published", percentage=33.9),
_row("all", "apprenticeship", 8, "published", percentage=4.4),
_row("all", "employment", 6, "published", percentage=3.3),
_row("all", "not_sustained", 5, "published", percentage=2.8),
_row("all", "not_captured", 4, "published", percentage=2.2),
]
block = _destinations_block(rows)
group = block["groups"]["all"]
published = sum(c["pupils"] for c in group["categories"] if c["pupils"] is not None)
for value in group["aggregates"].values():
if value is None or value.get("pupils") is None:
continue
assert value["pupils"] != group["cohort"] - published, (
"an aggregate that equals the residual identifies the suppressed cell"
)
def test_cohort_year_is_reported_so_the_page_can_date_itself():
block = _destinations_block([_row("all", "school_sixth_form", 75, "published")])
assert block["cohort_year"] == "2022/23"
def test_empty_rows_yield_none_not_an_empty_shell():
assert _destinations_block([]) is None
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd /Users/tudor/projects/school_compare && uv run --with fastapi --with 'httpx==0.27.0' --with sqlalchemy --with pandas --with pytest --with pydantic-settings python -m pytest backend/tests/test_destinations_api.py -v`
Expected: FAIL — `ImportError: cannot import name '_destinations_block'`
- [ ] **Step 3: Add the SQLAlchemy models**
Append to `backend/models.py` after `FactFinance`:
```python
class FactKs4Destinations(Base):
"""KS4 leavers destinations — one row per URN, year, pupil group, measure."""
__tablename__ = "fact_ks4_destinations"
__table_args__ = (
Index("ix_ks4_dest_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
pupil_group = Column(String(20), primary_key=True)
destination_measure = Column(String(40), primary_key=True)
cohort_pupils = Column(Integer)
pupils = Column(Integer)
percentage = Column(Float)
# 'published' | 'suppressed' | 'not_applicable'. Never collapse this to a
# null check: a suppressed cell prints "withheld", an absent one prints
# nothing, and the difference is what keeps the disclosure rules workable.
status = Column(String(20))
class FactKs5Destinations(Base):
"""16-18 study leavers destinations — same grain as FactKs4Destinations."""
__tablename__ = "fact_ks5_destinations"
__table_args__ = (
Index("ix_ks5_dest_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
pupil_group = Column(String(20), primary_key=True)
destination_measure = Column(String(40), primary_key=True)
cohort_pupils = Column(Integer)
pupils = Column(Integer)
percentage = Column(Float)
status = Column(String(20))
```
- [ ] **Step 4: Write the serialiser**
Add to `backend/data_loader.py`, above `_empty_supplementary`:
```python
_AGGREGATE_MEASURES = {"agg_sustained_education", "agg_sustained_all"}
def _format_cohort_year(year: int | None) -> str | None:
"""202223 -> '2022/23'. The page must date its own cohort: destinations run
two GCSE years behind the results shown above them."""
if not year:
return None
text = str(year)
return f"{text[:4]}/{text[6:8]}" if len(text) == 8 else f"{text[:4]}/{text[4:6]}"
def _destinations_block(rows: list[dict]) -> dict | None:
"""Shape destination rows for one phase into the API's block.
Carries `status` through untouched and emits no computed totals. The only
aggregates present are ones DfE published itself; the frontend decides
whether they are safe to show (see lib/destinations.ts, R2).
"""
if not rows:
return None
latest_year = max(r["year"] for r in rows if r.get("year") is not None)
rows = [r for r in rows if r.get("year") == latest_year]
groups: dict[str, dict] = {}
for row in rows:
group = groups.setdefault(
row["pupil_group"],
{"cohort": row.get("cohort_pupils"), "categories": [], "aggregates": {}},
)
measure = row["destination_measure"]
cell = {
"category": measure,
"pupils": row.get("pupils"),
"percentage": row.get("percentage"),
"status": row.get("status"),
}
if measure in _AGGREGATE_MEASURES:
group["aggregates"][measure.removeprefix("agg_")] = cell
else:
group["categories"].append(cell)
if not groups:
return None
return {"cohort_year": _format_cohort_year(latest_year), "groups": groups}
```
- [ ] **Step 5: Run test to verify it passes**
Run: `cd /Users/tudor/projects/school_compare && uv run --with fastapi --with 'httpx==0.27.0' --with sqlalchemy --with pandas --with pytest --with pydantic-settings python -m pytest backend/tests/test_destinations_api.py -v`
Expected: PASS, 4 tests
- [ ] **Step 6: Wire it into the batch loader**
In `backend/data_loader.py`, add `"destinations": None` to the dict `_empty_supplementary` returns. Then add two query functions inside `get_supplementary_data_batch`, following the `_ofsted` / `_census` pattern exactly, each wrapped in `_safe`:
```python
# Destinations — KS4 and 16-18, all years; _destinations_block picks the
# latest and shapes the groups.
def _destinations():
from collections import defaultdict
per_urn_ks4 = defaultdict(list)
for r in (db.query(FactKs4Destinations)
.filter(FactKs4Destinations.urn.in_(urns)).all()):
per_urn_ks4[r.urn].append({
"year": r.year, "pupil_group": r.pupil_group,
"destination_measure": r.destination_measure,
"cohort_pupils": r.cohort_pupils, "pupils": r.pupils,
"percentage": r.percentage, "status": r.status,
})
per_urn_ks5 = defaultdict(list)
for r in (db.query(FactKs5Destinations)
.filter(FactKs5Destinations.urn.in_(urns)).all()):
per_urn_ks5[r.urn].append({
"year": r.year, "pupil_group": r.pupil_group,
"destination_measure": r.destination_measure,
"cohort_pupils": r.cohort_pupils, "pupils": r.pupils,
"percentage": r.percentage, "status": r.status,
})
for urn in urns:
ks4 = _destinations_block(per_urn_ks4.get(urn, []))
ks5 = _destinations_block(per_urn_ks5.get(urn, []))
result[urn]["destinations"] = (
{"ks4": ks4, "ks5": ks5} if (ks4 or ks5) else None
)
_safe(_destinations)
```
Import `FactKs4Destinations` and `FactKs5Destinations` alongside the other mart models at the top of the file.
- [ ] **Step 7: Expose it on the endpoint**
In `backend/app.py`, in the `get_school_details` return dict, after `"finance": supplementary.get("finance"),`:
```python
"destinations": supplementary.get("destinations"),
```
- [ ] **Step 8: Run the full backend suite and commit**
Run: `cd /Users/tudor/projects/school_compare && uv run --with fastapi --with 'httpx==0.27.0' --with sqlalchemy --with pandas --with pytest --with pydantic-settings python -m pytest backend/tests/ -q`
Expected: all pass
```bash
git add backend/models.py backend/data_loader.py backend/app.py backend/tests/test_destinations_api.py
git commit -m "feat(destinations): serve destinations without closing the gaps"
```
---
### Task 7: Frontend types and section flags
**Files:**
- Modify: `nextjs-app/lib/types.ts`
- Modify: `nextjs-app/lib/schoolSections.ts:12-46` (the interfaces) and `:48-95` (`computeSchoolFlags`)
- Test: `nextjs-app/__tests__/lib/schoolSections.destinations.test.ts`
**Interfaces:**
- Consumes: `lib/destinations.ts` types from Task 1
- Produces: `SchoolDestinations` type; `SchoolFlags.hasKs4Destinations` and `.hasKs5Destinations`
- [ ] **Step 1: Write the failing test**
Create `nextjs-app/__tests__/lib/schoolSections.destinations.test.ts`:
```ts
import { computeSchoolFlags } from '@/lib/schoolSections';
const base = {
schoolInfo: { urn: 1, school_name: 'X', phase: 'Secondary', has_sixth_form: true } as any,
yearlyData: [], absenceData: null, census: null, deprivation: null, finance: null,
};
const ks4Only = {
ks4: { cohort_year: '2022/23', groups: { all: { cohort: 180, categories: [], aggregates: {} } } },
ks5: null,
} as any;
it('flags KS4 destinations when the block is present', () => {
const flags = computeSchoolFlags({ ...base, destinations: ks4Only });
expect(flags.hasKs4Destinations).toBe(true);
expect(flags.hasKs5Destinations).toBe(false);
});
it('flags neither when the block is absent', () => {
const flags = computeSchoolFlags({ ...base, destinations: null });
expect(flags.hasKs4Destinations).toBe(false);
expect(flags.hasKs5Destinations).toBe(false);
});
it('does not flag a group with no categories as renderable', () => {
const empty = { ks4: { cohort_year: '2022/23', groups: {} }, ks5: null } as any;
expect(computeSchoolFlags({ ...base, destinations: empty }).hasKs4Destinations).toBe(false);
});
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd nextjs-app && npx jest __tests__/lib/schoolSections.destinations.test.ts`
Expected: FAIL — `hasKs4Destinations` is undefined
- [ ] **Step 3: Add the types**
In `nextjs-app/lib/types.ts`:
```ts
import type { DestinationCell, PupilGroup } from './destinations';
export interface DestinationGroupPayload {
cohort: number | null;
categories: DestinationCell[];
aggregates: Partial<Record<'sustained_education' | 'sustained_all', DestinationCell>>;
}
export interface DestinationPhase {
cohort_year: string | null;
groups: Partial<Record<PupilGroup, DestinationGroupPayload>>;
}
export interface SchoolDestinations {
ks4: DestinationPhase | null;
ks5: DestinationPhase | null;
}
```
- [ ] **Step 4: Extend the flags**
In `nextjs-app/lib/schoolSections.ts`, add `destinations: SchoolDestinations | null` to `SchoolFlagsInput`, add `hasKs4Destinations: boolean` and `hasKs5Destinations: boolean` to `SchoolFlags`, destructure `destinations` in `computeSchoolFlags`, and compute:
```ts
// A phase counts as present only if some group actually carries categories —
// a block with an empty groups map is a pipeline artefact, not a section.
const phaseHasContent = (phase: DestinationPhase | null | undefined) =>
!!phase && Object.values(phase.groups ?? {}).some(g => (g?.categories?.length ?? 0) > 0);
const hasKs4Destinations = phaseHasContent(destinations?.ks4);
const hasKs5Destinations = phaseHasContent(destinations?.ks5);
```
Return both from `computeSchoolFlags`. Add nav items `{ id: 'destinations', label: 'After Year 11' }` and `{ id: 'post16-destinations', label: 'After the sixth form' }` in `buildNavItems`, gated on the two flags, positioned after the GCSE entry.
- [ ] **Step 5: Run test and typecheck**
Run: `cd nextjs-app && npx jest __tests__/lib/schoolSections.destinations.test.ts && npm run typecheck`
Expected: PASS
- [ ] **Step 6: Commit**
```bash
git add nextjs-app/lib/types.ts nextjs-app/lib/schoolSections.ts nextjs-app/__tests__/lib/schoolSections.destinations.test.ts
git commit -m "feat(destinations): types and section flags"
```
---
### Task 8: The After Year 11 section
**Files:**
- Create: `nextjs-app/components/school/DestinationsSection.tsx`
- Create: `nextjs-app/components/school/DestinationsView.tsx`
- Create: `nextjs-app/components/school/destinations.module.css`
- Modify: `nextjs-app/components/school/SecondarySchoolSections.tsx`
- Test: `nextjs-app/__tests__/components/DestinationsSection.test.tsx`
**Interfaces:**
- Consumes: `lib/destinations.ts` (Task 1), `SchoolDestinations` (Task 7), tokens (Task 2)
- Produces: `<DestinationsSection destinations={phase} schoolName={string} />`
- [ ] **Step 1: Write the failing test**
Create `nextjs-app/__tests__/components/DestinationsSection.test.tsx`:
```tsx
import { render, screen } from '@testing-library/react';
import { DestinationsSection } from '@/components/school/DestinationsSection';
const cell = (category: string, pupils: number | null, status = 'published') => ({
category, pupils, percentage: pupils === null ? null : (pupils / 180) * 100, status,
});
const fullPhase: any = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: [
cell('school_sixth_form', 75), cell('sixth_form_college', 21),
cell('further_education', 55), cell('other_education', 6),
cell('apprenticeship', 8), cell('employment', 6),
cell('not_sustained', 5), cell('not_captured', 4),
],
aggregates: {},
},
},
};
const suppressedPhase: any = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: [
cell('school_sixth_form', 75), cell('sixth_form_college', null, 'suppressed'),
cell('further_education', 55), cell('other_education', 6),
cell('apprenticeship', 8), cell('employment', 6),
cell('not_sustained', 5), cell('not_captured', 4),
],
aggregates: {},
},
},
};
it('dates its own cohort so it is not read as stale', () => {
render(<DestinationsSection destinations={fullPhase} schoolName="Northbrook Academy" />);
expect(screen.getByText(/2022\/23/)).toBeInTheDocument();
});
it('renders the bar when the group is fully published', () => {
const { container } = render(<DestinationsSection destinations={fullPhase} schoolName="X" />);
expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(8);
});
it('renders NO bar when a category is withheld', () => {
const { container } = render(<DestinationsSection destinations={suppressedPhase} schoolName="X" />);
expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(0);
expect(screen.getByText(/withheld/i)).toBeInTheDocument();
});
it('never states a remainder for a partially suppressed group', () => {
const { container } = render(<DestinationsSection destinations={suppressedPhase} schoolName="X" />);
// 180 cohort - 155 published = 25, the withheld figure. It must appear nowhere.
expect(container.textContent).not.toMatch(/\b25\b/);
});
it('never claims a pupil stayed at this school', () => {
const { container } = render(<DestinationsSection destinations={fullPhase} schoolName="Northbrook Academy" />);
expect(container.textContent).not.toMatch(/stayed on (here|at)/i);
});
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd nextjs-app && npx jest __tests__/components/DestinationsSection.test.tsx`
Expected: FAIL — cannot find module
- [ ] **Step 3: Write the server section**
Create `nextjs-app/components/school/DestinationsSection.tsx`. It renders the `Section` shell, the title, a subtitle naming the cohort year and the publication lag, and delegates the interactive body to `DestinationsView` with `all` as the server-rendered default. Mark each bar segment with `data-destination-segment` so the tests and the E2E journeys can assert on its absence.
Key structure:
```tsx
/**
* DestinationsSection — where a school's Year 11 leavers went. Server component.
*
* The headline is deliberately NOT the sustained-destination rate: that figure
* sits between 92% and 97% for nearly every school in England, so leading with
* it would say nothing. The mix is what varies.
*/
import type { DestinationPhase } from '@/lib/types';
import { Section, sectionStyles } from './sectionShared';
import { DestinationsView } from './DestinationsView';
export function DestinationsSection({
destinations, schoolName,
}: { destinations: DestinationPhase; schoolName: string }) {
const cohort = destinations.groups.all?.cohort ?? null;
return (
<Section id="destinations">
<h2 className={sectionStyles.sectionTitle}>After Year 11</h2>
<p className={sectionStyles.sectionSubtitle}>
Where {cohort ? `the ${cohort} pupils` : 'the pupils'} who left Year 11 in{' '}
{destinations.cohort_year ?? 'the most recent year published'} went next.
Destination measures are published about two years after the exams above.
</p>
<DestinationsView destinations={destinations} schoolName={schoolName} />
</Section>
);
}
```
- [ ] **Step 4: Write the client view**
Create `nextjs-app/components/school/DestinationsView.tsx` with `'use client'`. It owns the cohort switch (`role="radiogroup"`, arrow-key navigation), the card↔bar hover linkage, and the render decisions:
- Cards from `CARD_GROUPS` — `aggregateCells` for the value, or a "Not published" card when it returns `null`
- Bar only when `canRenderBar(group)`; otherwise a panel explaining that the categories add up to the cohort so the rest cannot be drawn
- A published aggregate is shown only when `canRenderPublishedAggregate(components)` is true
- The full table always, with withheld rows marked
- Segment widths from `toBarSegments`, each carrying `data-destination-segment`
Wrap the `toBarSegments` call in the `canRenderBar` guard rather than a try/catch — the throw is a backstop for programmer error, not control flow.
- [ ] **Step 5: Write the stylesheet**
Create `nextjs-app/components/school/destinations.module.css` using only the tokens from Task 2 plus the existing section tokens. The absence segment is `background-image: repeating-linear-gradient(45deg, var(--dest-none-hatch) 0 3px, transparent 3px 7px)` over `var(--bg-card)` with a `1px` inset ring in `var(--dest-none)`. Segments sit in a flex row with `gap: 2px`.
- [ ] **Step 6: Mount it on the secondary template**
In `nextjs-app/components/school/SecondarySchoolSections.tsx`, render `<DestinationsSection />` after the GCSE section and before admissions, gated on `flags.hasKs4Destinations`.
- [ ] **Step 7: Run tests, typecheck, commit**
Run: `cd nextjs-app && npx jest __tests__/components/DestinationsSection.test.tsx && npm run typecheck`
Expected: PASS, 5 tests
```bash
git add nextjs-app/components/school/DestinationsSection.tsx \
nextjs-app/components/school/DestinationsView.tsx \
nextjs-app/components/school/destinations.module.css \
nextjs-app/components/school/SecondarySchoolSections.tsx \
nextjs-app/__tests__/components/DestinationsSection.test.tsx
git commit -m "feat(destinations): the After Year 11 section"
```
---
### Task 9: The post-16 section, and removing the placeholder
**Files:**
- Create: `nextjs-app/components/school/Post16DestinationsSection.tsx`
- Modify: `nextjs-app/components/school/SecondaryAdmissionsSection.tsx:110-120`
- Modify: `nextjs-app/components/school/SecondarySchoolSections.tsx`
- Test: `nextjs-app/__tests__/components/Post16DestinationsSection.test.tsx`
**Interfaces:**
- Consumes: everything from Task 8; reuses `DestinationsView`
- Produces: `<Post16DestinationsSection destinations={phase} schoolName={string} />`
- [ ] **Step 1: Write the failing test**
Create `nextjs-app/__tests__/components/Post16DestinationsSection.test.tsx`:
```tsx
import { render, screen } from '@testing-library/react';
import { Post16DestinationsSection } from '@/components/school/Post16DestinationsSection';
const phase: any = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 96,
categories: [
{ category: 'higher_education', pupils: 56, percentage: 58.3, status: 'published' },
{ category: 'further_education', pupils: 12, percentage: 12.5, status: 'published' },
{ category: 'apprenticeship', pupils: 9, percentage: 9.4, status: 'published' },
{ category: 'employment', pupils: 13, percentage: 13.5, status: 'published' },
{ category: 'not_sustained', pupils: 6, percentage: 6.3, status: 'published' },
],
aggregates: {},
},
},
};
it('names the Year 13 cohort, not Year 11', () => {
render(<Post16DestinationsSection destinations={phase} schoolName="X" />);
expect(screen.getByText(/Year 13/)).toBeInTheDocument();
});
it('reports university destinations', () => {
render(<Post16DestinationsSection destinations={phase} schoolName="X" />);
expect(screen.getByText(/higher education|university/i)).toBeInTheDocument();
});
```
- [ ] **Step 2: Run test to verify it fails**
Run: `cd nextjs-app && npx jest __tests__/components/Post16DestinationsSection.test.tsx`
Expected: FAIL — cannot find module
- [ ] **Step 3: Write the section**
Create `nextjs-app/components/school/Post16DestinationsSection.tsx` mirroring `DestinationsSection` with `id="post16-destinations"`, the heading "After the sixth form", copy naming Year 13, and the same `DestinationsView` body.
- [ ] **Step 4: Remove the placeholder**
In `nextjs-app/components/school/SecondaryAdmissionsSection.tsx`, delete the "Post-16 destination data coming soon" paragraph at line ~117 and its surrounding conditional. The sixth-form badge in the header stays.
- [ ] **Step 5: Mount it**
In `SecondarySchoolSections.tsx`, render `<Post16DestinationsSection />` after `<DestinationsSection />`, gated on `flags.hasKs5Destinations`. Where the school has no sixth form the section is simply not rendered — no placeholder, because absence is the correct statement.
- [ ] **Step 6: Confirm the placeholder is gone, run tests, commit**
Run: `grep -rn "coming soon" nextjs-app/components/`
Expected: no output
Run: `cd nextjs-app && npx jest && npm run typecheck`
Expected: PASS
```bash
git add nextjs-app/components/school/Post16DestinationsSection.tsx \
nextjs-app/components/school/SecondaryAdmissionsSection.tsx \
nextjs-app/components/school/SecondarySchoolSections.tsx \
nextjs-app/__tests__/components/Post16DestinationsSection.test.tsx
git commit -m "feat(destinations): the post-16 section, replacing the placeholder"
```
---
### Task 10: E2E journeys
Per CLAUDE.md, user-facing behaviour extends `e2e/` in the same PR. Note the staging E2E gate runs post-merge — these journeys cannot pass in PR checks until the Airflow DAG has populated the marts on staging.
**Files:**
- Modify: `e2e/tests/journeys.spec.ts`
**Interfaces:**
- Consumes: the rendered pages from Tasks 8 and 9
- Produces: nothing
- [ ] **Step 1: Write the journeys**
Append to `e2e/tests/journeys.spec.ts`:
```ts
test('a secondary school page says where its Year 11 leavers went', async ({ page }) => {
await page.goto('/school/abbey-grange-church-of-england-academy-137083');
const section = page.locator('#destinations');
await expect(section).toBeVisible();
// The section must date its own cohort — destinations run two GCSE years
// behind the results above them, and an undated figure reads as stale.
await expect(section).toContainText(/20\d{2}\/\d{2}/);
});
test('the destinations section draws no bar for a group with withheld figures', async ({ page }) => {
await page.goto('/school/abbey-grange-church-of-england-academy-137083');
const section = page.locator('#destinations');
await section.getByRole('radio', { name: /disadvantaged/i }).click();
const withheld = section.getByText(/withheld/i);
if (await withheld.count() > 0) {
// R1: where anything is withheld, the bar must be absent entirely — a bar
// with a gap in it publishes the withheld figure by its width.
await expect(section.locator('[data-destination-segment]')).toHaveCount(0);
}
});
test('a school with no sixth form has no post-16 destinations section', async ({ page }) => {
await page.goto('/school/abbey-grange-church-of-england-academy-137083');
const hasSixthForm = await page.getByText(/sixth form/i).count() > 0;
if (!hasSixthForm) {
await expect(page.locator('#post16-destinations')).toHaveCount(0);
}
});
test('the destinations section never claims a pupil stayed at this school', async ({ page }) => {
await page.goto('/school/abbey-grange-church-of-england-academy-137083');
const text = await page.locator('#destinations').textContent();
// The published file reports destination TYPE, never destination institution.
expect(text ?? '').not.toMatch(/stayed on (here|at this school)/i);
});
```
- [ ] **Step 2: Verify the slug resolves**
Run: `grep -n "school/" e2e/tests/journeys.spec.ts | head -5`
Match the slug format the existing school-page journeys use. If they build slugs from an API call rather than hardcoding, follow that pattern instead of the literal above.
- [ ] **Step 3: Commit**
```bash
git add e2e/tests/journeys.spec.ts
git commit -m "test(e2e): destination journeys, including the no-bar rule"
```
---
## Self-Review
**Spec coverage.** Every section of the design maps to a task: disclosure rules → Task 1 (guards) and Task 5 (mart tests); availability/extraction → Task 3; staging and the `safe_numeric` prohibition → Task 4; marts and R3 → Task 5; API → Task 6; display → Tasks 2, 7, 8, 9; edge states → Tasks 8 and 9; testing → every task plus Task 10.
**Gap found and closed.** A first draft had Task 3 drop every non-school row, which would have left `fact_destination_national` (Task 5) reading an empty table — and a note telling the executor to go back and amend an earlier task. Task 3 now keeps national rows with `urn = None` from the start, and its tests cover both that and the LA rows that must still be dropped.
**Type consistency.** `row_to_record` reads `geographicLevel`, not the `locations` keys, because a school row also carries `NAT`, `LA` and `REG` entries for its parents — keying off `"NAT" in locations` would admit every LA row as national. `status` takes the same three values in the tap, the staging models, the marts, the SQLAlchemy models, the API and `lib/destinations.ts`. `pupil_group` is `all` / `disadvantaged` / `other` throughout.