From ac5b7ccd7f95b51f014d21788a4926307fd362e8 Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 10:08:25 +0100 Subject: [PATCH 01/11] docs: design for 2023/24 KS4/KS2 data and DfE LA averages (C2, H2) Co-Authored-By: Claude Opus 5.5 --- ...0-06-ks4-2023-24-and-la-averages-design.md | 300 ++++++++++++++++++ 1 file changed, 300 insertions(+) create mode 100644 docs/superpowers/specs/2026-10-06-ks4-2023-24-and-la-averages-design.md diff --git a/docs/superpowers/specs/2026-10-06-ks4-2023-24-and-la-averages-design.md b/docs/superpowers/specs/2026-10-06-ks4-2023-24-and-la-averages-design.md new file mode 100644 index 0000000..0777acf --- /dev/null +++ b/docs/superpowers/specs/2026-10-06-ks4-2023-24-and-la-averages-design.md @@ -0,0 +1,300 @@ +# 2023/24 Results and DfE LA Averages — Design + +**Date:** 2026-10-06 +**Status:** approved design, not yet implemented +**Scope:** `tap_uk_ees` (KS4 results, KS4 information, new LA stream), +`safe_numeric`, new `fact_ks4_la_averages` mart, annual EES DAG selector, +`/api/la-averages`, secondary search rows and map cards +**Fixes:** audit findings C2 and H2, from the 3 Oct 2026 accuracy audit + +## Goal + +Show every 2023/24 result and school-information figure DfE published, and +compare each state-funded secondary school with DfE's own local-authority +average. + +## The problem + +### C2: 2023/24 is empty + +Bishop Stopford School (137086) shows Attainment 8 60.7 for 2022/23, nothing +for 2023/24 and 58.7 for 2024/25. DfE's 2023/24 figures are Attainment 8 64.1, +Progress 8 +1.02 and English and maths grade 4+ 91.7%. 2023/24 is the last year +DfE published Progress 8 (2024/25 has no KS2 baseline), so the site shows no +recent Progress 8 for any school. Site-wide, 4,170 listed schools have a DfE +2023/24 Attainment 8 and 3,384 a Progress 8. + +There are three separate causes. + +1. **KS4 results.** The 2024/25 release's + `202425_performance_tables_schools_final.csv` is a time series: it holds + 2022/23, 2023/24 and 2024/25 under the current column names, with the right + values (Bishop Stopford 2023/24: 64.1, 1.02, 91.7). The 2023/24 release's own + `202324_performance_tables_schools_final.csv`, re-issued on 10 March 2026, + uses the older names (`t_pupils`, `avg_att8`, `avg_p8score`, + `pt_l2basics_94` …). `EESDatasetStream` reads releases in the API's order, + newest first, so the old file is read last. Its rows carry none of the + declared fields, and target-postgres upserts on the stream's primary key + (`append_only = not key_properties` in meltanolabs-target-postgres 0.8.0), + so they overwrite the good 2023/24 rows with nulls. + + The two files share the keys of every "Total" row, but 34,254 of the 57,090 + 2023/24 sub-group rows use different labels ("Low prior" against "Low prior + attainment"). Those old-label rows sit in `raw.ees_ks4_performance` with + null measures. Nothing reads them. + +2. **KS4 school information.** 2023/24 information exists only in the 2023/24 + release, in `202324_information_about_schools_final.csv`, with the older + names. `EESKS4InfoStream` declares the newer ones, so prior attainment, + SEN percentages, disadvantage gaps and Progress 8 banding are null for + 2023/24. + +3. **KS2 school information.** The 2023/24 file + `ks2_school_information_data.csv` uses the declared names, and pupil counts + load (school 147411: 818 pupils, 112 eligible). Its percentages are written + with a sign (`ptfsm6cla1a = "34%"`). `safe_numeric` accepts only + `^-?[0-9]+(\.[0-9]+)?$`, so disadvantaged, EAL, SEN and mobility percentages + are null for every school in 2023/24. + +DfE published no school-level KS2 information file for 2022/23 (the 2023/24 +release's attainment file carries 2022/23 attainment rows only). Those nulls are +a gap in the source, not a defect. + +### H2: "vs LA avg" uses the wrong average + +`/api/la-averages` (`backend/app.py`) takes an unweighted mean of every school +in the LA with an Attainment 8 score, independent and special schools included. +Independent schools score low because DfE measures exclude IGCSEs, and special +schools score low for other reasons, so the average is too low almost +everywhere. Kensington and Chelsea's "LA avg" is 35.2: the mean of 6 state +schools (54.9) and 8 independent schools (20.4). DfE's figure is 54.5. Of the +151 LAs the audit compared, ours was lower in 147, by 7.1 points on average and +by up to 19.3, so most secondary schools look better than their area. + +DfE's LA averages are already in the file the pipeline downloads for the +England averages: the "summary, all state-funded" data set +(`data-catalogue/data-set/1b649e16-01e8-435b-a814-56be2faf9054/csv`). Its +`Local authority` / `All state-funded` / Total rows match DfE's published +performance-table LA averages (RECTYPE 4) for all 152 LAs in 2024/25, with no +difference. `EESKs4NationalStream` keeps the England row and discards them. +"All state-funded" is the same population as the England benchmark the site +already shows. + +Computing the average ourselves from state-funded schools, weighted by pupils, +was tested and rejected: it differs from DfE's figure by 1.1 points on average +and is never exact. + +## Non-goals + +- 2022/23 KS2 school information (DfE published none). +- An "excludes IGCSEs" note wherever an independent school's Attainment 8 + appears. This change only stops comparing independent schools with the LA. +- Showing 2023/24 Progress 8 in the GCSE section's headline. The history section + shows it once the data loads; the 2024/25 banner stays true. +- Updating the fixed EES data-set id when DfE publishes 2025/26. It already + feeds the England averages; the LA stream shares it. +- LA comparisons for other measures or on other pages. +- Deleting the leftover old-label raw rows automatically. + +## The rules + +### The newest release owns every year it contains (KS4 results only) + +`EESDatasetStream` gets an opt-in class attribute, +`_newest_release_owns_period: bool = False`. `EESKS4PerformanceStream` sets it +to `True`. When it is on: + +- releases are processed newest first by `time_period` (from the release slug), + not in the API's order; +- the stream records each `time_period` it has emitted; +- in each older release, rows whose `time_period` a newer release already + emitted are dropped, and the stream logs how many it skipped and for which + years; +- years only an older release contains are emitted as before. + +The filter is a pure function, testable without a download. Other streams keep +the current behaviour. A general rule would be wrong: the 2024/25 KS2 file holds +98,448 of the 955,956 rows the 2023/24 release has for 2023/24. + +### KS4 information: old names + +`EESKS4InfoStream._column_renames` maps the 2023/24 names onto the declared +fields: + +| 2023/24 column | Declared field | +|---|---| +| `t_allks_pupils` | `allks_pupil_count` | +| `t_allks_boys` | `allks_boys_count` | +| `t_allks_girls` | `allks_girls_count` | +| `t_pupils` | `endks4_pupil_count` | +| `avg_ks2_scaledscore` | `ks2_scaledscore_average` | +| `pt_sen_with_ehcp` | `sen_with_ehcp_pupil_percent` | +| `pt_sen` | `sen_pupil_percent` | +| `pt_sen_no_ehcp` | `sen_no_ehcp_pupil_percent` | +| `diffn_att8` | `attainment8_diffn` | +| `diffn_p8mea` | `progress8_diffn` | +| `p8_banding` | `progress8_banding` | + +Newer files contain none of the old names, so they are unaffected. + +### `safe_numeric` accepts a trailing `%` + +The pattern becomes `^-?[0-9]+(\.[0-9]+)?%?$` and the cast reads +`rtrim(col, '%')`. A value such as `34%` can only mean 34. Suppression codes +(`c`, `z`, `x` …) still become null. + +### LA averages + +A new stream, `ees_ks4_la`, reads the same CSV as `ees_ks4_national` and keeps +rows where `geographic_level = 'Local authority'`, +`establishment_type_group = 'All state-funded'`, `breakdown_topic = 'Total'` and +`breakdown = 'Total'` (case-insensitive, as the national stream compares). It +emits `time_period`, `old_la_code`, `new_la_code`, `la_name` and the 8 headline +measures the national stream emits (`_KS4_NATIONAL_COL_MAP`). Primary key: +(`time_period`, `old_la_code`). The two streams share one download-and-filter +helper. + +`old_la_code` is the GIAS LA code (`local_authority_code`), so schools join on +the code, not the name. Names match today for every LA DfE publishes; DfE +publishes no figure for City of London. + +### Which schools get a gap + +The search row and the map card show "vs LA avg" only when the school has an +Attainment 8 score, is neither special (`isSpecialSchool`) nor independent +(`isIndependentSchool`: "independent" in the GIAS type, which covers "Other +independent school" and "Other independent special school"), and its LA has a +DfE figure for the year the endpoint serves. + +## Delivery + +Two PRs, as for C1. + +### PR 1: pipeline + +- `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py`: release + precedence, KS4 information renames, `ees_ks4_la` stream registered in + `discover_streams`, shared national/LA helper. +- `pipeline/transform/macros/safe_numeric.sql`: trailing `%`. +- `pipeline/transform/models/staging/`: source `raw.ees_ks4_la` and + `stg_ees_ks4_la` (view): `cast(old_la_code as integer) as la_code`, + `cast(time_period as integer) as year`, `la_name`, measures via + `safe_numeric`, named as in `stg_ees_ks4_national`. +- `pipeline/transform/models/marts/fact_ks4_la_averages.sql` (table): one row + per (`year`, `la_code`) with `la_name` and the columns of + `fact_ks4_national_averages`. Schema tests: unique (`year`, `la_code`); + `year`, `la_code` and `la_name` not null. +- Data tests in `pipeline/transform/tests/`: + - `assert_ks4_years_have_results`: every year in `stg_ees_ks4` has a non-null + Attainment 8 for at least 50% of its rows. DfE's files reach 82% each year; + 2023/24 loads at 0% today. Pre-2019 years come from the legacy model and are + not tested. + - `assert_ks2_info_percentages_loaded`: for every year in `stg_ees_ks2` where + at least 1,000 rows have `total_pupils`, at least 90% of those rows have + `disadvantaged_pct`. DfE's files reach 96–97%. 2022/23 has no pupil counts + and is skipped. + - `assert_ks4_la_averages_cover_las`: the latest year in + `fact_ks4_la_averages` has at least 145 LAs (DfE: 152). +- `pipeline/dags/school_data_pipeline.py`: the annual EES build selects + `stg_ees_ks4_la+`. +- `docs/ARCHITECTURE.md`: LA averages come from DfE's data set. + +### PR 2: backend and UI (after the EES DAG has run on PR 1) + +- `backend/models.py`: `Ks4LaAverage` for `marts.fact_ks4_la_averages`. +- `backend/app.py` `/api/la-averages`: the year is the latest with any school + Attainment 8 (as now). It reads that year's rows from the mart and keys each + `attainment_8_score` by our LA name, through the `local_authority_code` → + `local_authority` pairs in the school data. The response shape is unchanged. + No rows for that year, a missing table or a query error give an empty map, + logged, never another year's figures and never a computed mean. +- `nextjs-app/lib/utils.ts`: `isIndependentSchool(school)`. +- `nextjs-app/components/SecondarySchoolRow.tsx` and + `nextjs-app/components/LeafletMapInner.tsx`: the rule in "Which schools get a + gap". +- `e2e/tests/journeys.spec.ts`: the two journeys under Testing. + +## Testing + +**Extractor (pytest, `pipeline/tests/`, new):** + +- precedence: with the flag on, a year in a newer release is emitted once, from + the newer release; a year only an older release has is kept; with the flag + off, every row passes; +- releases arriving oldest first are still processed newest first; +- KS4 information: an old-format row yields `endks4_pupil_count`, + `ks2_scaledscore_average`, `progress8_banding` and the rest of the table; +- LA filter: a small CSV with national, regional, LA and sub-group rows yields + one row per LA and year, with the declared fields. + +**dbt (local `pgserver`, as for C1):** + +- unit test on `stg_ees_ks2`: `34%` → 34, `34` → 34, `c` → null; +- unit test on `stg_ees_ks4_la` or the mart: codes and years cast, measures + carried; +- the three data tests and the schema tests above; +- `pipeline/tests/test_dag_selectors.py` passes with the new selector. + +**Backend (pytest):** + +- the response gives the mart's figure where the fixture's plain mean differs; +- matching works by code when the mart's `la_name` differs from ours; +- a mart without the served year, and a missing table, give an empty map. + +**Front end (Jest):** + +- `isIndependentSchool` for "Other independent school", "Other independent + special school", "Academy converter" and null; +- `SecondarySchoolRow` and the map card show no gap for an independent school + and show one for a state school with an LA figure. + +**E2E (`journeys.spec.ts`, PR 2):** + +- C2: Bishop Stopford (137086) history shows 2023/24 Attainment 8 64.1 and + Progress 8 +1.02. These are final figures, so the journey stays stable. +- H2: a search returning an independent and a state secondary: the independent + row has no "vs LA avg", the state row has one. The exact gap is not asserted: + DfE's 2025/26 provisional KS4 data is due and would change it. + +## Rollout and verification + +1. PR 1 merges to staging. Tudor runs `school_data_annual_ees` on staging. +2. Through the staging API: Bishop Stopford 2023/24 Attainment 8 64.1 and + Progress 8 1.02; school 147411 2023/24 disadvantaged 34; the LA mart has 152 + LAs for 2024/25. If a data test fails, the build stops before search sync; + investigate before going further. +3. PR 2 merges, so the post-merge E2E gate runs against loaded data. +4. Production, Tudor's decision, in order: promote PR 1, run the EES DAG on + production, promote PR 2. If PR 2 arrives first, the endpoint returns an + empty map and rows show no gap, never a wrong one. +5. Optional cleanup, in the PR 1 description for Tudor: + `delete from raw.ees_ks4_performance where time_period = '202324' and pupil_count is null`. + New-format rows always have `pupil_count` (suppressed values are `c` or `z`, + not null), so this removes exactly the old-label leftovers. +6. Validation on production against DfE's files: every listed school's 2023/24 + Attainment 8, Progress 8 and English and maths figures match; the 2023/24 KS4 + and KS2 information fields match; `/api/la-averages` equals DfE for all 152 + LAs; independent rows show no gap. Then mark C2 and H2 resolved in the audit + report. + +## Expected visible change + +- Secondary history charts and tables run unbroken from 2022/23 to 2024/25, and + 2023/24 Progress 8 appears for about 3,400 schools. +- 2023/24 school information (KS2 and KS4) fills in. +- "vs LA avg" drops for most state secondaries. In Kensington and Chelsea, a + school with Attainment 8 60.0 moves from +24.8 to +5.5. +- Independent schools and City of London schools show no LA gap. + +## Risks + +- **Data-test thresholds stop a good build.** They were set from DfE's own + files (82% and 96–97% against 50% and 90%). A failure means the data changed + shape, which is what they are for. +- **`safe_numeric` is shared by 14 models.** Values written as `n%` were null + and become numbers. No DfE column uses `%` for anything but a percentage. + The staging EES run is the check. +- **Precedence drops data an older release holds more completely.** It is + opt-in for KS4 results, where both files hold the same 57,090 2023/24 keys. +- **The fixed EES data-set id goes stale** when 2025/26 is published. The + endpoint's year check then gives no gap rather than a mismatched one. -- 2.54.0 From 0e987ef06e9904df151e9a6392ed27d4c5bede6f Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 10:34:00 +0100 Subject: [PATCH 02/11] docs: plan for 2023/24 KS4/KS2 data and DfE LA averages (C2, H2) Co-Authored-By: Claude Opus 5.5 --- .../2026-10-06-ks4-2023-24-and-la-averages.md | 1766 +++++++++++++++++ 1 file changed, 1766 insertions(+) create mode 100644 docs/superpowers/plans/2026-10-06-ks4-2023-24-and-la-averages.md diff --git a/docs/superpowers/plans/2026-10-06-ks4-2023-24-and-la-averages.md b/docs/superpowers/plans/2026-10-06-ks4-2023-24-and-la-averages.md new file mode 100644 index 0000000..955e58b --- /dev/null +++ b/docs/superpowers/plans/2026-10-06-ks4-2023-24-and-la-averages.md @@ -0,0 +1,1766 @@ +# 2023/24 Results and DfE LA Averages Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Load every 2023/24 KS4 result and KS2/KS4 school-information figure DfE published (C2), and compare state-funded secondaries with DfE's own LA average (H2). + +**Architecture:** PR 1 changes the EES extractor (the KS4 results stream lets the newest release own each year; the KS4 information stream maps DfE's older 2023/24 names; a new `ees_ks4_la` stream keeps the LA rows of the file the England stream already reads), teaches `safe_numeric` to read `34%`, and adds `marts.fact_ks4_la_averages` plus data tests that fail when a year loads mostly empty. PR 2, after the EES DAG has run on staging, makes `/api/la-averages` serve that mart and drops the LA gap for independent schools. + +**Tech Stack:** Singer SDK tap (Python, pandas), meltanolabs target-postgres, dbt-postgres ~1.10 (YAML unit tests, singular data tests), FastAPI + SQLAlchemy, Next.js / React / TypeScript, Jest + Testing Library, Playwright. + +**Spec:** `docs/superpowers/specs/2026-10-06-ks4-2023-24-and-la-averages-design.md` + +## Global Constraints + +- Release precedence is opt-in and set only on `EESKS4PerformanceStream`. A general rule would wipe KS2: the 2024/25 KS2 file holds 98,448 of the 955,956 rows the 2023/24 release has for 2023/24. +- The KS4 information renames are exactly the spec's 11 pairs; newer files must be unaffected. +- `safe_numeric` pattern: `^-?[0-9]+(\.[0-9]+)?%?$`, cast `rtrim(col, '%')::numeric`; suppression codes stay null. +- LA rows: `geographic_level = 'Local authority'`, `establishment_type_group = 'All state-funded'`, `breakdown_topic = 'Total'`, `breakdown = 'Total'`, compared case-insensitively. Key (`time_period`, `old_la_code`); `old_la_code` is the GIAS LA code. +- Data-test thresholds: KS4 Attainment 8 present for ≥ 50% of a year's rows in `stg_ees_ks4`; KS2 `disadvantaged_pct` present for ≥ 90% of rows with `total_pupils`, in years with ≥ 1,000 such rows; latest year of `fact_ks4_la_averages` has ≥ 145 LAs. +- CI's pytest installs only `requirements.txt` + pytest, httpx<0.28, pyyaml: no `singer_sdk`. Extractor logic under test lives in modules that import only pandas, loaded by file path (as `pipeline/tests/test_gias_encoding.py` does). +- `/api/la-averages` keeps its response shape `{"year", "secondary": {"attainment_8_by_la": {: number}}}`. It never computes a mean and never serves another year's figures; anything missing gives an empty map. +- "vs LA avg" shows only for a school with Attainment 8 that is neither special (`isSpecialSchool`) nor independent (`isIndependentSchool`) and whose LA has a figure. +- No new UI copy. Site copy rules (factual, no generated sentences) unchanged. +- PR 2 must not be merged before `school_data_annual_ees` has run on staging with PR 1, nor promoted before it has run on production. Promotion is Tudor's; never trigger it. +- Never push to `main`. Commits end with the Co-Authored-By trailer the executing session's harness specifies; PR bodies end with the Claude Code line. Do not start a local app server (CLAUDE.md). + +## Review Focus + +1. A mart row whose Attainment 8 is suppressed (null) must be absent from the API map, never `null` in the JSON (a null reaches `att8 - null` in the row). Pinned in Task 10 (`test_an_la_without_a_dfe_figure_is_absent`). +2. An LA we list that DfE does not publish (City of London, code 201) must be absent, not zero. Pinned in Task 10 (same test). +3. A `time_period` with surrounding spaces must still match the year a newer release owns, or the old blanks come back. Pinned in Task 2 (`test_periods_match_despite_surrounding_spaces`). +4. A release whose year cannot be read from its slug must still be processed, after the dated ones, and never drop rows on its own account. Pinned in Task 2 (`test_releases_are_taken_newest_first_whatever_order_the_api_gives`, `test_a_file_without_time_period_is_left_alone`). +5. A school with no `school_type` must keep today's behaviour (gap shown when it has an LA figure). Pinned in Task 11 (`isIndependentSchool` null case and the row test). + +## File Map + +**PR 1 — branch `fix/ks4-2023-24-pipeline`** (exists; spec committed as `ac5b7cc`) + +- Create `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py` — newest release owns the year (pure). +- Create `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_info.py` — declared KS4 information fields and the 2023/24 renames (pure). +- Create `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_summary.py` — the KS4 summary data set's URL, measure map and filters (pure). +- Modify `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py` — wire the three modules; new `EESKs4LaStream`. +- Create `pipeline/tests/test_ees_release_precedence.py`, `pipeline/tests/test_ees_ks4_info_names.py`, `pipeline/tests/test_ees_ks4_summary.py`. +- Modify `pipeline/transform/macros/safe_numeric.sql`. +- Create `pipeline/transform/models/staging/stg_ees_ks2.yml` — unit test for `34%`. +- Modify `pipeline/transform/models/staging/_stg_sources.yml` — source `raw.ees_ks4_la`. +- Create `pipeline/transform/models/staging/stg_ees_ks4_la.sql` and `stg_ees_ks4_la.yml` (unit test). +- Create `pipeline/transform/models/marts/fact_ks4_la_averages.sql`; modify `_marts_schema.yml`. +- Create `pipeline/transform/tests/assert_ks4_years_have_results.sql`, `assert_ks2_info_percentages_loaded.sql`, `assert_ks4_la_averages_cover_las.sql`. +- Modify `pipeline/dags/school_data_pipeline.py` — EES build selects `stg_ees_ks4_la+`. +- Modify `docs/ARCHITECTURE.md`. + +**PR 2 — branch `fix/la-averages-site`** (from `origin/main`; touches no PR 1 file) + +- Modify `backend/models.py` (`Ks4LaAverage`), `backend/app.py` (`_la_averages_payload`, endpoint). +- Create `backend/tests/test_la_averages_marts.py`. +- Modify `nextjs-app/lib/utils.ts` (`isIndependentSchool`), `nextjs-app/components/SecondarySchoolRow.tsx`, `nextjs-app/components/LeafletMapInner.tsx`. +- Modify `nextjs-app/__tests__/lib/utils.test.ts`, `nextjs-app/__tests__/components/SecondarySchoolRow.test.tsx`, `nextjs-app/__tests__/components/LeafletMapInner.test.tsx`. +- Modify `e2e/tests/journeys.spec.ts`. + +**Commands used throughout.** `$SCRATCH` is the session scratchpad directory. Pipeline/backend pytest (CI's environment): + +```bash +uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest -q +``` + +dbt goes through the wrapper `$SCRATCH/dbt_local/dbt.sh ` (Task 1), which runs dbt-postgres 1.10 against a local `pgserver` database and exits with dbt's status. + +--- + +## PR 1 — pipeline + +### Task 1: Local dbt harness + +There is no Docker or Postgres on this machine; `pgserver` bundles Postgres in a pip wheel. `$SCRATCH/dbt_local/` already holds `pgdata/` and `dbt.sh` from the C1 work. + +**Files:** none in the repo. + +- [ ] **Step 1: Start Postgres** + +```bash +cd "$SCRATCH/dbt_local" +cat > start_pg.py <<'EOF' +import pgserver, sys +srv = pgserver.get_server(sys.argv[1], cleanup_mode=None) +print(srv.get_uri()) +if "school_compare" not in srv.psql("select datname from pg_database;"): + srv.psql("create database school_compare;") +EOF +uv run --no-project --with pgserver python start_pg.py "$SCRATCH/dbt_local/pgdata" +``` + +Expected: a URI ending `?host=`. If the socket dir differs from `PG_HOST` in `dbt.sh`, edit `dbt.sh` to match. + +- [ ] **Step 2: Create the raw tables from the tap's own schemas** + +```bash +cat > "$SCRATCH/dbt_local/raw_tables.py" <<'EOF' +"""Create raw. with a text column per declared field (what target-postgres loads).""" +import os, pgserver +from tap_uk_ees.tap import TapUKEES +WANTED = {"ees_ks2_attainment", "ees_ks2_info", "ees_ks4_performance", "ees_ks4_info", "ees_ks4_national", "ees_ks4_la"} +tap = TapUKEES(config={}) +stmts = ["create schema if not exists raw;"] +for stream in tap.streams.values(): + if stream.name in WANTED: + cols = ", ".join(f'"{c}" text' for c in stream.schema["properties"]) + stmts.append(f"drop table if exists raw.{stream.name} cascade; create table raw.{stream.name} ({cols});") + print("raw." + stream.name) +srv = pgserver.get_server(os.environ["SCRATCH"] + "/dbt_local/pgdata", cleanup_mode=None) +srv.psql("\\c school_compare\n" + "\n".join(stmts)) +EOF +cd /Users/tudor/projects/school_compare +SCRATCH="$SCRATCH" uv run --no-project --with ./pipeline/plugins/extractors/tap-uk-ees --with pgserver python "$SCRATCH/dbt_local/raw_tables.py" +``` + +Expected: five `raw.…` lines (`ees_ks4_la` does not exist yet; Task 4 adds it and this script is re-run). + +Then a helper that runs SQL from stdin against the local database (used in Tasks 6 and 8): + +```bash +cat > "$SCRATCH/dbt_local/sql.sh" <<'EOF' +#!/bin/bash +# Run SQL from stdin against the local school_compare database; print the output. +D=$(cd "$(dirname "$0")" && pwd) +SQL=$(cat) D="$D" uv run --no-project --with pgserver python -c ' +import os, pgserver +srv = pgserver.get_server(os.environ["D"] + "/pgdata", cleanup_mode=None) +print(srv.psql("\\c school_compare\n" + os.environ["SQL"]))' +EOF +chmod +x "$SCRATCH/dbt_local/sql.sh" +echo "select count(*) from raw.ees_ks4_performance;" | "$SCRATCH/dbt_local/sql.sh" +``` + +Expected: a count of `0`. + +- [ ] **Step 3: Baseline build of the models this PR touches** + +```bash +$SCRATCH/dbt_local/dbt.sh build --no-partial-parse --select stg_ees_ks2 stg_ees_ks4 stg_ees_ks4_national fact_ks4_national_averages +``` + +Expected: exit 0, the models created (empty). If `pgserver` cannot start, stop and report: the dbt tests will then be proven only by the staging EES DAG run, and the PR description must say so. + +### Task 2: The newest release owns every year it contains (KS4 results) + +**Files:** +- Create: `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py` +- Modify: `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py` (imports; `EESDatasetStream`; `EESKS4PerformanceStream`) +- Test: `pipeline/tests/test_ees_release_precedence.py` + +**Interfaces:** +- Produces: `newest_first(releases: list[dict]) -> list[dict]`; `periods_in(df: pd.DataFrame) -> set[str]`; `drop_owned_periods(df: pd.DataFrame, owned: set[str]) -> tuple[pd.DataFrame, dict[str, int]]`; class attribute `EESDatasetStream._newest_release_owns_period: bool = False`, `True` on `EESKS4PerformanceStream` only. + +- [ ] **Step 1: Write the failing tests** + +`pipeline/tests/test_ees_release_precedence.py`: + +```python +"""DfE re-publishes earlier years inside later KS4 releases. + +The 2024/25 results file holds 2022/23, 2023/24 and 2024/25 under current +column names. The 2023/24 release's own file, re-issued in March 2026 under +older names, was read after it and overwrote every 2023/24 row with blanks +(audit C2). For the KS4 results stream the newest release owns every year it +contains. +""" +import importlib.util +import re +from pathlib import Path + +import pandas as pd +import pytest + +TAP_DIR = (Path(__file__).resolve().parents[1] / 'plugins' / 'extractors' / 'tap-uk-ees' + / 'tap_uk_ees') + + +@pytest.fixture +def precedence(): + spec = importlib.util.spec_from_file_location( + 'release_precedence', TAP_DIR / 'release_precedence.py') + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def _release(period): + return {'id': f'release-{period}', 'time_period': period} + + +def test_releases_are_taken_newest_first_whatever_order_the_api_gives(precedence): + releases = [_release('202223'), _release('202425'), _release(None), _release('202324')] + ordered = precedence.newest_first(releases) + assert [r['time_period'] for r in ordered] == ['202425', '202324', '202223', None] + + +def test_a_year_a_newer_release_supplied_is_dropped_from_an_older_one(precedence): + newer = pd.DataFrame({'time_period': ['202425', '202324', '202223'], 'school_urn': ['1'] * 3}) + older = pd.DataFrame({'time_period': ['202324', '202324', '201920'], 'school_urn': ['1', '2', '1']}) + + kept, skipped = precedence.drop_owned_periods(older, precedence.periods_in(newer)) + + assert list(kept['time_period']) == ['201920'] + assert skipped == {'202324': 2} + + +def test_periods_match_despite_surrounding_spaces(precedence): + owned = precedence.periods_in(pd.DataFrame({'time_period': [' 202324 ']})) + kept, skipped = precedence.drop_owned_periods(pd.DataFrame({'time_period': ['202324']}), owned) + assert owned == {'202324'} + assert kept.empty + assert skipped == {'202324': 1} + + +def test_nothing_is_dropped_before_any_year_is_owned(precedence): + df = pd.DataFrame({'time_period': ['202324'], 'school_urn': ['1']}) + kept, skipped = precedence.drop_owned_periods(df, set()) + assert kept.equals(df) + assert skipped == {} + + +def test_a_file_without_time_period_is_left_alone(precedence): + df = pd.DataFrame({'school_urn': ['1']}) + kept, skipped = precedence.drop_owned_periods(df, {'202324'}) + assert kept.equals(df) + assert skipped == {} + assert precedence.periods_in(df) == set() + + +def test_only_the_ks4_results_stream_opts_in(): + # A general rule would wipe KS2: the 2024/25 KS2 file holds 98,448 of the + # 955,956 rows the 2023/24 release has for 2023/24. + source = (TAP_DIR / 'tap.py').read_text() + opted_in = [chunk.split('(')[0] for chunk in source.split('\nclass ')[1:] + if re.search(r'_newest_release_owns_period\s*=\s*True', chunk)] + assert opted_in == ['EESKS4PerformanceStream'] +``` + +- [ ] **Step 2: Run them to see them fail** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest pipeline/tests/test_ees_release_precedence.py -q` +Expected: 5 errors (`FileNotFoundError` for `release_precedence.py`) and 1 failure (`test_only_the_ks4_results_stream_opts_in`: `[] == ['EESKS4PerformanceStream']`). + +- [ ] **Step 3: Write the module** + +`pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py`: + +```python +"""Which release a row comes from when DfE re-publishes a year. + +DfE re-publishes earlier years inside later releases: the 2024/25 KS4 results +file holds 2022/23, 2023/24 and 2024/25. The 2023/24 release's own file, +re-issued on 10 March 2026 under older column names, was read after it and +overwrote every 2023/24 row with blanks (audit C2). A stream that opts in +treats the newest release as the authority for every year it contains. + +Free of the Singer SDK so CI's pytest, which installs only the backend's +requirements, can load it. +""" +from __future__ import annotations + +import pandas as pd + + +def newest_first(releases: list[dict]) -> list[dict]: + """Releases by time_period, newest first. A release whose time_period is + unknown goes last, in the order given.""" + dated = [r for r in releases if r.get("time_period")] + undated = [r for r in releases if not r.get("time_period")] + return sorted(dated, key=lambda r: r["time_period"], reverse=True) + undated + + +def periods_in(df: pd.DataFrame) -> set[str]: + """The years a release's rows cover.""" + if "time_period" not in df.columns: + return set() + return set(df["time_period"].astype(str).str.strip()) - {""} + + +def drop_owned_periods( + df: pd.DataFrame, owned: set[str] +) -> tuple[pd.DataFrame, dict[str, int]]: + """Drop the rows for years a newer release already supplied. + + Returns the rows kept and, for each year dropped, how many rows went. + """ + if "time_period" not in df.columns or not owned: + return df, {} + periods = df["time_period"].astype(str).str.strip() + dropped = periods.isin(owned) + skipped = {str(k): int(v) for k, v in periods[dropped].value_counts().items()} + return df[~dropped], skipped +``` + +- [ ] **Step 4: Wire it into the tap** + +In `tap.py`, below `from singer_sdk import typing as th`, add: + +```python +from tap_uk_ees.release_precedence import drop_owned_periods, newest_first, periods_in +``` + +In `EESDatasetStream`, extend the class docstring's last paragraph and add the attribute after `_column_renames`: + +```python + Subclasses may set _column_renames to map messy CSV column names to + clean Singer field names before yielding records. + Subclasses may set _newest_release_owns_period when DfE re-publishes + earlier years in later releases and the newest copy is the authority. + """ +``` + +```python + _column_renames: dict = {} # CSV column name → Singer field name + _newest_release_owns_period: bool = False # see release_precedence.py +``` + +In `get_records`, directly after the `self.logger.info("Found %d release(s) for %s", …)` call: + +```python + if self._newest_release_owns_period: + releases = newest_first(releases) + owned_periods: set[str] = set() +``` + +and directly after the "Drop rows with no URN" block (before `self.logger.info("Emitting %d school-level rows …")`): + +```python + if self._newest_release_owns_period: + df, skipped = drop_owned_periods(df, owned_periods) + for period, count in sorted(skipped.items()): + self.logger.info( + "Skipping %d rows for %s from release %s: a newer release supplied that year", + count, period, release_id, + ) + owned_periods |= periods_in(df) +``` + +In `EESKS4PerformanceStream`, after `_target_filename = "performance_tables_schools"`: + +```python + # DfE's 2024/25 file re-publishes 2022/23 and 2023/24 under current names; + # the 2023/24 release's own file uses older ones (audit C2). + _newest_release_owns_period = True +``` + +- [ ] **Step 5: Run the tests and an import check** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest pipeline/tests/test_ees_release_precedence.py -q` +Expected: `6 passed`. + +Run: `uv run --no-project --with ./pipeline/plugins/extractors/tap-uk-ees python -c "from tap_uk_ees.tap import TapUKEES; t = TapUKEES(config={}); print(t.streams['ees_ks4_performance']._newest_release_owns_period, t.streams['ees_ks2_attainment']._newest_release_owns_period)"` +Expected: `True False`. + +- [ ] **Step 6: Commit** + +```bash +git add pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py pipeline/tests/test_ees_release_precedence.py +git commit -m "fix(ees): newest KS4 release owns every year it contains (C2)" +``` + +### Task 3: KS4 information keeps its 2023/24 fields + +**Files:** +- Create: `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_info.py` +- Modify: `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py` (`EESKS4InfoStream` and the comment above it) +- Test: `pipeline/tests/test_ees_ks4_info_names.py` + +**Interfaces:** +- Produces: `KS4_INFO_FIELDS: tuple[str, ...]` (18 names, excluding `time_period`, `school_urn`); `KS4_INFO_RENAMES: dict[str, str]` (11 pairs). The stream's schema is unchanged: same 20 fields, same order. + +- [ ] **Step 1: Write the failing tests** + +`pipeline/tests/test_ees_ks4_info_names.py`: + +```python +"""KS4 school information for 2023/24 exists only in the 2023/24 release, +whose file uses DfE's older column names. The stream declared only the newer +ones, so every 2023/24 field loaded as null (audit C2). The headers below are +DfE's, copied from 202324_information_about_schools_final.csv and +202425_information_about_schools_final.csv. +""" +import importlib.util +from pathlib import Path + +import pytest + +MODULE = (Path(__file__).resolve().parents[1] / 'plugins' / 'extractors' / 'tap-uk-ees' + / 'tap_uk_ees' / 'ks4_info.py') + +HEADER_2023_24 = ( + 'time_period', 'time_identifier', 'geographic_level', 'country_code', 'country_name', + 'school_laestab', 'school_urn', 'school_name', 'old_la_code', 'new_la_code', 'la_name', + 'version', 'establishment_type_group', 'full_address', 'telnum', 'pcon_code', 'pcon_name', + 'contflag', 'iclose', 'reldenom', 'admpol_pt', 'egender', 'feeder', 'agerange', + 't_allks_pupils', 't_allks_boys', 't_allks_girls', 't_pupils', 't_boys', 'pt_boys', + 't_girls', 'pt_girls', 'avg_ks2_scaledscore', 't_prior_lo', 'pt_prior_lo', 't_prior_av', + 'pt_prior_av', 't_prior_hi', 'pt_prior_hi', 't_disadvantaged', 'pt_disadvantaged', + 't_not_disadvantaged', 'pt_not_disadvantaged', 't_language_not_english', + 'pt_language_not_english', 't_language_english', 'pt_language_english', + 't_language_unknown', 'pt_language_unknown', 't_not_mobile', 'pt_not_mobile', + 't_sen_with_ehcp', 'pt_sen_with_ehcp', 't_sen', 'pt_sen', 't_sen_no_ehcp', + 'pt_sen_no_ehcp', 'diffn_att8', 'diffn_p8mea', 'p8_banding', +) + +HEADER_2024_25 = ( + 'time_period', 'time_identifier', 'geographic_level', 'country_code', 'country_name', + 'school_laestab', 'school_urn', 'school_name', 'old_la_code', 'new_la_code', 'la_name', + 'version', 'establishment_type_group', 'full_address', 'telnum', 'pcon_code', 'pcon_name', + 'contflag', 'iclose', 'reldenom', 'admpol_pt', 'egender', 'feeder', 'agerange', + 'allks_pupil_count', 'allks_boys_count', 'allks_girls_count', 'endks4_pupil_count', + 'ks2_scaledscore_average', 'sen_with_ehcp_pupil_count', 'sen_with_ehcp_pupil_percent', + 'sen_pupil_count', 'sen_pupil_percent', 'sen_no_ehcp_pupil_count', + 'sen_no_ehcp_pupil_percent', 'attainment8_diffn', 'progress8_diffn', 'progress8_banding', +) + + +@pytest.fixture +def ks4_info(): + spec = importlib.util.spec_from_file_location('ks4_info', MODULE) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def test_every_declared_field_is_in_the_2023_24_file_once_renamed(ks4_info): + renamed = {ks4_info.KS4_INFO_RENAMES.get(c, c) for c in HEADER_2023_24} + assert set(ks4_info.KS4_INFO_FIELDS) - renamed == set() + + +def test_every_declared_field_is_in_the_2024_25_file(ks4_info): + assert set(ks4_info.KS4_INFO_FIELDS) - set(HEADER_2024_25) == set() + + +def test_the_renames_cannot_collide_with_either_file(ks4_info): + # An old name in the current file, or a new name already in the old file, + # would let a rename overwrite a real column. + assert set(ks4_info.KS4_INFO_RENAMES) & set(HEADER_2024_25) == set() + assert set(ks4_info.KS4_INFO_RENAMES.values()) & set(HEADER_2023_24) == set() + + +def test_every_rename_names_a_declared_field(ks4_info): + assert set(ks4_info.KS4_INFO_RENAMES.values()) <= set(ks4_info.KS4_INFO_FIELDS) +``` + +- [ ] **Step 2: Run them to see them fail** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest pipeline/tests/test_ees_ks4_info_names.py -q` +Expected: 4 errors, `FileNotFoundError` for `ks4_info.py`. + +- [ ] **Step 3: Write the module** + +`pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_info.py`: + +```python +"""KS4 school information: the fields the stream declares, and the older +names DfE used for them. + +2023/24 information exists only in the 2023/24 release, whose +202324_information_about_schools_final.csv (re-issued 10 March 2026) uses the +older names. Without the renames every 2023/24 field loaded as null (audit C2). +Newer files contain none of the older names, so the renames leave them alone. + +Free of the Singer SDK so CI's pytest can load it. +""" + +# Declared Singer fields besides the required time_period and school_urn. +KS4_INFO_FIELDS = ( + "school_laestab", + "school_name", + "establishment_type_group", + "reldenom", + "admpol_pt", + "egender", + "agerange", + "allks_pupil_count", + "allks_boys_count", + "allks_girls_count", + "endks4_pupil_count", + "ks2_scaledscore_average", + "sen_with_ehcp_pupil_percent", + "sen_pupil_percent", + "sen_no_ehcp_pupil_percent", + "attainment8_diffn", + "progress8_diffn", + "progress8_banding", +) + +# 2023/24 column name → declared field. +KS4_INFO_RENAMES = { + "t_allks_pupils": "allks_pupil_count", + "t_allks_boys": "allks_boys_count", + "t_allks_girls": "allks_girls_count", + "t_pupils": "endks4_pupil_count", + "avg_ks2_scaledscore": "ks2_scaledscore_average", + "pt_sen_with_ehcp": "sen_with_ehcp_pupil_percent", + "pt_sen": "sen_pupil_percent", + "pt_sen_no_ehcp": "sen_no_ehcp_pupil_percent", + "diffn_att8": "attainment8_diffn", + "diffn_p8mea": "progress8_diffn", + "p8_banding": "progress8_banding", +} +``` + +- [ ] **Step 4: Use it in the stream** + +In `tap.py`, add to the imports: + +```python +from tap_uk_ees.ks4_info import KS4_INFO_FIELDS, KS4_INFO_RENAMES +``` + +Replace the comment above `class EESKS4InfoStream` and the class's schema: + +```python +# ── KS4 Information (wide format: one row per school, context/demographics) ── +# Files: 202425_information_about_schools_final.csv (38 cols, current names); +# 202324_information_about_schools_final.csv (60 cols, older names — the only +# source of 2023/24 information). Field list and renames: ks4_info.py. + +class EESKS4InfoStream(EESDatasetStream): + name = "ees_ks4_info" + primary_keys = ["school_urn", "time_period"] + _publication_slug = "key-stage-4-performance" + _target_filename = "information_about_schools" + _column_renames = KS4_INFO_RENAMES + schema = th.PropertiesList( + th.Property("time_period", th.StringType, required=True), + th.Property("school_urn", th.StringType, required=True), + *[th.Property(field, th.StringType) for field in KS4_INFO_FIELDS], + ).to_dict() +``` + +- [ ] **Step 5: Run the tests and check the schema is unchanged** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest pipeline/tests/test_ees_ks4_info_names.py -q` +Expected: `4 passed`. + +Run: `uv run --no-project --with ./pipeline/plugins/extractors/tap-uk-ees python -c "from tap_uk_ees.tap import TapUKEES; print(list(TapUKEES(config={}).streams['ees_ks4_info'].schema['properties']))"` +Expected: `['time_period', 'school_urn', 'school_laestab', 'school_name', 'establishment_type_group', 'reldenom', 'admpol_pt', 'egender', 'agerange', 'allks_pupil_count', 'allks_boys_count', 'allks_girls_count', 'endks4_pupil_count', 'ks2_scaledscore_average', 'sen_with_ehcp_pupil_percent', 'sen_pupil_percent', 'sen_no_ehcp_pupil_percent', 'attainment8_diffn', 'progress8_diffn', 'progress8_banding']` — the same 20 fields in the same order as before. + +- [ ] **Step 6: Commit** + +```bash +git add pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_info.py pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py pipeline/tests/test_ees_ks4_info_names.py +git commit -m "fix(ees): read 2023/24 KS4 school information under DfE's older names (C2)" +``` + +### Task 4: KS4 LA averages stream + +**Files:** +- Create: `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_summary.py` +- Modify: `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py` (KS4 national block and class; new class; `discover_streams`) +- Test: `pipeline/tests/test_ees_ks4_summary.py` + +**Interfaces:** +- Produces: `KS4_SUMMARY_CSV_URL: str`; `KS4_HEADLINE_COL_MAP: dict[str, str]` (the 8 pairs formerly `_KS4_NATIONAL_COL_MAP`); `headline_rows(df, geographic_level: str) -> pd.DataFrame`; `headline_record(row, keys: tuple[str, ...]) -> dict[str, str]`; Singer stream `ees_ks4_la` with fields `time_period`, `old_la_code`, `new_la_code`, `la_name`, `attainment_8_score`, `progress_8_score`, `english_maths_standard_pass_pct`, `english_maths_strong_pass_pct`, `ebacc_entry_pct`, `ebacc_standard_pass_pct`, `ebacc_strong_pass_pct`, `ebacc_avg_score`. + +- [ ] **Step 1: Record the national stream's current output (baseline for the refactor)** + +```bash +uv run --no-project --with ./pipeline/plugins/extractors/tap-uk-ees python -c " +from tap_uk_ees.tap import TapUKEES +rows = sorted(TapUKEES(config={}).streams['ees_ks4_national'].get_records(None), key=lambda r: r['time_period']) +print(len(rows), rows[-1])" | tee "$SCRATCH/ks4_national_before.txt" +``` + +Expected: `7` rows (2018/19–2024/25); the last is 2024/25 with `progress_8_score: 'z'`. + +- [ ] **Step 2: Write the failing tests** + +`pipeline/tests/test_ees_ks4_summary.py`: + +```python +"""DfE's KS4 "summary, all state-funded" data set holds England, regional and +LA rows. The England stream kept only the England row; its LA rows match +DfE's published LA averages exactly, where the API's own mean was 7 points +low (audit H2). Values below are DfE's for Kensington and Chelsea (207) and +Wandsworth (212). +""" +import importlib.util +import io +from pathlib import Path + +import pandas as pd +import pytest + +MODULE = (Path(__file__).resolve().parents[1] / 'plugins' / 'extractors' / 'tap-uk-ees' + / 'tap_uk_ees' / 'ks4_summary.py') + +CSV = """time_period,geographic_level,old_la_code,new_la_code,la_name,establishment_type_group,breakdown_topic,breakdown,attainment8_average,progress8_average,engmath_94_percent,engmath_95_percent,ebacc_entering_percent,ebacc_94_percent,ebacc_95_percent,ebacc_aps_average +202425,National,,,,All state-funded,Total,Total,46.1,z,64.5,45.7,40.5,26.9,17.7,4.1 +202425,National,,,,All state-funded,Sex,Boys,44.1,z,62.0,43.0,38.0,24.0,16.0,3.9 +202425,Regional,,,,All state-funded,Total,Total,47.2,z,66.0,47.0,41.0,28.0,18.0,4.2 +202425,Local authority,207,E09000020,Kensington and Chelsea,All state-funded,Total,Total,54.5,z,77,61.4,45.6,32,26.6,4.89 +202425,Local authority,207,E09000020,Kensington and Chelsea,All state-funded,Sex,Girls,57.0,z,80,64.0,48.0,35,28.0,5.1 +202324,Local authority,207,E09000020,Kensington and Chelsea,All state-funded,Total,Total,54.5,0.29,76,60.0,44.0,31,25.0,4.8 +202425,Local authority,212,E09000032,Wandsworth,All state-funded,Total,Total,51.8,z,72,55.0,50.0,33,24.0,4.6 +202425,Local authority,212,E09000032,Wandsworth,Academies and free schools,Total,Total,52.0,z,73,56.0,51.0,34,25.0,4.7 +""" + +LA_KEYS = ('time_period', 'old_la_code', 'new_la_code', 'la_name') + + +def _df(): + return pd.read_csv(io.StringIO(CSV), dtype=str, keep_default_na=False) + + +@pytest.fixture +def summary(): + spec = importlib.util.spec_from_file_location('ks4_summary', MODULE) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def test_la_rows_are_one_per_la_and_year(summary): + rows = summary.headline_rows(_df(), 'Local authority') + assert sorted(zip(rows['time_period'], rows['old_la_code'])) == [ + ('202324', '207'), ('202425', '207'), ('202425', '212')] + + +def test_an_la_record_carries_codes_name_and_the_headline_measures(summary): + rows = summary.headline_rows(_df(), 'Local authority') + row = rows[(rows['time_period'] == '202425') & (rows['old_la_code'] == '207')].iloc[0] + assert summary.headline_record(row, LA_KEYS) == { + 'time_period': '202425', 'old_la_code': '207', 'new_la_code': 'E09000020', + 'la_name': 'Kensington and Chelsea', + 'attainment_8_score': '54.5', 'progress_8_score': 'z', + 'english_maths_standard_pass_pct': '77', 'english_maths_strong_pass_pct': '61.4', + 'ebacc_entry_pct': '45.6', 'ebacc_standard_pass_pct': '32', + 'ebacc_strong_pass_pct': '26.6', 'ebacc_avg_score': '4.89', + } + + +def test_the_england_rows_are_one_per_year(summary): + rows = summary.headline_rows(_df(), 'National') + assert list(rows['time_period']) == ['202425'] + assert summary.headline_record(rows.iloc[0], ('time_period',))['attainment_8_score'] == '46.1' + + +def test_column_names_and_labels_match_whatever_their_case(summary): + df = _df() + df.columns = [c.upper() for c in df.columns] + df['GEOGRAPHIC_LEVEL'] = df['GEOGRAPHIC_LEVEL'].str.upper() + assert len(summary.headline_rows(df, 'Local authority')) == 3 +``` + +- [ ] **Step 3: Run them to see them fail** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest pipeline/tests/test_ees_ks4_summary.py -q` +Expected: 4 errors, `FileNotFoundError` for `ks4_summary.py`. + +- [ ] **Step 4: Write the module** + +`pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_summary.py`: + +```python +"""DfE's KS4 "summary, all state-funded" data set (Key stage 4 performance). + +One CSV holds England, regional and local-authority headline rows for every +year since 2018/19. The England stream (ees_ks4_national) and the LA stream +(ees_ks4_la) both read it. Its LA rows match DfE's published performance-table +LA averages exactly (audit H2). Suppressed values ('z', 'x') become NULL in +dbt; Progress 8 is 'z' in years with no KS2 baseline (2024/25): DfE policy, +not missing data. + +Free of the Singer SDK so CI's pytest can load it. +""" +from __future__ import annotations + +import pandas as pd + +KS4_SUMMARY_CSV_URL = ( + "https://explore-education-statistics.service.gov.uk/data-catalogue/" + "data-set/1b649e16-01e8-435b-a814-56be2faf9054/csv" +) + +# CSV column → Singer field: the same 8 headline measures at every level. +KS4_HEADLINE_COL_MAP = { + "attainment8_average": "attainment_8_score", + "progress8_average": "progress_8_score", + "engmath_94_percent": "english_maths_standard_pass_pct", + "engmath_95_percent": "english_maths_strong_pass_pct", + "ebacc_entering_percent": "ebacc_entry_pct", + "ebacc_94_percent": "ebacc_standard_pass_pct", + "ebacc_95_percent": "ebacc_strong_pass_pct", + "ebacc_aps_average": "ebacc_avg_score", +} + + +def headline_rows(df: pd.DataFrame, geographic_level: str) -> pd.DataFrame: + """All-pupil rows for all state-funded schools at one geographic level + ("National" or "Local authority"). Column names are lower-cased first; + a filter column the file lacks is not applied.""" + df = df.copy() + df.columns = [c.strip().lower() for c in df.columns] + for col, want in ( + ("geographic_level", geographic_level), + ("establishment_type_group", "All state-funded"), + ("breakdown_topic", "Total"), + ("breakdown", "Total"), + ): + if col in df.columns: + df = df[df[col].str.strip().str.lower() == want.lower()] + return df + + +def headline_record(row: pd.Series, keys: tuple[str, ...]) -> dict[str, str]: + """A Singer record: the identifying columns, then the headline measures.""" + record = {key: str(row.get(key, "")).strip() for key in keys} + for csv_col, field in KS4_HEADLINE_COL_MAP.items(): + record[field] = str(row.get(csv_col, "")).strip() + return record +``` + +- [ ] **Step 5: Use it in the tap** + +Add to the imports in `tap.py`: + +```python +from tap_uk_ees.ks4_summary import ( + KS4_HEADLINE_COL_MAP, + KS4_SUMMARY_CSV_URL, + headline_record, + headline_rows, +) +``` + +Replace everything from the line `# ── KS4 National Headlines (national level only — one row per year) ──────────` through the end of `class EESKs4NationalStream` (the line ` yield record` before the `# ── Legacy KS2` banner) with: + +```python +# ── KS4 National and LA Headlines (one data set, two streams) ──────────────── +# DfE's "summary, all state-funded" data set: England, regional and LA rows, +# 2018/19 → latest. URL, measures and filters: ks4_summary.py. + +def _read_ks4_summary(logger): + """Download DfE's KS4 summary data set.""" + import pandas as pd + + logger.info("Downloading KS4 summary data set: %s", KS4_SUMMARY_CSV_URL) + resp = requests.get(KS4_SUMMARY_CSV_URL, timeout=60) + resp.raise_for_status() + return pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False) + + +class EESKs4NationalStream(Stream): + """National KS4 headline averages — one row per academic year (England, + all state-funded schools, all pupils).""" + + name = "ees_ks4_national" + primary_keys = ["time_period"] + replication_key = None + + schema = th.PropertiesList( + th.Property("time_period", th.StringType, required=True), + *[th.Property(out, th.StringType) for out in KS4_HEADLINE_COL_MAP.values()], + ).to_dict() + + def get_records(self, context): + df = headline_rows(_read_ks4_summary(self.logger), "National") + self.logger.info("Emitting %d national KS4 rows", len(df)) + for _, row in df.iterrows(): + yield headline_record(row, ("time_period",)) + + +class EESKs4LaStream(Stream): + """DfE's KS4 local-authority averages — one row per academic year and LA + (all state-funded schools, all pupils), from the same data set as + ees_ks4_national. They match DfE's published performance-table LA averages + and replace a mean the API took over every school, independent and special + included (audit H2). old_la_code is the GIAS LA code: schools join on it, + not on the name.""" + + name = "ees_ks4_la" + primary_keys = ["time_period", "old_la_code"] + replication_key = None + + schema = th.PropertiesList( + th.Property("time_period", th.StringType, required=True), + th.Property("old_la_code", th.StringType, required=True), + th.Property("new_la_code", th.StringType), + th.Property("la_name", th.StringType), + *[th.Property(out, th.StringType) for out in KS4_HEADLINE_COL_MAP.values()], + ).to_dict() + + def get_records(self, context): + df = headline_rows(_read_ks4_summary(self.logger), "Local authority") + self.logger.info("Emitting %d LA KS4 rows", len(df)) + for _, row in df.iterrows(): + yield headline_record(row, ("time_period", "old_la_code", "new_la_code", "la_name")) +``` + +In `discover_streams`, add `EESKs4LaStream(self),` after `EESKs4NationalStream(self),`. + +Then confirm nothing else used the old names: `grep -rn "_KS4_NATIONAL" pipeline/` → no output. + +- [ ] **Step 6: Run the tests, then both streams against DfE** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest pipeline/tests/test_ees_ks4_summary.py -q` +Expected: `4 passed`. + +```bash +uv run --no-project --with ./pipeline/plugins/extractors/tap-uk-ees python -c " +from tap_uk_ees.tap import TapUKEES +tap = TapUKEES(config={}) +nat = sorted(tap.streams['ees_ks4_national'].get_records(None), key=lambda r: r['time_period']) +la = list(tap.streams['ees_ks4_la'].get_records(None)) +print(len(nat), nat[-1]) +print(len(la), len({(r['time_period'], r['old_la_code']) for r in la})) +print([r for r in la if (r['time_period'], r['old_la_code']) in {('202425', '207'), ('202425', '212')}])" +``` + +Expected: first line identical to `$SCRATCH/ks4_national_before.txt`; then `1058 1058`; then Kensington and Chelsea with `attainment_8_score: '54.5'` and Wandsworth with `'51.8'`. + +- [ ] **Step 7: Re-create the raw tables, now including `raw.ees_ks4_la`** + +Run Task 1 Step 2's command again. Expected: six `raw.…` lines. + +- [ ] **Step 8: Commit** + +```bash +git add pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_summary.py pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py pipeline/tests/test_ees_ks4_summary.py +git commit -m "feat(ees): keep DfE's KS4 LA averages from the summary data set (H2)" +``` + +### Task 5: LA averages staging model, mart, schema tests and DAG selector + +**Files:** +- Modify: `pipeline/transform/models/staging/_stg_sources.yml` +- Create: `pipeline/transform/models/staging/stg_ees_ks4_la.sql`, `pipeline/transform/models/staging/stg_ees_ks4_la.yml` +- Create: `pipeline/transform/models/marts/fact_ks4_la_averages.sql` +- Modify: `pipeline/transform/models/marts/_marts_schema.yml` +- Modify: `pipeline/dags/school_data_pipeline.py:196` +- Modify: `docs/ARCHITECTURE.md` + +**Interfaces:** +- Consumes: `raw.ees_ks4_la` (Task 4's fields). +- Produces: `marts.fact_ks4_la_averages(year integer, la_code integer, la_name text, attainment_8_score numeric, progress_8_score numeric, english_maths_standard_pass_pct numeric, english_maths_strong_pass_pct numeric, ebacc_entry_pct numeric, ebacc_standard_pass_pct numeric, ebacc_strong_pass_pct numeric, ebacc_avg_score numeric)`, unique on (`year`, `la_code`). PR 2's backend reads `year`, `la_code`, `la_name`, `attainment_8_score`. + +- [ ] **Step 1: Declare the source and write the unit test** + +In `_stg_sources.yml`, after the `ees_ks4_national` entry: + +```yaml + - name: ees_ks4_la + description: Official KS4 local-authority averages (all state-funded schools) from the same DfE EES data set as ees_ks4_national — one row per academic year and LA +``` + +`pipeline/transform/models/staging/stg_ees_ks4_la.yml`: + +```yaml +version: 2 + +unit_tests: + - name: stg_ees_ks4_la_casts_codes_years_and_measures + description: DfE's LA rows, keyed by the GIAS LA code. Suppressed values ('z') become null. + model: stg_ees_ks4_la + given: + - input: source('raw', 'ees_ks4_la') + rows: + - {time_period: '202425', old_la_code: '207', new_la_code: 'E09000020', la_name: 'Kensington and Chelsea', attainment_8_score: '54.5', progress_8_score: 'z', english_maths_standard_pass_pct: '77'} + - {time_period: '202324', old_la_code: '207', new_la_code: 'E09000020', la_name: 'Kensington and Chelsea', attainment_8_score: '54.5', progress_8_score: '0.29', english_maths_standard_pass_pct: '76'} + expect: + rows: + - {year: 202425, la_code: 207, la_name: 'Kensington and Chelsea', attainment_8_score: 54.5, progress_8_score: null, english_maths_standard_pass_pct: 77} + - {year: 202324, la_code: 207, la_name: 'Kensington and Chelsea', attainment_8_score: 54.5, progress_8_score: 0.29, english_maths_standard_pass_pct: 76} +``` + +- [ ] **Step 2: Run it to see it fail** + +Run: `$SCRATCH/dbt_local/dbt.sh test --no-partial-parse --select "stg_ees_ks4_la,test_type:unit"` +Expected: non-zero exit; the selector or compilation fails because `stg_ees_ks4_la` does not exist. + +- [ ] **Step 3: Write the staging model and the mart** + +`pipeline/transform/models/staging/stg_ees_ks4_la.sql`: + +```sql +-- Staging model: official DfE KS4 local-authority averages — one row per +-- academic year and LA (all state-funded schools, all pupils). Same EES data +-- set as stg_ees_ks4_national. la_code is the GIAS LA code +-- (dim_location.local_authority_code). Suppressed values ('z', 'x') are +-- coerced to NULL by safe_numeric. + +select + cast(trim(time_period) as integer) as year, + cast(trim(old_la_code) as integer) as la_code, + trim(la_name) as la_name, + {{ safe_numeric('attainment_8_score') }} as attainment_8_score, + {{ safe_numeric('progress_8_score') }} as progress_8_score, + {{ safe_numeric('english_maths_standard_pass_pct') }} as english_maths_standard_pass_pct, + {{ safe_numeric('english_maths_strong_pass_pct') }} as english_maths_strong_pass_pct, + {{ safe_numeric('ebacc_entry_pct') }} as ebacc_entry_pct, + {{ safe_numeric('ebacc_standard_pass_pct') }} as ebacc_standard_pass_pct, + {{ safe_numeric('ebacc_strong_pass_pct') }} as ebacc_strong_pass_pct, + {{ safe_numeric('ebacc_avg_score') }} as ebacc_avg_score +from {{ source('raw', 'ees_ks4_la') }} +where time_period ~ '^[0-9]+$' + and old_la_code ~ '^[0-9]+$' +``` + +`pipeline/transform/models/marts/fact_ks4_la_averages.sql`: + +```sql +{{ config(materialized='table') }} + +-- Mart: OFFICIAL DfE KS4 local-authority averages — one row per academic year +-- and LA (all state-funded schools, all pupils), from the same EES data set +-- as fact_ks4_national_averages. Feeds "vs LA avg" on search rows and map +-- cards, replacing a mean the API took over every school in the LA, +-- independent and special included (audit H2). la_code is the GIAS LA code. + +select + year, + la_code, + la_name, + attainment_8_score, + progress_8_score, + english_maths_standard_pass_pct, + english_maths_strong_pass_pct, + ebacc_entry_pct, + ebacc_standard_pass_pct, + ebacc_strong_pass_pct, + ebacc_avg_score +from {{ ref('stg_ees_ks4_la') }} +order by year, la_code +``` + +In `_marts_schema.yml`, correct the stale description of `fact_ks4_national_averages` and add the new mart after it: + +```yaml + - name: fact_ks4_national_averages + description: Official DfE KS4 national headline averages (England, all state-funded schools, all pupils) — one row per academic year + columns: + - name: year + tests: [not_null, unique] + + - name: fact_ks4_la_averages + description: Official DfE KS4 local-authority averages (all state-funded schools, all pupils) — one row per academic year and LA; la_code is the GIAS LA code + columns: + - name: year + tests: [not_null] + - name: la_code + tests: [not_null] + - name: la_name + tests: [not_null] + tests: + - unique: + column_name: "year || '-' || la_code" +``` + +- [ ] **Step 4: Run the unit test and build** + +Run: `$SCRATCH/dbt_local/dbt.sh build --no-partial-parse --select stg_ees_ks4_la fact_ks4_la_averages` +Expected: exit 0; the unit test PASS, both models created, the four schema tests PASS (on an empty table). + +- [ ] **Step 5: Select the new models in the annual EES build** + +In `pipeline/dags/school_data_pipeline.py`, in the `dbt_build_ees` command, change `stg_ees_ks4_national+` to `stg_ees_ks4_national+ stg_ees_ks4_la+`. + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest pipeline/tests/test_dag_selectors.py -q` +Expected: all pass (the new mart reads only `stg_ees_ks4_la`, which the same build selects). + +- [ ] **Step 6: Document the benchmark source** + +In `docs/ARCHITECTURE.md`, after the paragraph that begins "Search starts from a cached latest-row-per-school snapshot.", add: + +```markdown +Benchmarks are DfE's own figures, read from marts: `fact_ks2_national_averages` +and `fact_ks4_national_averages` for England, and `fact_ks4_la_averages` (all +state-funded schools) for the search rows' "vs LA avg". The API never averages +school rows to make a benchmark. +``` + +- [ ] **Step 7: Commit** + +```bash +git add pipeline/transform/models/staging/_stg_sources.yml pipeline/transform/models/staging/stg_ees_ks4_la.sql pipeline/transform/models/staging/stg_ees_ks4_la.yml pipeline/transform/models/marts/fact_ks4_la_averages.sql pipeline/transform/models/marts/_marts_schema.yml pipeline/dags/school_data_pipeline.py docs/ARCHITECTURE.md +git commit -m "feat(dbt): fact_ks4_la_averages from DfE's LA rows (H2)" +``` + +### Task 6: Data tests that fail when a year loads empty + +The KS2 test stays red at the end of this task: the `%` fix is Task 7. + +**Files:** +- Create: `pipeline/transform/tests/assert_ks4_years_have_results.sql` +- Create: `pipeline/transform/tests/assert_ks2_info_percentages_loaded.sql` +- Create: `pipeline/transform/tests/assert_ks4_la_averages_cover_las.sql` + +**Interfaces:** +- Consumes: `stg_ees_ks4` (`year`, `attainment_8_score`), `stg_ees_ks2` (`year`, `total_pupils`, `disadvantaged_pct`), `fact_ks4_la_averages` (Task 5). + +- [ ] **Step 1: Write the three tests** + +`pipeline/transform/tests/assert_ks4_years_have_results.sql`: + +```sql +-- Every year loaded from EES has an Attainment 8 for at least half its +-- schools. DfE's files reach 82% each year; 2023/24 loaded at 0% after DfE +-- re-issued its file under older column names (audit C2). Pre-2019 years come +-- from stg_legacy_ks4 and are not checked here. + +select + year, + count(*) as schools, + count(attainment_8_score) as with_attainment_8 +from {{ ref('stg_ees_ks4') }} +group by year +having count(attainment_8_score) < 0.5 * count(*) +``` + +`pipeline/transform/tests/assert_ks2_info_percentages_loaded.sql`: + +```sql +-- Where a year's KS2 school information loaded (pupil counts present), its +-- percentages loaded too. DfE's files give a disadvantaged % for 96–97% of +-- those schools; 2023/24 loaded none because the file writes "34%" (audit +-- C2). A year without an information file (2022/23) has no pupil counts and +-- is skipped. + +select + year, + count(total_pupils) as with_pupils, + count(*) filter (where total_pupils is not null and disadvantaged_pct is not null) + as with_disadvantaged +from {{ ref('stg_ees_ks2') }} +group by year +having count(total_pupils) >= 1000 + and count(*) filter (where total_pupils is not null and disadvantaged_pct is not null) + < 0.9 * count(total_pupils) +``` + +`pipeline/transform/tests/assert_ks4_la_averages_cover_las.sql`: + +```sql +-- DfE publishes an average for about 152 LAs a year. Fewer than 145 in the +-- latest year (or none at all) means the LA filter in tap_uk_ees +-- (ks4_summary.headline_rows) stopped matching, e.g. after DfE renamed a label. + +with latest as ( + select max(year) as year from {{ ref('fact_ks4_la_averages') }} +) + +select l.year, count(f.la_code) as las +from latest l +left join {{ ref('fact_ks4_la_averages') }} f on f.year = l.year +group by l.year +having count(f.la_code) < 145 +``` + +- [ ] **Step 2: Load fixtures that reproduce each failure, and watch the tests fail** + +```bash +"$SCRATCH/dbt_local/sql.sh" <<'SQL' +truncate raw.ees_ks4_performance, raw.ees_ks4_info, raw.ees_ks2_attainment, raw.ees_ks2_info, raw.ees_ks4_la; +-- KS4: 2024/25 loaded; 2023/24 as the old-format file left it (keys, no measures) +insert into raw.ees_ks4_performance (school_urn, time_period, breakdown_topic, breakdown, sex, pupil_count, attainment8_average) + select (100000 + g)::text, '202425', 'Total', 'Total', 'Total', '200', '50.0' from generate_series(1, 20) g; +insert into raw.ees_ks4_performance (school_urn, time_period, breakdown_topic, breakdown, sex) + select (100000 + g)::text, '202324', 'Total', 'Total', 'Total' from generate_series(1, 20) g; +-- KS2: results for three years; information for 2023/24 ("34%") and 2024/25 ("32"); none for 2022/23 +insert into raw.ees_ks2_attainment (school_urn, time_period, subject, breakdown_topic, breakdown, expected_standard_pupil_percent) + select (200000 + g)::text, y, 'Reading, writing and maths', 'All pupils', 'Total', '70' + from generate_series(1, 1000) g, (values ('202223'), ('202324'), ('202425')) v(y); +insert into raw.ees_ks2_info (school_urn, time_period, totpups, telig, ptfsm6cla1a) + select (200000 + g)::text, '202324', '300', '45', '34%' from generate_series(1, 1000) g; +insert into raw.ees_ks2_info (school_urn, time_period, totpups, telig, ptfsm6cla1a) + select (200000 + g)::text, '202425', '300', '45', '32' from generate_series(1, 1000) g; +-- LA: nothing loaded +SQL +$SCRATCH/dbt_local/dbt.sh build --no-partial-parse --select stg_ees_ks4 stg_ees_ks2 stg_ees_ks4_la fact_ks4_la_averages assert_ks4_years_have_results assert_ks2_info_percentages_loaded assert_ks4_la_averages_cover_las +``` + +Expected: non-zero exit; `FAIL 1 assert_ks4_years_have_results` (202324), `FAIL 1 assert_ks2_info_percentages_loaded` (202324), `FAIL 1 assert_ks4_la_averages_cover_las` (empty mart). + +- [ ] **Step 3: Load the good data and watch the KS4 and LA tests pass** + +```bash +"$SCRATCH/dbt_local/sql.sh" <<'SQL' +update raw.ees_ks4_performance set pupil_count = '217', attainment8_average = '64.1' where time_period = '202324'; +insert into raw.ees_ks4_la (time_period, old_la_code, la_name, attainment_8_score) + select '202425', (800 + g)::text, 'LA ' || g, '46.0' from generate_series(1, 152) g; +SQL +$SCRATCH/dbt_local/dbt.sh build --no-partial-parse --select stg_ees_ks4 stg_ees_ks2 stg_ees_ks4_la fact_ks4_la_averages assert_ks4_years_have_results assert_ks2_info_percentages_loaded assert_ks4_la_averages_cover_las +``` + +Expected: non-zero exit only because of `FAIL 1 assert_ks2_info_percentages_loaded`; `PASS assert_ks4_years_have_results`, `PASS assert_ks4_la_averages_cover_las`. + +- [ ] **Step 4: Check the LA threshold bites below 145** + +```bash +"$SCRATCH/dbt_local/sql.sh" <<'SQL' +insert into raw.ees_ks4_la (time_period, old_la_code, la_name, attainment_8_score) + select '202526', (800 + g)::text, 'LA ' || g, '46.0' from generate_series(1, 144) g; +SQL +$SCRATCH/dbt_local/dbt.sh build --no-partial-parse --select stg_ees_ks4_la fact_ks4_la_averages assert_ks4_la_averages_cover_las +``` + +Expected: `FAIL 1 assert_ks4_la_averages_cover_las` (202526, 144). Then remove those rows (`echo "delete from raw.ees_ks4_la where time_period = '202526';" | "$SCRATCH/dbt_local/sql.sh"`) and re-run the same build → `PASS`. + +- [ ] **Step 5: Commit** + +```bash +git add pipeline/transform/tests/assert_ks4_years_have_results.sql pipeline/transform/tests/assert_ks2_info_percentages_loaded.sql pipeline/transform/tests/assert_ks4_la_averages_cover_las.sql +git commit -m "test(dbt): fail the build when a published year loads empty (C2, H2)" +``` + +### Task 7: `safe_numeric` reads a trailing `%` + +**Files:** +- Modify: `pipeline/transform/macros/safe_numeric.sql` +- Create: `pipeline/transform/models/staging/stg_ees_ks2.yml` + +- [ ] **Step 1: Write the failing unit test** + +`pipeline/transform/models/staging/stg_ees_ks2.yml`: + +```yaml +version: 2 + +unit_tests: + - name: stg_ees_ks2_reads_percentages_written_with_a_sign + description: > + DfE's 2023/24 KS2 information file writes percentages as "34%", and every + 2023/24 percentage loaded as null (audit C2). Plain numbers still load; + suppression codes stay null. School 147411's 2023/24 figures are DfE's. + model: stg_ees_ks2 + given: + - input: source('raw', 'ees_ks2_attainment') + rows: + - {school_urn: '147411', time_period: '202324', subject: 'Reading, writing and maths', breakdown_topic: 'All pupils', breakdown: 'Total', expected_standard_pupil_percent: '70'} + - {school_urn: '100000', time_period: '202425', subject: 'Reading, writing and maths', breakdown_topic: 'All pupils', breakdown: 'Total', expected_standard_pupil_percent: '79'} + - input: source('raw', 'ees_ks2_info') + rows: + - {school_urn: '147411', time_period: '202324', totpups: '818', telig: '112', ptfsm6cla1a: '34%', ptealgrp2: '56%', psenelk: '20%', psenele: 'c', ptmobn: '87.5%'} + - {school_urn: '100000', time_period: '202425', totpups: '792', telig: '105', ptfsm6cla1a: '32', ptealgrp2: 'x', psenelk: '14', psenele: '2', ptmobn: '87'} + expect: + rows: + - {urn: 147411, year: 202324, total_pupils: 818, rwm_expected_pct: 70, disadvantaged_pct: 34, eal_pct: 56, sen_support_pct: 20, sen_ehcp_pct: null, stability_pct: 87.5} + - {urn: 100000, year: 202425, total_pupils: 792, rwm_expected_pct: 79, disadvantaged_pct: 32, eal_pct: null, sen_support_pct: 14, sen_ehcp_pct: 2, stability_pct: 87} +``` + +- [ ] **Step 2: Run it to see it fail** + +Run: `$SCRATCH/dbt_local/dbt.sh test --no-partial-parse --select "stg_ees_ks2,test_type:unit"` +Expected: `FAIL` — for 147411 the actual `disadvantaged_pct`, `eal_pct`, `sen_support_pct` and `stability_pct` are null. + +- [ ] **Step 3: Change the macro** + +Replace the whole of `pipeline/transform/macros/safe_numeric.sql` with: + +```sql +{# + safe_numeric(col) + Casts a string column to numeric, treating any non-numeric value as NULL. + Handles all EES suppression codes (z, c, x, q, u, etc.) without needing + an explicit list — any string that doesn't look like a number becomes NULL. + A trailing percent sign is accepted: DfE's 2023/24 KS2 information file + writes percentages as "34%" (audit C2). +#} +{% macro safe_numeric(col) -%} + CASE WHEN {{ col }} ~ '^-?[0-9]+(\.[0-9]+)?%?$' THEN rtrim({{ col }}, '%')::numeric ELSE NULL END +{%- endmacro %} +``` + +- [ ] **Step 4: Run the unit test and the KS2 data test** + +Run: `$SCRATCH/dbt_local/dbt.sh test --no-partial-parse --select "stg_ees_ks2,test_type:unit"` +Expected: `PASS`. + +Run: `$SCRATCH/dbt_local/dbt.sh build --no-partial-parse --select stg_ees_ks2 assert_ks2_info_percentages_loaded` +Expected: exit 0 — the Task 6 fixtures still loaded, `PASS assert_ks2_info_percentages_loaded`. + +- [ ] **Step 5: Check nothing else that uses the macro breaks** + +Run: `$SCRATCH/dbt_local/dbt.sh compile --no-partial-parse` +Expected: exit 0 (all 14 models that call `safe_numeric` compile). + +- [ ] **Step 6: Commit** + +```bash +git add pipeline/transform/macros/safe_numeric.sql pipeline/transform/models/staging/stg_ees_ks2.yml +git commit -m "fix(dbt): read percentages DfE writes as \"34%\" (C2)" +``` + +### Task 8: Load the real files end to end and verify PR 1 + +Runs the changed tap against DfE into the local database through the same loader production uses, then the dbt models and tests. It downloads about 2 GB; run it in the background. + +**Files:** none in the repo. + +- [ ] **Step 1: Prepare the catalog and loader config** + +```bash +mkdir -p "$SCRATCH/e2e_load" && cd "$SCRATCH/e2e_load" +TAP=/Users/tudor/projects/school_compare/pipeline/plugins/extractors/tap-uk-ees +echo '{}' > tap.json +uv run --no-project --with "$TAP" tap-uk-ees --config tap.json --discover > catalog.json +python3 - <<'EOF' +import json +want = {"ees_ks2_attainment", "ees_ks2_info", "ees_ks4_performance", "ees_ks4_info", "ees_ks4_national", "ees_ks4_la"} +catalog = json.load(open("catalog.json")) +for stream in catalog["streams"]: + for entry in stream["metadata"]: + if entry["breadcrumb"] == []: + entry["metadata"]["selected"] = stream["tap_stream_id"] in want +json.dump(catalog, open("catalog.json", "w")) +print(sorted(s["tap_stream_id"] for s in catalog["streams"])) +EOF +SOCK=$(grep -o 'PG_HOST=[^ ]*' "$SCRATCH/dbt_local/dbt.sh" | cut -d= -f2) +cat > target.json < run_load.sh <<'EOF' +#!/bin/bash +TAP=/Users/tudor/projects/school_compare/pipeline/plugins/extractors/tap-uk-ees +uv run --no-project --with "$TAP" tap-uk-ees --config tap.json --catalog catalog.json 2> tap.log \ + | uv run --no-project --with meltanolabs-target-postgres target-postgres --config target.json > target.log 2>&1 +echo "exit ${PIPESTATUS[*]}" +EOF +bash run_load.sh +``` + +Expected: `exit 0 0`. `grep "Skipping" tap.log` shows `Skipping 57090 rows for 202324 from release b76a938a-7875-4542-af20-0b23ecb99a49`. If the loader cannot connect to the socket, record a ruling and continue with Step 4 only on staging (Task 9's description must then say so). + +- [ ] **Step 3: Build and test the affected models** + +```bash +$SCRATCH/dbt_local/dbt.sh build --no-partial-parse --select stg_ees_ks2 stg_ees_ks4 stg_ees_ks4_national fact_ks4_national_averages stg_ees_ks4_la fact_ks4_la_averages assert_ks4_years_have_results assert_ks2_info_percentages_loaded assert_ks4_la_averages_cover_las +``` + +Expected: exit 0; all unit, schema and data tests PASS. + +- [ ] **Step 4: Check the values against DfE** + +Run through `"$SCRATCH/dbt_local/sql.sh" <<'SQL' … SQL`: + +```sql +-- C2 results and information: Bishop Stopford School +select year, attainment_8_score, progress_8_score, english_maths_standard_pass_pct, + eligible_pupils, prior_attainment_avg, sen_pct, progress_8_banding, + attainment_8_disadvantage_gap, progress_8_disadvantage_gap +from staging.stg_ees_ks4 where urn = 137086 order by year; +-- C2 KS2 information +select year, total_pupils, disadvantaged_pct, eal_pct, sen_support_pct +from staging.stg_ees_ks2 where urn = 147411 and year = 202324; +-- No old-format rows came in +select count(*) from raw.ees_ks4_performance where time_period = '202324' and pupil_count is null; +-- Coverage per year +select year, count(*), count(attainment_8_score), count(progress_8_score) from staging.stg_ees_ks4 group by year order by year; +-- H2 +select year, count(*) from marts.fact_ks4_la_averages group by year order by year; +select la_code, la_name, attainment_8_score from marts.fact_ks4_la_averages where year = 202425 and la_code in (207, 212, 213); +``` + +Expected: +- 137086: 202223 60.7 / 0.88 / 84; 202324 64.1 / 1.02 / 91.7 with 217, 108.1, 6.5, `Well above average`, 4.6, 0.14; 202425 58.7 / null / 84.2. +- 147411 2023/24: 818, 34, 56, 20. +- old-format rows: `0`. +- coverage: 202223 5657 / 4631 / 3624; 202324 5709 / 4682 / 3663; 202425 5755 / 4736 / 0. +- LA rows per year: 150, 150, 151, 151, 152, 152, 152 (2018/19 → 2024/25). +- 207 Kensington and Chelsea 54.5; 212 Wandsworth 51.8; 213 Westminster 53.2. + +- [ ] **Step 5: Run CI's Python suite** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest backend/tests pipeline/tests scripts/ci/tests -q` +Expected: all pass. + +### Task 9: Open PR 1 + +- [ ] **Step 1:** `git status --short` → only `.DS_Store`. `git log --oneline origin/main..HEAD` → the spec, the plan and Tasks 2–7's commits. +- [ ] **Step 2:** `git push -u origin fix/ks4-2023-24-pipeline`. +- [ ] **Step 3:** Open the PR through the Gitea API with basic auth from `git credential fill` (host `privaterepo.sitaru.org`; never print the secret): `POST /api/v1/repos/tudor/school_compare/pulls`, head `fix/ks4-2023-24-pipeline`, base `main`, title `fix(pipeline): load 2023/24 KS4/KS2 data and DfE LA averages (C2, H2, part 1 of 2)`. Body: + - what changed (the three C2 causes; the LA stream and mart; the data tests), with Task 8's evidence; + - **After merge:** run `school_data_annual_ees` on staging; expected staging API values: `/api/schools/137086` 2023/24 Attainment 8 64.1, Progress 8 1.02; `/api/schools/147411` 2023/24 disadvantaged 34; + - **Optional cleanup** (production and staging, once): `delete from raw.ees_ks4_performance where time_period = '202324' and pupil_count is null;` — removes the 34,254 old-label rows the re-issued file left; new-format rows always have `pupil_count`; + - PR 2 (site) must merge only after the staging EES run, and be promoted only after the production EES run; + - ends with the Claude Code line. +- [ ] **Step 4:** Report the PR URL. Do not merge. + +--- + +## PR 2 — site + +Start from an up-to-date main: `git fetch origin && git switch -c fix/la-averages-site origin/main`. + +### Task 10: `/api/la-averages` serves DfE's figures + +**Files:** +- Modify: `backend/models.py` (after `Ks4NationalAverage`) +- Modify: `backend/app.py` (`get_la_averages`) +- Test: `backend/tests/test_la_averages_marts.py` + +**Interfaces:** +- Consumes: `marts.fact_ks4_la_averages` (`year`, `la_code`, `la_name`, `attainment_8_score`). +- Produces: `Ks4LaAverage` ORM model; `_la_averages_payload(df: pd.DataFrame) -> dict` with the endpoint's unchanged shape. + +- [ ] **Step 1: Write the failing tests** + +`backend/tests/test_la_averages_marts.py`: + +```python +"""/api/la-averages serves DfE's own LA averages (fact_ks4_la_averages). + +It used to average the dataframe: every school with an Attainment 8, +independent and special schools included. Kensington and Chelsea came out at +35.2 against DfE's 54.5, and most LAs about 7 points low (audit H2). +""" + +import numpy as np +import pandas as pd +import pytest +from fastapi.testclient import TestClient + +LATEST = 202425 + + +def _df(): + return pd.DataFrame([ + # A state school and an independent: their mean, 37.65, is not DfE's figure. + dict(year=LATEST, local_authority="Kensington and Chelsea", local_authority_code=207, attainment_8_score=54.9), + dict(year=LATEST, local_authority="Kensington and Chelsea", local_authority_code=207, attainment_8_score=20.4), + dict(year=LATEST, local_authority="Bristol, City of", local_authority_code=801, attainment_8_score=45.0), + dict(year=LATEST, local_authority="West Sussex", local_authority_code=938, attainment_8_score=48.0), + # DfE publishes no LA figure for City of London. + dict(year=LATEST, local_authority="City of London", local_authority_code=201, attainment_8_score=30.0), + # A newer year with primary results only. + dict(year=202526, local_authority="Kensington and Chelsea", local_authority_code=207, attainment_8_score=np.nan), + ]) + + +class _Row: + def __init__(self, year, la_code, la_name, attainment_8_score): + self.year = year + self.la_code = la_code + self.la_name = la_name + self.attainment_8_score = attainment_8_score + + +class _StubSession: + rows = [ + _Row(LATEST, 207, "Kensington and Chelsea", 54.5), + _Row(LATEST, 801, "Bristol City", 46.3), # DfE's spelling, not ours + _Row(LATEST, 938, "West Sussex", None), # suppressed + _Row(LATEST, 330, "Birmingham", 44.0), # no school of ours there + _Row(202324, 207, "Kensington and Chelsea", 54.5), + ] + + def query(self, model): + assert model.__name__ == "Ks4LaAverage" + return self + + def filter(self, condition): + # The payload filters on year == ; apply it as Postgres would. + self._year = condition.right.value + return self + + def all(self): + return [r for r in self.rows if r.year == self._year] + + def rollback(self): + pass + + def close(self): + pass + + +class _OldYearOnly(_StubSession): + rows = [_Row(202324, 207, "Kensington and Chelsea", 54.5)] + + +class _NoMart(_StubSession): + def all(self): + raise RuntimeError('relation "marts.fact_ks4_la_averages" does not exist') + + +@pytest.fixture() +def payload(monkeypatch): + from backend import app as app_module + from backend import database as database_module + + def _run(session_cls): + monkeypatch.setattr(database_module, "SessionLocal", session_cls) + return app_module._la_averages_payload(_df()) + + return _run + + +def test_serves_dfe_figures_keyed_by_our_la_names(payload): + out = payload(_StubSession) + assert out["year"] == LATEST + assert out["secondary"]["attainment_8_by_la"] == { + "Kensington and Chelsea": 54.5, + "Bristol, City of": 46.3, + } + + +def test_an_la_without_a_dfe_figure_is_absent(payload): + by_la = payload(_StubSession)["secondary"]["attainment_8_by_la"] + assert "City of London" not in by_la + assert "West Sussex" not in by_la + assert None not in by_la.values() + + +def test_no_dfe_figures_for_the_year_give_an_empty_map(payload): + assert payload(_OldYearOnly) == {"year": LATEST, "secondary": {"attainment_8_by_la": {}}} + + +def test_a_missing_mart_gives_an_empty_map(payload): + assert payload(_NoMart)["secondary"]["attainment_8_by_la"] == {} + + +def test_the_endpoint_serves_dfe_figures(monkeypatch): + from backend import app as app_module + from backend import database as database_module + + monkeypatch.setattr(app_module, "load_school_data", _df) + monkeypatch.setattr(database_module, "SessionLocal", _StubSession) + resp = TestClient(app_module.app).get("/api/la-averages") + assert resp.status_code == 200, resp.text + assert resp.json()["secondary"]["attainment_8_by_la"]["Kensington and Chelsea"] == 54.5 +``` + +- [ ] **Step 2: Run them to see them fail** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest backend/tests/test_la_averages_marts.py -q` +Expected: 4 failures on `_la_averages_payload` (AttributeError), and the endpoint test failing with `KeyError: 'Kensington and Chelsea'` — today's code takes the newest year (202526), which has no Attainment 8, and returns an empty map. + +- [ ] **Step 3: Add the ORM model** + +In `backend/models.py`, after `class Ks4NationalAverage`: + +```python +class Ks4LaAverage(Base): + """Official DfE KS4 local-authority averages (all state-funded schools) — + one row per academic year and LA. la_code is the GIAS LA code.""" + __tablename__ = "fact_ks4_la_averages" + __table_args__ = MARTS + + year = Column(Integer, primary_key=True) + la_code = Column(Integer, primary_key=True) + la_name = Column(String) + attainment_8_score = Column(Float) +``` + +- [ ] **Step 4: Replace the endpoint** + +In `backend/app.py`, replace the whole `get_la_averages` function (decorators included) with: + +```python +def _la_averages_payload(df: pd.DataFrame) -> dict: + """Per-LA Attainment 8 for the "vs LA avg" comparison: DfE's own LA + averages (fact_ks4_la_averages, all state-funded schools), never a mean of + the dataframe. That mean counted independent and special schools and put + most LAs about 7 points low (audit H2). + + The year is the latest with any school Attainment 8, so the average and + the scores set against it are the same year. Figures are keyed by our LA + name through the LA code. An LA without a DfE figure for that year (City of + London) is absent; no figures for the year, or no mart, give an empty map, + so rows show no comparison, never another year's figure. + """ + empty = {"year": 0, "secondary": {"attainment_8_by_la": {}}} + if df.empty or "attainment_8_score" not in df.columns: + return empty + scored = df[df["attainment_8_score"].notna()] + if scored.empty: + return empty + year = int(scored["year"].max()) + + la = (df[["local_authority_code", "local_authority"]] + .dropna() + .drop_duplicates("local_authority_code")) + name_by_code = {int(code): name for code, name in + zip(la["local_authority_code"], la["local_authority"])} + + from . import database + from .models import Ks4LaAverage + + rows: list = [] + db = None + try: + db = database.SessionLocal() + rows = db.query(Ks4LaAverage).filter(Ks4LaAverage.year == year).all() + except Exception: + logging.getLogger(__name__).warning( + "DfE LA averages unavailable for %s", year, exc_info=True) + if db is not None: + db.rollback() + finally: + if db is not None: + db.close() + + by_la = { + name_by_code[row.la_code]: row.attainment_8_score + for row in rows + if row.attainment_8_score is not None and row.la_code in name_by_code + } + return {"year": year, "secondary": {"attainment_8_by_la": by_la}} + + +@app.get("/api/la-averages") +@limiter.limit(f"{settings.rate_limit_per_minute}/minute") +async def get_la_averages(request: Request): + """DfE's per-LA Attainment 8 averages for the latest year with results.""" + return _la_averages_payload(load_school_data()) +``` + +- [ ] **Step 5: Run the tests and the whole backend suite** + +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest backend/tests/test_la_averages_marts.py -q` → `5 passed`. +Run: `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest backend/tests -q` → all pass. + +- [ ] **Step 6: Commit** + +```bash +git add backend/models.py backend/app.py backend/tests/test_la_averages_marts.py +git commit -m "fix(api): serve DfE's LA averages for \"vs LA avg\" (H2)" +``` + +### Task 11: No LA gap for independent schools + +**Files:** +- Modify: `nextjs-app/lib/utils.ts` (after `isSpecialSchool`) +- Modify: `nextjs-app/components/SecondarySchoolRow.tsx` +- Modify: `nextjs-app/components/LeafletMapInner.tsx` +- Test: `nextjs-app/__tests__/lib/utils.test.ts`, `nextjs-app/__tests__/components/SecondarySchoolRow.test.tsx`, `nextjs-app/__tests__/components/LeafletMapInner.test.tsx` + +**Interfaces:** +- Produces: `isIndependentSchool(school: { school_type?: string | null }): boolean`. + +- [ ] **Step 1: Write the failing tests** + +Append to `__tests__/lib/utils.test.ts`: + +```ts +describe('isIndependentSchool', () => { + const { isIndependentSchool } = require('@/lib/utils'); + + it('matches both GIAS independent types', () => { + expect(isIndependentSchool({ school_type: 'Other independent school' })).toBe(true); + expect(isIndependentSchool({ school_type: 'Other independent special school' })).toBe(true); + }); + + it('does not match state-funded types or a missing type', () => { + for (const t of ['Academy converter', 'Community school', 'Free schools', 'Non-maintained special school']) { + expect(isIndependentSchool({ school_type: t })).toBe(false); + } + expect(isIndependentSchool({ school_type: null })).toBe(false); + expect(isIndependentSchool({})).toBe(false); + }); +}); +``` + +Append to `__tests__/components/SecondarySchoolRow.test.tsx`: + +```tsx +describe('SecondarySchoolRow LA comparison', () => { + it('compares a state school with its LA average', () => { + render( + , + ); + expect(screen.getByText(/\+5\.5 vs LA avg/)).toBeInTheDocument(); + }); + + it('keeps the comparison for a school whose type is unknown', () => { + render( + , + ); + expect(screen.getByText(/\+5\.5 vs LA avg/)).toBeInTheDocument(); + }); + + it("shows an independent school's Attainment 8 without an LA comparison", () => { + // DfE's LA average covers state-funded schools; an independent's + // Attainment 8 leaves out IGCSEs (audit H2). + render( + , + ); + expect(screen.getByText('20.4')).toBeInTheDocument(); + expect(screen.queryByText(/vs LA avg/)).not.toBeInTheDocument(); + }); +}); +``` + +Append to `__tests__/components/LeafletMapInner.test.tsx`: + +```tsx +it('compares a state secondary with its LA average on the card, but not an independent one', () => { + const state: School = { ...base, urn: 3, school_name: 'Holland Park School', school_type: 'Academy converter', + phase: 'Secondary', local_authority: 'Kensington and Chelsea', attainment_8_score: 60, + latitude: 51.5, longitude: -0.2, distance: 0.3 }; + const independent: School = { ...state, urn: 4, school_name: 'Abbey Gate College', + school_type: 'Other independent school', attainment_8_score: 20.4 }; + const laAverages = { 'Kensington and Chelsea': 54.5 }; + + const { container, rerender } = renderMap({ schools: [state, independent], laAverages, selectedUrn: 3 }); + expect(container.querySelector('.sc-popup')).toHaveTextContent('60.0 Att 8 +5.5 vs LA'); + + rerender({ schools: [state, independent], laAverages, selectedUrn: 4 }); + const card = container.querySelector('.sc-popup')!; + expect(card).toHaveTextContent('20.4 Att 8'); + expect(card).not.toHaveTextContent(/vs LA/); +}); +``` + +- [ ] **Step 2: Run them to see them fail** + +Run: `cd nextjs-app && npx jest __tests__/lib/utils.test.ts __tests__/components/SecondarySchoolRow.test.tsx __tests__/components/LeafletMapInner.test.tsx` +Expected: the `isIndependentSchool` tests fail (`isIndependentSchool is not a function`); the independent row and independent card tests fail (a "vs LA" comparison is shown); the state-school and unknown-type tests pass already. + +- [ ] **Step 3: Implement** + +In `lib/utils.ts`, after `isSpecialSchool`: + +```ts +/** + * Independent (fee-paying) schools: GIAS types "Other independent school" and + * "Other independent special school". DfE's Attainment 8 for them leaves out + * IGCSEs, and DfE's LA averages cover state-funded schools only, so callers + * drop the "vs LA avg" comparison for them (audit H2). + */ +export function isIndependentSchool(school: { school_type?: string | null }): boolean { + return /\bindependent\b/i.test(school.school_type ?? ''); +} +``` + +In `components/SecondarySchoolRow.tsx`, add `isIndependentSchool` to the `@/lib/utils` import and replace the comment and `laDelta`: + +```tsx + // The school's own Attainment 8 is a same-school figure — shown whenever it + // exists (special schools included; their type tag on line 2 gives context). + // Only the vs-LA-average delta, a benchmark comparison, is dropped: for + // special schools / PRUs / AP, whose pupils aren't measured against it + // fairly, and for independent schools, because DfE's LA average covers + // state-funded schools and an independent's Attainment 8 leaves out IGCSEs. + const laDelta = + att8 != null && !isSpecialSchool(school) && !isIndependentSchool(school) && laAvgAttainment8 != null + ? att8 - laAvgAttainment8 + : null; +``` + +In `components/LeafletMapInner.tsx`, add `isIndependentSchool` to the `@/lib/utils` import, and in `metricHtml` change `if (!special && laAvg != null) {` to: + +```ts + if (!special && !isIndependentSchool(school) && laAvg != null) { +``` + +- [ ] **Step 4: Run the tests, the whole suite and the type check** + +Run: `npx jest __tests__/lib/utils.test.ts __tests__/components/SecondarySchoolRow.test.tsx __tests__/components/LeafletMapInner.test.tsx` → all pass. +Run: `npm run typecheck && npm test` → all pass. + +- [ ] **Step 5: Commit** + +```bash +git add nextjs-app/lib/utils.ts nextjs-app/components/SecondarySchoolRow.tsx nextjs-app/components/LeafletMapInner.tsx nextjs-app/__tests__/lib/utils.test.ts nextjs-app/__tests__/components/SecondarySchoolRow.test.tsx nextjs-app/__tests__/components/LeafletMapInner.test.tsx +git commit -m "fix(search): no LA comparison for independent schools (H2)" +``` + +### Task 12: E2E journeys + +**Files:** +- Modify: `e2e/tests/journeys.spec.ts` + +- [ ] **Step 1: Keep the existing LA journey on a state school** + +In `test('a secondary search row compares its Attainment 8 with the LA average', …)`, change the school filter's last line to exclude independents: + +```ts + && !/special|pupil referral|alternative provision|independent/i.test(s.school_type ?? '')); +``` + +- [ ] **Step 2: Add the two journeys after it** + +```ts +test('an independent secondary shows its Attainment 8 without an LA comparison', async ({ page }) => { + // DfE's LA averages cover state-funded schools, and an independent school's + // Attainment 8 leaves out IGCSEs, so a gap would mislead (audit H2). + const res = await page.request.get('/api/schools?school_type=independent&phase=secondary&page_size=50'); + expect(res.ok()).toBeTruthy(); + const school = ((await res.json()).schools ?? []).find( + (s: { attainment_8_score?: number | null; school_type?: string }) => + s.attainment_8_score != null && !/special/i.test(s.school_type ?? '')); + test.skip(!school, 'no independent secondary with an Attainment 8 here'); + + const averagesLoaded = page.waitForResponse(r => r.url().includes('/la-averages')); + await searchByName(page, school.school_name); + await averagesLoaded; + const link = page.locator(`a[href^="/school/${school.urn}-"]`).first(); + await expect(link).toBeVisible({ timeout: 15_000 }); + const stats = link.locator('xpath=ancestor::div[contains(@class, "__rowContent")][1]') + .locator('[class*="__line3"]'); + await expect(stats.getByText(school.attainment_8_score.toFixed(1))).toBeVisible(); + await expect(stats.getByText(/vs LA avg/)).toHaveCount(0); +}); + +test('a secondary shows its 2023/24 GCSE results, the last year DfE published Progress 8', async ({ page }) => { + // DfE re-issued its 2023/24 file under older column names, and every + // school's 2023/24 row loaded empty (audit C2). These are DfE's final + // figures for Bishop Stopford School, so they do not change. + await page.goto('/school/137086'); + const history = page.locator('#history'); + await history.getByText('View raw year-by-year data').click(); + const row = history.getByRole('row', { name: /2023\/24/ }); + await expect(row).toContainText('64.1'); + await expect(row).toContainText('+1.0'); + await expect(row).toContainText('91.7%'); +}); +``` + +- [ ] **Step 3: Check the file compiles and lists the new tests** + +Run: `cd e2e && npx playwright test --list tests/journeys.spec.ts -g "independent secondary|2023/24 GCSE"` → no compile errors; exactly the two new tests are listed. (`e2e/` has no tsconfig; Playwright compiles the spec itself.) + +- [ ] **Step 4: Run them against staging to see the right ones fail** + +Only once PR 1 has merged and `school_data_annual_ees` has run on staging (ask Tudor; the staging API must show `/api/schools/137086` 2023/24 Attainment 8 64.1 first): + +Run: `cd e2e && BASE_URL=https://stx.schoolcompare.co.uk npx playwright test tests/journeys.spec.ts -g "independent secondary|2023/24 GCSE|compares its Attainment 8"` +Expected: the 2023/24 journey PASSES (data loaded, no UI change needed); the independent journey FAILS (staging still runs the old backend and row, which show a gap); the existing LA journey PASSES. If the EES run has not happened yet, record that and rely on the post-merge staging gate. + +- [ ] **Step 5: Commit** + +```bash +git add e2e/tests/journeys.spec.ts +git commit -m "test(e2e): 2023/24 results and no LA gap for independents (C2, H2)" +``` + +### Task 13: Verify and open PR 2 + +- [ ] **Step 1:** CI's checks locally — `cd nextjs-app && npm run typecheck && npm test`; from the repo root, `uv run --no-project --with-requirements requirements.txt python -c "from backend.app import app; print('backend imports OK')"` and `uv run --no-project --with-requirements requirements.txt --with pytest --with "httpx<0.28" --with pyyaml python -m pytest backend/tests pipeline/tests scripts/ci/tests -q`. Expected: all pass. +- [ ] **Step 2:** `git status --short` → only `.DS_Store`. `git push -u origin fix/la-averages-site`. +- [ ] **Step 3:** Open the PR through the Gitea API (as Task 9): head `fix/la-averages-site`, base `main`, title `fix(site): DfE LA averages for "vs LA avg"; no gap for independents (H2, part 2 of 2)`. Body: what changed; Task 12 Step 4's result; **Merge only after `school_data_annual_ees` has run on staging with PR 1; promote only after it has run on production** (before that the endpoint returns an empty map and rows show no gap); ends with the Claude Code line. +- [ ] **Step 4:** Report the PR URL. Do not merge. + +--- + +## After both PRs are live in production (Tudor promotes) + +Validate against DfE's files, as in the audit: every listed school's 2023/24 Attainment 8, Progress 8 and English and maths figures match `202425_performance_tables_schools_final.csv`; the 2023/24 KS4 and KS2 information fields match their files; `/api/la-averages` equals DfE for all 152 LAs; independent rows show no gap. Then mark C2 and H2 resolved in the audit report artifact. -- 2.54.0 From 4db1131d0fb440b2bcebde87499dd581f25a3408 Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 10:45:14 +0100 Subject: [PATCH 03/11] fix(ees): newest KS4 release owns every year it contains (C2) Co-Authored-By: Claude Opus 5.5 --- .../tap_uk_ees/release_precedence.py | 44 +++++++++++ .../extractors/tap-uk-ees/tap_uk_ees/tap.py | 20 +++++ pipeline/tests/test_ees_release_precedence.py | 78 +++++++++++++++++++ 3 files changed, 142 insertions(+) create mode 100644 pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py create mode 100644 pipeline/tests/test_ees_release_precedence.py diff --git a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py new file mode 100644 index 0000000..52080c6 --- /dev/null +++ b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py @@ -0,0 +1,44 @@ +"""Which release a row comes from when DfE re-publishes a year. + +DfE re-publishes earlier years inside later releases: the 2024/25 KS4 results +file holds 2022/23, 2023/24 and 2024/25. The 2023/24 release's own file, +re-issued on 10 March 2026 under older column names, was read after it and +overwrote every 2023/24 row with blanks (audit C2). A stream that opts in +treats the newest release as the authority for every year it contains. + +Free of the Singer SDK so CI's pytest, which installs only the backend's +requirements, can load it. +""" +from __future__ import annotations + +import pandas as pd + + +def newest_first(releases: list[dict]) -> list[dict]: + """Releases by time_period, newest first. A release whose time_period is + unknown goes last, in the order given.""" + dated = [r for r in releases if r.get("time_period")] + undated = [r for r in releases if not r.get("time_period")] + return sorted(dated, key=lambda r: r["time_period"], reverse=True) + undated + + +def periods_in(df: pd.DataFrame) -> set[str]: + """The years a release's rows cover.""" + if "time_period" not in df.columns: + return set() + return set(df["time_period"].astype(str).str.strip()) - {""} + + +def drop_owned_periods( + df: pd.DataFrame, owned: set[str] +) -> tuple[pd.DataFrame, dict[str, int]]: + """Drop the rows for years a newer release already supplied. + + Returns the rows kept and, for each year dropped, how many rows went. + """ + if "time_period" not in df.columns or not owned: + return df, {} + periods = df["time_period"].astype(str).str.strip() + dropped = periods.isin(owned) + skipped = {str(k): int(v) for k, v in periods[dropped].value_counts().items()} + return df[~dropped], skipped diff --git a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py index b8bdf53..1b399df 100644 --- a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py +++ b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py @@ -17,6 +17,8 @@ import requests from singer_sdk import Stream, Tap from singer_sdk import typing as th +from tap_uk_ees.release_precedence import drop_owned_periods, newest_first, periods_in + CONTENT_API_BASE = ( "https://content.explore-education-statistics.service.gov.uk/api" ) @@ -88,6 +90,8 @@ class EESDatasetStream(Stream): target CSV path inside the ZIP (substring match, not exact). Subclasses may set _column_renames to map messy CSV column names to clean Singer field names before yielding records. + Subclasses may set _newest_release_owns_period when DfE re-publishes + earlier years in later releases and the newest copy is the authority. """ replication_key = None @@ -96,6 +100,7 @@ class EESDatasetStream(Stream): _urn_column: str = "school_urn" # column name for URN in the CSV _encoding: str = "utf-8" # CSV file encoding (some DfE files use latin-1) _column_renames: dict = {} # CSV column name → Singer field name + _newest_release_owns_period: bool = False # see release_precedence.py def get_records(self, context): import pandas as pd @@ -110,6 +115,9 @@ class EESDatasetStream(Stream): self.logger.info( "Found %d release(s) for %s", len(releases), self._publication_slug ) + if self._newest_release_owns_period: + releases = newest_first(releases) + owned_periods: set[str] = set() for release in releases: release_id = release["id"] @@ -163,6 +171,15 @@ class EESDatasetStream(Stream): if urn_col in df.columns: df = df[df[urn_col].notna() & (df[urn_col] != "")] + if self._newest_release_owns_period: + df, skipped = drop_owned_periods(df, owned_periods) + for period, count in sorted(skipped.items()): + self.logger.info( + "Skipping %d rows for %s from release %s: a newer release supplied that year", + count, period, release_id, + ) + owned_periods |= periods_in(df) + self.logger.info("Emitting %d school-level rows from release %s", len(df), release_id) for _, row in df.iterrows(): @@ -251,6 +268,9 @@ class EESKS4PerformanceStream(EESDatasetStream): primary_keys = ["school_urn", "time_period", "breakdown_topic", "breakdown", "sex"] _publication_slug = "key-stage-4-performance" _target_filename = "performance_tables_schools" + # DfE's 2024/25 file re-publishes 2022/23 and 2023/24 under current names; + # the 2023/24 release's own file uses older ones (audit C2). + _newest_release_owns_period = True schema = th.PropertiesList( th.Property("time_period", th.StringType, required=True), th.Property("school_urn", th.StringType, required=True), diff --git a/pipeline/tests/test_ees_release_precedence.py b/pipeline/tests/test_ees_release_precedence.py new file mode 100644 index 0000000..09ce3ae --- /dev/null +++ b/pipeline/tests/test_ees_release_precedence.py @@ -0,0 +1,78 @@ +"""DfE re-publishes earlier years inside later KS4 releases. + +The 2024/25 results file holds 2022/23, 2023/24 and 2024/25 under current +column names. The 2023/24 release's own file, re-issued in March 2026 under +older names, was read after it and overwrote every 2023/24 row with blanks +(audit C2). For the KS4 results stream the newest release owns every year it +contains. +""" +import importlib.util +import re +from pathlib import Path + +import pandas as pd +import pytest + +TAP_DIR = (Path(__file__).resolve().parents[1] / 'plugins' / 'extractors' / 'tap-uk-ees' + / 'tap_uk_ees') + + +@pytest.fixture +def precedence(): + spec = importlib.util.spec_from_file_location( + 'release_precedence', TAP_DIR / 'release_precedence.py') + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def _release(period): + return {'id': f'release-{period}', 'time_period': period} + + +def test_releases_are_taken_newest_first_whatever_order_the_api_gives(precedence): + releases = [_release('202223'), _release('202425'), _release(None), _release('202324')] + ordered = precedence.newest_first(releases) + assert [r['time_period'] for r in ordered] == ['202425', '202324', '202223', None] + + +def test_a_year_a_newer_release_supplied_is_dropped_from_an_older_one(precedence): + newer = pd.DataFrame({'time_period': ['202425', '202324', '202223'], 'school_urn': ['1'] * 3}) + older = pd.DataFrame({'time_period': ['202324', '202324', '201920'], 'school_urn': ['1', '2', '1']}) + + kept, skipped = precedence.drop_owned_periods(older, precedence.periods_in(newer)) + + assert list(kept['time_period']) == ['201920'] + assert skipped == {'202324': 2} + + +def test_periods_match_despite_surrounding_spaces(precedence): + owned = precedence.periods_in(pd.DataFrame({'time_period': [' 202324 ']})) + kept, skipped = precedence.drop_owned_periods(pd.DataFrame({'time_period': ['202324']}), owned) + assert owned == {'202324'} + assert kept.empty + assert skipped == {'202324': 1} + + +def test_nothing_is_dropped_before_any_year_is_owned(precedence): + df = pd.DataFrame({'time_period': ['202324'], 'school_urn': ['1']}) + kept, skipped = precedence.drop_owned_periods(df, set()) + assert kept.equals(df) + assert skipped == {} + + +def test_a_file_without_time_period_is_left_alone(precedence): + df = pd.DataFrame({'school_urn': ['1']}) + kept, skipped = precedence.drop_owned_periods(df, {'202324'}) + assert kept.equals(df) + assert skipped == {} + assert precedence.periods_in(df) == set() + + +def test_only_the_ks4_results_stream_opts_in(): + # A general rule would wipe KS2: the 2024/25 KS2 file holds 98,448 of the + # 955,956 rows the 2023/24 release has for 2023/24. + source = (TAP_DIR / 'tap.py').read_text() + opted_in = [chunk.split('(')[0] for chunk in source.split('\nclass ')[1:] + if re.search(r'_newest_release_owns_period\s*=\s*True', chunk)] + assert opted_in == ['EESKS4PerformanceStream'] -- 2.54.0 From 967b1f0eed316ce22a779dcae69954f1b81eb385 Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 10:45:57 +0100 Subject: [PATCH 04/11] fix(ees): read 2023/24 KS4 school information under DfE's older names (C2) Co-Authored-By: Claude Opus 5.5 --- .../tap-uk-ees/tap_uk_ees/ks4_info.py | 47 +++++++++++++ .../extractors/tap-uk-ees/tap_uk_ees/tap.py | 25 ++----- pipeline/tests/test_ees_ks4_info_names.py | 67 +++++++++++++++++++ 3 files changed, 120 insertions(+), 19 deletions(-) create mode 100644 pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_info.py create mode 100644 pipeline/tests/test_ees_ks4_info_names.py diff --git a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_info.py b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_info.py new file mode 100644 index 0000000..eff6b94 --- /dev/null +++ b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_info.py @@ -0,0 +1,47 @@ +"""KS4 school information: the fields the stream declares, and the older +names DfE used for them. + +2023/24 information exists only in the 2023/24 release, whose +202324_information_about_schools_final.csv (re-issued 10 March 2026) uses the +older names. Without the renames every 2023/24 field loaded as null (audit C2). +Newer files contain none of the older names, so the renames leave them alone. + +Free of the Singer SDK so CI's pytest can load it. +""" + +# Declared Singer fields besides the required time_period and school_urn. +KS4_INFO_FIELDS = ( + "school_laestab", + "school_name", + "establishment_type_group", + "reldenom", + "admpol_pt", + "egender", + "agerange", + "allks_pupil_count", + "allks_boys_count", + "allks_girls_count", + "endks4_pupil_count", + "ks2_scaledscore_average", + "sen_with_ehcp_pupil_percent", + "sen_pupil_percent", + "sen_no_ehcp_pupil_percent", + "attainment8_diffn", + "progress8_diffn", + "progress8_banding", +) + +# 2023/24 column name → declared field. +KS4_INFO_RENAMES = { + "t_allks_pupils": "allks_pupil_count", + "t_allks_boys": "allks_boys_count", + "t_allks_girls": "allks_girls_count", + "t_pupils": "endks4_pupil_count", + "avg_ks2_scaledscore": "ks2_scaledscore_average", + "pt_sen_with_ehcp": "sen_with_ehcp_pupil_percent", + "pt_sen": "sen_pupil_percent", + "pt_sen_no_ehcp": "sen_no_ehcp_pupil_percent", + "diffn_att8": "attainment8_diffn", + "diffn_p8mea": "progress8_diffn", + "p8_banding": "progress8_banding", +} diff --git a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py index 1b399df..cb7816a 100644 --- a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py +++ b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py @@ -17,6 +17,7 @@ import requests from singer_sdk import Stream, Tap from singer_sdk import typing as th +from tap_uk_ees.ks4_info import KS4_INFO_FIELDS, KS4_INFO_RENAMES from tap_uk_ees.release_precedence import drop_owned_periods, newest_first, periods_in CONTENT_API_BASE = ( @@ -330,34 +331,20 @@ class EESKS4PerformanceStream(EESDatasetStream): # ── KS4 Information (wide format: one row per school, context/demographics) ── -# File: 202425_information_about_schools_provisional.csv (38 cols) +# Files: 202425_information_about_schools_final.csv (38 cols, current names); +# 202324_information_about_schools_final.csv (60 cols, older names — the only +# source of 2023/24 information). Field list and renames: ks4_info.py. class EESKS4InfoStream(EESDatasetStream): name = "ees_ks4_info" primary_keys = ["school_urn", "time_period"] _publication_slug = "key-stage-4-performance" _target_filename = "information_about_schools" + _column_renames = KS4_INFO_RENAMES schema = th.PropertiesList( th.Property("time_period", th.StringType, required=True), th.Property("school_urn", th.StringType, required=True), - th.Property("school_laestab", th.StringType), - th.Property("school_name", th.StringType), - th.Property("establishment_type_group", th.StringType), - th.Property("reldenom", th.StringType), - th.Property("admpol_pt", th.StringType), - th.Property("egender", th.StringType), - th.Property("agerange", th.StringType), - th.Property("allks_pupil_count", th.StringType), - th.Property("allks_boys_count", th.StringType), - th.Property("allks_girls_count", th.StringType), - th.Property("endks4_pupil_count", th.StringType), - th.Property("ks2_scaledscore_average", th.StringType), - th.Property("sen_with_ehcp_pupil_percent", th.StringType), - th.Property("sen_pupil_percent", th.StringType), - th.Property("sen_no_ehcp_pupil_percent", th.StringType), - th.Property("attainment8_diffn", th.StringType), - th.Property("progress8_diffn", th.StringType), - th.Property("progress8_banding", th.StringType), + *[th.Property(field, th.StringType) for field in KS4_INFO_FIELDS], ).to_dict() diff --git a/pipeline/tests/test_ees_ks4_info_names.py b/pipeline/tests/test_ees_ks4_info_names.py new file mode 100644 index 0000000..ccdb102 --- /dev/null +++ b/pipeline/tests/test_ees_ks4_info_names.py @@ -0,0 +1,67 @@ +"""KS4 school information for 2023/24 exists only in the 2023/24 release, +whose file uses DfE's older column names. The stream declared only the newer +ones, so every 2023/24 field loaded as null (audit C2). The headers below are +DfE's, copied from 202324_information_about_schools_final.csv and +202425_information_about_schools_final.csv. +""" +import importlib.util +from pathlib import Path + +import pytest + +MODULE = (Path(__file__).resolve().parents[1] / 'plugins' / 'extractors' / 'tap-uk-ees' + / 'tap_uk_ees' / 'ks4_info.py') + +HEADER_2023_24 = ( + 'time_period', 'time_identifier', 'geographic_level', 'country_code', 'country_name', + 'school_laestab', 'school_urn', 'school_name', 'old_la_code', 'new_la_code', 'la_name', + 'version', 'establishment_type_group', 'full_address', 'telnum', 'pcon_code', 'pcon_name', + 'contflag', 'iclose', 'reldenom', 'admpol_pt', 'egender', 'feeder', 'agerange', + 't_allks_pupils', 't_allks_boys', 't_allks_girls', 't_pupils', 't_boys', 'pt_boys', + 't_girls', 'pt_girls', 'avg_ks2_scaledscore', 't_prior_lo', 'pt_prior_lo', 't_prior_av', + 'pt_prior_av', 't_prior_hi', 'pt_prior_hi', 't_disadvantaged', 'pt_disadvantaged', + 't_not_disadvantaged', 'pt_not_disadvantaged', 't_language_not_english', + 'pt_language_not_english', 't_language_english', 'pt_language_english', + 't_language_unknown', 'pt_language_unknown', 't_not_mobile', 'pt_not_mobile', + 't_sen_with_ehcp', 'pt_sen_with_ehcp', 't_sen', 'pt_sen', 't_sen_no_ehcp', + 'pt_sen_no_ehcp', 'diffn_att8', 'diffn_p8mea', 'p8_banding', +) + +HEADER_2024_25 = ( + 'time_period', 'time_identifier', 'geographic_level', 'country_code', 'country_name', + 'school_laestab', 'school_urn', 'school_name', 'old_la_code', 'new_la_code', 'la_name', + 'version', 'establishment_type_group', 'full_address', 'telnum', 'pcon_code', 'pcon_name', + 'contflag', 'iclose', 'reldenom', 'admpol_pt', 'egender', 'feeder', 'agerange', + 'allks_pupil_count', 'allks_boys_count', 'allks_girls_count', 'endks4_pupil_count', + 'ks2_scaledscore_average', 'sen_with_ehcp_pupil_count', 'sen_with_ehcp_pupil_percent', + 'sen_pupil_count', 'sen_pupil_percent', 'sen_no_ehcp_pupil_count', + 'sen_no_ehcp_pupil_percent', 'attainment8_diffn', 'progress8_diffn', 'progress8_banding', +) + + +@pytest.fixture +def ks4_info(): + spec = importlib.util.spec_from_file_location('ks4_info', MODULE) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def test_every_declared_field_is_in_the_2023_24_file_once_renamed(ks4_info): + renamed = {ks4_info.KS4_INFO_RENAMES.get(c, c) for c in HEADER_2023_24} + assert set(ks4_info.KS4_INFO_FIELDS) - renamed == set() + + +def test_every_declared_field_is_in_the_2024_25_file(ks4_info): + assert set(ks4_info.KS4_INFO_FIELDS) - set(HEADER_2024_25) == set() + + +def test_the_renames_cannot_collide_with_either_file(ks4_info): + # An old name in the current file, or a new name already in the old file, + # would let a rename overwrite a real column. + assert set(ks4_info.KS4_INFO_RENAMES) & set(HEADER_2024_25) == set() + assert set(ks4_info.KS4_INFO_RENAMES.values()) & set(HEADER_2023_24) == set() + + +def test_every_rename_names_a_declared_field(ks4_info): + assert set(ks4_info.KS4_INFO_RENAMES.values()) <= set(ks4_info.KS4_INFO_FIELDS) -- 2.54.0 From dbb60acff327e6423cad14bf5f61c6eb05f064df Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 10:46:53 +0100 Subject: [PATCH 05/11] feat(ees): keep DfE's KS4 LA averages from the summary data set (H2) Co-Authored-By: Claude Opus 5.5 --- .../tap-uk-ees/tap_uk_ees/ks4_summary.py | 56 +++++++++++ .../extractors/tap-uk-ees/tap_uk_ees/tap.py | 98 +++++++++---------- pipeline/tests/test_ees_ks4_summary.py | 72 ++++++++++++++ 3 files changed, 177 insertions(+), 49 deletions(-) create mode 100644 pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_summary.py create mode 100644 pipeline/tests/test_ees_ks4_summary.py diff --git a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_summary.py b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_summary.py new file mode 100644 index 0000000..3ef7eec --- /dev/null +++ b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/ks4_summary.py @@ -0,0 +1,56 @@ +"""DfE's KS4 "summary, all state-funded" data set (Key stage 4 performance). + +One CSV holds England, regional and local-authority headline rows for every +year since 2018/19. The England stream (ees_ks4_national) and the LA stream +(ees_ks4_la) both read it. Its LA rows match DfE's published performance-table +LA averages exactly (audit H2). Suppressed values ('z', 'x') become NULL in +dbt; Progress 8 is 'z' in years with no KS2 baseline (2024/25): DfE policy, +not missing data. + +Free of the Singer SDK so CI's pytest can load it. +""" +from __future__ import annotations + +import pandas as pd + +KS4_SUMMARY_CSV_URL = ( + "https://explore-education-statistics.service.gov.uk/data-catalogue/" + "data-set/1b649e16-01e8-435b-a814-56be2faf9054/csv" +) + +# CSV column → Singer field: the same 8 headline measures at every level. +KS4_HEADLINE_COL_MAP = { + "attainment8_average": "attainment_8_score", + "progress8_average": "progress_8_score", + "engmath_94_percent": "english_maths_standard_pass_pct", + "engmath_95_percent": "english_maths_strong_pass_pct", + "ebacc_entering_percent": "ebacc_entry_pct", + "ebacc_94_percent": "ebacc_standard_pass_pct", + "ebacc_95_percent": "ebacc_strong_pass_pct", + "ebacc_aps_average": "ebacc_avg_score", +} + + +def headline_rows(df: pd.DataFrame, geographic_level: str) -> pd.DataFrame: + """All-pupil rows for all state-funded schools at one geographic level + ("National" or "Local authority"). Column names are lower-cased first; + a filter column the file lacks is not applied.""" + df = df.copy() + df.columns = [c.strip().lower() for c in df.columns] + for col, want in ( + ("geographic_level", geographic_level), + ("establishment_type_group", "All state-funded"), + ("breakdown_topic", "Total"), + ("breakdown", "Total"), + ): + if col in df.columns: + df = df[df[col].str.strip().str.lower() == want.lower()] + return df + + +def headline_record(row: pd.Series, keys: tuple[str, ...]) -> dict[str, str]: + """A Singer record: the identifying columns, then the headline measures.""" + record = {key: str(row.get(key, "")).strip() for key in keys} + for csv_col, field in KS4_HEADLINE_COL_MAP.items(): + record[field] = str(row.get(csv_col, "")).strip() + return record diff --git a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py index cb7816a..d550b80 100644 --- a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py +++ b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py @@ -18,6 +18,12 @@ from singer_sdk import Stream, Tap from singer_sdk import typing as th from tap_uk_ees.ks4_info import KS4_INFO_FIELDS, KS4_INFO_RENAMES +from tap_uk_ees.ks4_summary import ( + KS4_HEADLINE_COL_MAP, + KS4_SUMMARY_CSV_URL, + headline_record, + headline_rows, +) from tap_uk_ees.release_precedence import drop_owned_periods, newest_first, periods_in CONTENT_API_BASE = ( @@ -571,37 +577,23 @@ class EESKs2NationalStream(Stream): yield record -# ── KS4 National Headlines (national level only — one row per year) ────────── -# Dataset: "National characteristics summary data" (Key stage 4 performance). -# Official England state-funded headline measures, 2018/19 → latest. -# Suppressed values ('z', 'x') → NULL downstream. Progress 8 is legitimately -# absent in years with no KS2 baseline (e.g. 2024/25) — that is DfE policy, -# not missing data. +# ── KS4 National and LA Headlines (one data set, two streams) ──────────────── +# DfE's "summary, all state-funded" data set: England, regional and LA rows, +# 2018/19 → latest. URL, measures and filters: ks4_summary.py. -_KS4_NATIONAL_CSV_URL = ( - "https://explore-education-statistics.service.gov.uk/data-catalogue/" - "data-set/1b649e16-01e8-435b-a814-56be2faf9054/csv" -) +def _read_ks4_summary(logger): + """Download DfE's KS4 summary data set.""" + import pandas as pd -_KS4_NATIONAL_COL_MAP = { - "attainment8_average": "attainment_8_score", - "progress8_average": "progress_8_score", - "engmath_94_percent": "english_maths_standard_pass_pct", - "engmath_95_percent": "english_maths_strong_pass_pct", - "ebacc_entering_percent": "ebacc_entry_pct", - "ebacc_94_percent": "ebacc_standard_pass_pct", - "ebacc_95_percent": "ebacc_strong_pass_pct", - "ebacc_aps_average": "ebacc_avg_score", -} + logger.info("Downloading KS4 summary data set: %s", KS4_SUMMARY_CSV_URL) + resp = requests.get(KS4_SUMMARY_CSV_URL, timeout=60) + resp.raise_for_status() + return pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False) class EESKs4NationalStream(Stream): - """National KS4 headline averages — one row per academic year. - - Filters to geographic_level == 'National', establishment_type_group == - 'All state-funded', breakdown_topic == 'Total', breakdown == 'Total' - so only the England-wide all-pupils row per year is emitted. - """ + """National KS4 headline averages — one row per academic year (England, + all state-funded schools, all pupils).""" name = "ees_ks4_national" primary_keys = ["time_period"] @@ -609,34 +601,41 @@ class EESKs4NationalStream(Stream): schema = th.PropertiesList( th.Property("time_period", th.StringType, required=True), - *[th.Property(out, th.StringType) for out in _KS4_NATIONAL_COL_MAP.values()], + *[th.Property(out, th.StringType) for out in KS4_HEADLINE_COL_MAP.values()], ).to_dict() def get_records(self, context): - import pandas as pd - - self.logger.info("Downloading KS4 national headlines: %s", _KS4_NATIONAL_CSV_URL) - resp = requests.get(_KS4_NATIONAL_CSV_URL, timeout=60) - resp.raise_for_status() - - df = pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False) - df.columns = [c.strip().lower() for c in df.columns] - - for col, want in [ - ("geographic_level", "national"), - ("establishment_type_group", "all state-funded"), - ("breakdown_topic", "total"), - ("breakdown", "total"), - ]: - if col in df.columns: - df = df[df[col].str.strip().str.lower() == want] - + df = headline_rows(_read_ks4_summary(self.logger), "National") self.logger.info("Emitting %d national KS4 rows", len(df)) for _, row in df.iterrows(): - record = {"time_period": row.get("time_period", "").strip()} - for csv_col, field in _KS4_NATIONAL_COL_MAP.items(): - record[field] = row.get(csv_col, "").strip() - yield record + yield headline_record(row, ("time_period",)) + + +class EESKs4LaStream(Stream): + """DfE's KS4 local-authority averages — one row per academic year and LA + (all state-funded schools, all pupils), from the same data set as + ees_ks4_national. They match DfE's published performance-table LA averages + and replace a mean the API took over every school, independent and special + included (audit H2). old_la_code is the GIAS LA code: schools join on it, + not on the name.""" + + name = "ees_ks4_la" + primary_keys = ["time_period", "old_la_code"] + replication_key = None + + schema = th.PropertiesList( + th.Property("time_period", th.StringType, required=True), + th.Property("old_la_code", th.StringType, required=True), + th.Property("new_la_code", th.StringType), + th.Property("la_name", th.StringType), + *[th.Property(out, th.StringType) for out in KS4_HEADLINE_COL_MAP.values()], + ).to_dict() + + def get_records(self, context): + df = headline_rows(_read_ks4_summary(self.logger), "Local authority") + self.logger.info("Emitting %d LA KS4 rows", len(df)) + for _, row in df.iterrows(): + yield headline_record(row, ("time_period", "old_la_code", "new_la_code", "la_name")) # ── Legacy KS2 (pre-COVID wide format from DfE performance tables) ──────────── @@ -979,6 +978,7 @@ class TapUKEES(Tap): LegacyKS4Stream(self), EESKs2NationalStream(self), EESKs4NationalStream(self), + EESKs4LaStream(self), ] diff --git a/pipeline/tests/test_ees_ks4_summary.py b/pipeline/tests/test_ees_ks4_summary.py new file mode 100644 index 0000000..da8a1a9 --- /dev/null +++ b/pipeline/tests/test_ees_ks4_summary.py @@ -0,0 +1,72 @@ +"""DfE's KS4 "summary, all state-funded" data set holds England, regional and +LA rows. The England stream kept only the England row; its LA rows match +DfE's published LA averages exactly, where the API's own mean was 7 points +low (audit H2). Values below are DfE's for Kensington and Chelsea (207) and +Wandsworth (212). +""" +import importlib.util +import io +from pathlib import Path + +import pandas as pd +import pytest + +MODULE = (Path(__file__).resolve().parents[1] / 'plugins' / 'extractors' / 'tap-uk-ees' + / 'tap_uk_ees' / 'ks4_summary.py') + +CSV = """time_period,geographic_level,old_la_code,new_la_code,la_name,establishment_type_group,breakdown_topic,breakdown,attainment8_average,progress8_average,engmath_94_percent,engmath_95_percent,ebacc_entering_percent,ebacc_94_percent,ebacc_95_percent,ebacc_aps_average +202425,National,,,,All state-funded,Total,Total,46.1,z,64.5,45.7,40.5,26.9,17.7,4.1 +202425,National,,,,All state-funded,Sex,Boys,44.1,z,62.0,43.0,38.0,24.0,16.0,3.9 +202425,Regional,,,,All state-funded,Total,Total,47.2,z,66.0,47.0,41.0,28.0,18.0,4.2 +202425,Local authority,207,E09000020,Kensington and Chelsea,All state-funded,Total,Total,54.5,z,77,61.4,45.6,32,26.6,4.89 +202425,Local authority,207,E09000020,Kensington and Chelsea,All state-funded,Sex,Girls,57.0,z,80,64.0,48.0,35,28.0,5.1 +202324,Local authority,207,E09000020,Kensington and Chelsea,All state-funded,Total,Total,54.5,0.29,76,60.0,44.0,31,25.0,4.8 +202425,Local authority,212,E09000032,Wandsworth,All state-funded,Total,Total,51.8,z,72,55.0,50.0,33,24.0,4.6 +202425,Local authority,212,E09000032,Wandsworth,Academies and free schools,Total,Total,52.0,z,73,56.0,51.0,34,25.0,4.7 +""" + +LA_KEYS = ('time_period', 'old_la_code', 'new_la_code', 'la_name') + + +def _df(): + return pd.read_csv(io.StringIO(CSV), dtype=str, keep_default_na=False) + + +@pytest.fixture +def summary(): + spec = importlib.util.spec_from_file_location('ks4_summary', MODULE) + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +def test_la_rows_are_one_per_la_and_year(summary): + rows = summary.headline_rows(_df(), 'Local authority') + assert sorted(zip(rows['time_period'], rows['old_la_code'])) == [ + ('202324', '207'), ('202425', '207'), ('202425', '212')] + + +def test_an_la_record_carries_codes_name_and_the_headline_measures(summary): + rows = summary.headline_rows(_df(), 'Local authority') + row = rows[(rows['time_period'] == '202425') & (rows['old_la_code'] == '207')].iloc[0] + assert summary.headline_record(row, LA_KEYS) == { + 'time_period': '202425', 'old_la_code': '207', 'new_la_code': 'E09000020', + 'la_name': 'Kensington and Chelsea', + 'attainment_8_score': '54.5', 'progress_8_score': 'z', + 'english_maths_standard_pass_pct': '77', 'english_maths_strong_pass_pct': '61.4', + 'ebacc_entry_pct': '45.6', 'ebacc_standard_pass_pct': '32', + 'ebacc_strong_pass_pct': '26.6', 'ebacc_avg_score': '4.89', + } + + +def test_the_england_rows_are_one_per_year(summary): + rows = summary.headline_rows(_df(), 'National') + assert list(rows['time_period']) == ['202425'] + assert summary.headline_record(rows.iloc[0], ('time_period',))['attainment_8_score'] == '46.1' + + +def test_column_names_and_labels_match_whatever_their_case(summary): + df = _df() + df.columns = [c.upper() for c in df.columns] + df['GEOGRAPHIC_LEVEL'] = df['GEOGRAPHIC_LEVEL'].str.upper() + assert len(summary.headline_rows(df, 'Local authority')) == 3 -- 2.54.0 From 59dd20ab00dadd61669f8e0bda1091903fb00324 Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 12:32:22 +0100 Subject: [PATCH 06/11] feat(dbt): fact_ks4_la_averages from DfE's LA rows (H2) Co-Authored-By: Claude Opus 5.5 --- docs/ARCHITECTURE.md | 5 +++++ pipeline/dags/school_data_pipeline.py | 2 +- .../transform/models/marts/_marts_schema.yml | 15 ++++++++++++- .../models/marts/fact_ks4_la_averages.sql | 22 +++++++++++++++++++ .../transform/models/staging/_stg_sources.yml | 3 +++ .../models/staging/stg_ees_ks4_la.sql | 21 ++++++++++++++++++ .../models/staging/stg_ees_ks4_la.yml | 15 +++++++++++++ 7 files changed, 81 insertions(+), 2 deletions(-) create mode 100644 pipeline/transform/models/marts/fact_ks4_la_averages.sql create mode 100644 pipeline/transform/models/staging/stg_ees_ks4_la.sql create mode 100644 pipeline/transform/models/staging/stg_ees_ks4_la.yml diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 6d72959..23ae57c 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -61,6 +61,11 @@ history from the full DataFrame and supplementary data from marts. Comparisons batch supplementary queries across selected URNs. Async routes still contain synchronous dependency calls; a fully asynchronous database layer is not present. +Benchmarks are DfE's own figures, read from marts: `fact_ks2_national_averages` +and `fact_ks4_national_averages` for England, and `fact_ks4_la_averages` (all +state-funded schools) for the search rows' "vs LA avg". The API never averages +school rows to make a benchmark. + ## Frontend boundaries `app/(frontend)` owns the public root layout and pages. `app/(payload)` owns the diff --git a/pipeline/dags/school_data_pipeline.py b/pipeline/dags/school_data_pipeline.py index 2b27af7..68a683e 100644 --- a/pipeline/dags/school_data_pipeline.py +++ b/pipeline/dags/school_data_pipeline.py @@ -193,7 +193,7 @@ with DAG( dbt_build_ees = BashOperator( task_id="dbt_build", - bash_command=f"cd {PIPELINE_DIR}/transform && {DBT_BIN} build --profiles-dir . --target production --select stg_ees_ks2+ stg_legacy_ks2+ stg_ees_ks4+ stg_legacy_ks4+ stg_ees_census+ stg_ees_admissions+ stg_ees_ks2_national+ stg_ees_ks4_national+ stg_ees_ks4_destinations+ stg_ees_ks5_destinations+ stg_ees_ks4_destinations_national+ stg_ees_ks5_destinations_national+", + bash_command=f"cd {PIPELINE_DIR}/transform && {DBT_BIN} build --profiles-dir . --target production --select stg_ees_ks2+ stg_legacy_ks2+ stg_ees_ks4+ stg_legacy_ks4+ stg_ees_census+ stg_ees_admissions+ stg_ees_ks2_national+ stg_ees_ks4_national+ stg_ees_ks4_la+ stg_ees_ks4_destinations+ stg_ees_ks5_destinations+ stg_ees_ks4_destinations_national+ stg_ees_ks5_destinations_national+", ) sync_typesense_ees = BashOperator( diff --git a/pipeline/transform/models/marts/_marts_schema.yml b/pipeline/transform/models/marts/_marts_schema.yml index b198b6b..4355775 100644 --- a/pipeline/transform/models/marts/_marts_schema.yml +++ b/pipeline/transform/models/marts/_marts_schema.yml @@ -279,11 +279,24 @@ models: tests: [not_null, unique] - name: fact_ks4_national_averages - description: Computed national KS4 averages (means across state schools in our dataset — not official DfE figures) — one row per academic year + description: Official DfE KS4 national headline averages (England, all state-funded schools, all pupils) — one row per academic year columns: - name: year tests: [not_null, unique] + - name: fact_ks4_la_averages + description: Official DfE KS4 local-authority averages (all state-funded schools, all pupils) — one row per academic year and LA; la_code is the GIAS LA code + columns: + - name: year + tests: [not_null] + - name: la_code + tests: [not_null] + - name: la_name + tests: [not_null] + tests: + - unique: + column_name: "year || '-' || la_code" + - name: fact_deprivation description: IDACI deprivation index — one row per URN columns: diff --git a/pipeline/transform/models/marts/fact_ks4_la_averages.sql b/pipeline/transform/models/marts/fact_ks4_la_averages.sql new file mode 100644 index 0000000..40ab45c --- /dev/null +++ b/pipeline/transform/models/marts/fact_ks4_la_averages.sql @@ -0,0 +1,22 @@ +{{ config(materialized='table') }} + +-- Mart: OFFICIAL DfE KS4 local-authority averages — one row per academic year +-- and LA (all state-funded schools, all pupils), from the same EES data set +-- as fact_ks4_national_averages. Feeds "vs LA avg" on search rows and map +-- cards, replacing a mean the API took over every school in the LA, +-- independent and special included (audit H2). la_code is the GIAS LA code. + +select + year, + la_code, + la_name, + attainment_8_score, + progress_8_score, + english_maths_standard_pass_pct, + english_maths_strong_pass_pct, + ebacc_entry_pct, + ebacc_standard_pass_pct, + ebacc_strong_pass_pct, + ebacc_avg_score +from {{ ref('stg_ees_ks4_la') }} +order by year, la_code diff --git a/pipeline/transform/models/staging/_stg_sources.yml b/pipeline/transform/models/staging/_stg_sources.yml index 2d5880a..7aa577f 100644 --- a/pipeline/transform/models/staging/_stg_sources.yml +++ b/pipeline/transform/models/staging/_stg_sources.yml @@ -79,6 +79,9 @@ sources: - name: ees_ks4_national description: Official KS4 national headline averages from DfE EES data catalogue — one row per academic year + - name: ees_ks4_la + description: Official KS4 local-authority averages (all state-funded schools) from the same DfE EES data set as ees_ks4_national — one row per academic year and LA + # Phonics: no school-level data on EES (only national/LA level) - name: fbit_finance diff --git a/pipeline/transform/models/staging/stg_ees_ks4_la.sql b/pipeline/transform/models/staging/stg_ees_ks4_la.sql new file mode 100644 index 0000000..225ce7e --- /dev/null +++ b/pipeline/transform/models/staging/stg_ees_ks4_la.sql @@ -0,0 +1,21 @@ +-- Staging model: official DfE KS4 local-authority averages — one row per +-- academic year and LA (all state-funded schools, all pupils). Same EES data +-- set as stg_ees_ks4_national. la_code is the GIAS LA code +-- (dim_location.local_authority_code). Suppressed values ('z', 'x') are +-- coerced to NULL by safe_numeric. + +select + cast(trim(time_period) as integer) as year, + cast(trim(old_la_code) as integer) as la_code, + trim(la_name) as la_name, + {{ safe_numeric('attainment_8_score') }} as attainment_8_score, + {{ safe_numeric('progress_8_score') }} as progress_8_score, + {{ safe_numeric('english_maths_standard_pass_pct') }} as english_maths_standard_pass_pct, + {{ safe_numeric('english_maths_strong_pass_pct') }} as english_maths_strong_pass_pct, + {{ safe_numeric('ebacc_entry_pct') }} as ebacc_entry_pct, + {{ safe_numeric('ebacc_standard_pass_pct') }} as ebacc_standard_pass_pct, + {{ safe_numeric('ebacc_strong_pass_pct') }} as ebacc_strong_pass_pct, + {{ safe_numeric('ebacc_avg_score') }} as ebacc_avg_score +from {{ source('raw', 'ees_ks4_la') }} +where time_period ~ '^[0-9]+$' + and old_la_code ~ '^[0-9]+$' diff --git a/pipeline/transform/models/staging/stg_ees_ks4_la.yml b/pipeline/transform/models/staging/stg_ees_ks4_la.yml new file mode 100644 index 0000000..6e5b52b --- /dev/null +++ b/pipeline/transform/models/staging/stg_ees_ks4_la.yml @@ -0,0 +1,15 @@ +version: 2 + +unit_tests: + - name: stg_ees_ks4_la_casts_codes_years_and_measures + description: DfE's LA rows, keyed by the GIAS LA code. Suppressed values ('z') become null. + model: stg_ees_ks4_la + given: + - input: source('raw', 'ees_ks4_la') + rows: + - {time_period: '202425', old_la_code: '207', new_la_code: 'E09000020', la_name: 'Kensington and Chelsea', attainment_8_score: '54.5', progress_8_score: 'z', english_maths_standard_pass_pct: '77'} + - {time_period: '202324', old_la_code: '207', new_la_code: 'E09000020', la_name: 'Kensington and Chelsea', attainment_8_score: '54.5', progress_8_score: '0.29', english_maths_standard_pass_pct: '76'} + expect: + rows: + - {year: 202425, la_code: 207, la_name: 'Kensington and Chelsea', attainment_8_score: 54.5, progress_8_score: null, english_maths_standard_pass_pct: 77} + - {year: 202324, la_code: 207, la_name: 'Kensington and Chelsea', attainment_8_score: 54.5, progress_8_score: 0.29, english_maths_standard_pass_pct: 76} -- 2.54.0 From 242c603aecce35c81825a831385e0c5a300dbcd7 Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 12:33:03 +0100 Subject: [PATCH 07/11] test(dbt): fail the build when a published year loads empty (C2, H2) Co-Authored-By: Claude Opus 5.5 --- .../tests/assert_ks2_info_percentages_loaded.sql | 16 ++++++++++++++++ .../tests/assert_ks4_la_averages_cover_las.sql | 13 +++++++++++++ .../tests/assert_ks4_years_have_results.sql | 12 ++++++++++++ 3 files changed, 41 insertions(+) create mode 100644 pipeline/transform/tests/assert_ks2_info_percentages_loaded.sql create mode 100644 pipeline/transform/tests/assert_ks4_la_averages_cover_las.sql create mode 100644 pipeline/transform/tests/assert_ks4_years_have_results.sql diff --git a/pipeline/transform/tests/assert_ks2_info_percentages_loaded.sql b/pipeline/transform/tests/assert_ks2_info_percentages_loaded.sql new file mode 100644 index 0000000..c994dea --- /dev/null +++ b/pipeline/transform/tests/assert_ks2_info_percentages_loaded.sql @@ -0,0 +1,16 @@ +-- Where a year's KS2 school information loaded (pupil counts present), its +-- percentages loaded too. DfE's files give a disadvantaged % for 96–97% of +-- those schools; 2023/24 loaded none because the file writes "34%" (audit +-- C2). A year without an information file (2022/23) has no pupil counts and +-- is skipped. + +select + year, + count(total_pupils) as with_pupils, + count(*) filter (where total_pupils is not null and disadvantaged_pct is not null) + as with_disadvantaged +from {{ ref('stg_ees_ks2') }} +group by year +having count(total_pupils) >= 1000 + and count(*) filter (where total_pupils is not null and disadvantaged_pct is not null) + < 0.9 * count(total_pupils) diff --git a/pipeline/transform/tests/assert_ks4_la_averages_cover_las.sql b/pipeline/transform/tests/assert_ks4_la_averages_cover_las.sql new file mode 100644 index 0000000..bad12df --- /dev/null +++ b/pipeline/transform/tests/assert_ks4_la_averages_cover_las.sql @@ -0,0 +1,13 @@ +-- DfE publishes an average for about 152 LAs a year. Fewer than 145 in the +-- latest year (or none at all) means the LA filter in tap_uk_ees +-- (ks4_summary.headline_rows) stopped matching, e.g. after DfE renamed a label. + +with latest as ( + select max(year) as year from {{ ref('fact_ks4_la_averages') }} +) + +select l.year, count(f.la_code) as las +from latest l +left join {{ ref('fact_ks4_la_averages') }} f on f.year = l.year +group by l.year +having count(f.la_code) < 145 diff --git a/pipeline/transform/tests/assert_ks4_years_have_results.sql b/pipeline/transform/tests/assert_ks4_years_have_results.sql new file mode 100644 index 0000000..04b9df2 --- /dev/null +++ b/pipeline/transform/tests/assert_ks4_years_have_results.sql @@ -0,0 +1,12 @@ +-- Every year loaded from EES has an Attainment 8 for at least half its +-- schools. DfE's files reach 82% each year; 2023/24 loaded at 0% after DfE +-- re-issued its file under older column names (audit C2). Pre-2019 years come +-- from stg_legacy_ks4 and are not checked here. + +select + year, + count(*) as schools, + count(attainment_8_score) as with_attainment_8 +from {{ ref('stg_ees_ks4') }} +group by year +having count(attainment_8_score) < 0.5 * count(*) -- 2.54.0 From 4ae3853277611f46846ba91fe2ae054a50e4c909 Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 12:33:36 +0100 Subject: [PATCH 08/11] fix(dbt): read percentages DfE writes as "34%" (C2) Co-Authored-By: Claude Opus 5.5 --- pipeline/transform/macros/safe_numeric.sql | 4 +++- .../transform/models/staging/stg_ees_ks2.yml | 22 +++++++++++++++++++ 2 files changed, 25 insertions(+), 1 deletion(-) create mode 100644 pipeline/transform/models/staging/stg_ees_ks2.yml diff --git a/pipeline/transform/macros/safe_numeric.sql b/pipeline/transform/macros/safe_numeric.sql index 5f14cf1..a1d8d29 100644 --- a/pipeline/transform/macros/safe_numeric.sql +++ b/pipeline/transform/macros/safe_numeric.sql @@ -3,7 +3,9 @@ Casts a string column to numeric, treating any non-numeric value as NULL. Handles all EES suppression codes (z, c, x, q, u, etc.) without needing an explicit list — any string that doesn't look like a number becomes NULL. + A trailing percent sign is accepted: DfE's 2023/24 KS2 information file + writes percentages as "34%" (audit C2). #} {% macro safe_numeric(col) -%} - CASE WHEN {{ col }} ~ '^-?[0-9]+(\.[0-9]+)?$' THEN {{ col }}::numeric ELSE NULL END + CASE WHEN {{ col }} ~ '^-?[0-9]+(\.[0-9]+)?%?$' THEN rtrim({{ col }}, '%')::numeric ELSE NULL END {%- endmacro %} diff --git a/pipeline/transform/models/staging/stg_ees_ks2.yml b/pipeline/transform/models/staging/stg_ees_ks2.yml new file mode 100644 index 0000000..59c9ecd --- /dev/null +++ b/pipeline/transform/models/staging/stg_ees_ks2.yml @@ -0,0 +1,22 @@ +version: 2 + +unit_tests: + - name: stg_ees_ks2_reads_percentages_written_with_a_sign + description: > + DfE's 2023/24 KS2 information file writes percentages as "34%", and every + 2023/24 percentage loaded as null (audit C2). Plain numbers still load; + suppression codes stay null. School 147411's 2023/24 figures are DfE's. + model: stg_ees_ks2 + given: + - input: source('raw', 'ees_ks2_attainment') + rows: + - {school_urn: '147411', time_period: '202324', subject: 'Reading, writing and maths', breakdown_topic: 'All pupils', breakdown: 'Total', expected_standard_pupil_percent: '70'} + - {school_urn: '100000', time_period: '202425', subject: 'Reading, writing and maths', breakdown_topic: 'All pupils', breakdown: 'Total', expected_standard_pupil_percent: '79'} + - input: source('raw', 'ees_ks2_info') + rows: + - {school_urn: '147411', time_period: '202324', totpups: '818', telig: '112', ptfsm6cla1a: '34%', ptealgrp2: '56%', psenelk: '20%', psenele: 'c', ptmobn: '87.5%'} + - {school_urn: '100000', time_period: '202425', totpups: '792', telig: '105', ptfsm6cla1a: '32', ptealgrp2: 'x', psenelk: '14', psenele: '2', ptmobn: '87'} + expect: + rows: + - {urn: 147411, year: 202324, total_pupils: 818, rwm_expected_pct: 70, disadvantaged_pct: 34, eal_pct: 56, sen_support_pct: 20, sen_ehcp_pct: null, stability_pct: 87.5} + - {urn: 100000, year: 202425, total_pupils: 792, rwm_expected_pct: 79, disadvantaged_pct: 32, eal_pct: null, sen_support_pct: 14, sen_ehcp_pct: 2, stability_pct: 87} -- 2.54.0 From e09d7a202b26101bbf77c2cc8fd3d699c6d0e8cc Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 12:50:10 +0100 Subject: [PATCH 09/11] fix(ees): read the year from suffixed release slugs (review) A '2025-26-revised' release read as an unknown year went last, behind the '2025-26' release, which then owned the year and dropped every revised row. Co-Authored-By: Claude Opus 5.5 --- .../tap_uk_ees/release_precedence.py | 15 +++++++++++++ .../extractors/tap-uk-ees/tap_uk_ees/tap.py | 17 ++++++-------- pipeline/tests/test_ees_release_precedence.py | 22 +++++++++++++++++++ 3 files changed, 44 insertions(+), 10 deletions(-) diff --git a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py index 52080c6..1a63779 100644 --- a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py +++ b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/release_precedence.py @@ -11,8 +11,23 @@ requirements, can load it. """ from __future__ import annotations +import re + import pandas as pd +_SLUG_YEAR = re.compile(r"^(\d{4})-(\d{2})(?:-|$)") + + +def slug_to_time_period(slug: str) -> str | None: + """A release slug's academic year as a time_period: '2022-23' → '202223'. + + Suffixed slugs ('2024-25-revised', '2025-26-provisional') give the same + year, so newest_first places them by year; read as unknown, a revised + release went last and lost its year to the first release. + """ + match = _SLUG_YEAR.match(slug or "") + return match.group(1) + match.group(2) if match else None + def newest_first(releases: list[dict]) -> list[dict]: """Releases by time_period, newest first. A release whose time_period is diff --git a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py index d550b80..541ef8b 100644 --- a/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py +++ b/pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py @@ -24,7 +24,12 @@ from tap_uk_ees.ks4_summary import ( headline_record, headline_rows, ) -from tap_uk_ees.release_precedence import drop_owned_periods, newest_first, periods_in +from tap_uk_ees.release_precedence import ( + drop_owned_periods, + newest_first, + periods_in, + slug_to_time_period, +) CONTENT_API_BASE = ( "https://content.explore-education-statistics.service.gov.uk/api" @@ -40,14 +45,6 @@ def get_content_release_id(publication_slug: str) -> str: return resp.json()["id"] -def _slug_to_time_period(slug: str) -> str | None: - """Convert a release slug like '2022-23' to a time_period like '202223'.""" - parts = slug.split("-") - if len(parts) == 2 and len(parts[0]) == 4 and len(parts[1]) == 2: - return parts[0] + parts[1] - return None - - def get_all_releases(publication_slug: str) -> list[dict]: """Return all releases for a publication as dicts with 'id' and 'time_period'. @@ -72,7 +69,7 @@ def get_all_releases(publication_slug: str) -> list[dict]: total_pages = paging.get("totalPages", 1) for r in releases: - time_period = _slug_to_time_period(r.get("slug", "")) + time_period = slug_to_time_period(r.get("slug", "")) result.append({"id": r["id"], "time_period": time_period}) if page >= total_pages: diff --git a/pipeline/tests/test_ees_release_precedence.py b/pipeline/tests/test_ees_release_precedence.py index 09ce3ae..104fcae 100644 --- a/pipeline/tests/test_ees_release_precedence.py +++ b/pipeline/tests/test_ees_release_precedence.py @@ -76,3 +76,25 @@ def test_only_the_ks4_results_stream_opts_in(): opted_in = [chunk.split('(')[0] for chunk in source.split('\nclass ')[1:] if re.search(r'_newest_release_owns_period\s*=\s*True', chunk)] assert opted_in == ['EESKS4PerformanceStream'] + + +@pytest.mark.parametrize('slug, period', [ + ('2024-25', '202425'), + ('2024-25-revised', '202425'), + ('2025-26-provisional', '202526'), + ('latest', None), + ('', None), +]) +def test_a_release_slug_gives_its_year_whatever_its_suffix(precedence, slug, period): + # KS2 already publishes "2024-25-revised" and "2025-26-provisional". An + # unread suffix sent the release last, behind the provisional one for the + # same year, which then owned the year and dropped every revised row. + assert precedence.slug_to_time_period(slug) == period + + +def test_within_a_year_the_api_order_is_kept(precedence): + # The API lists a year's revised release before its first release. + releases = [{'id': 'revised', 'time_period': '202526'}, + {'id': 'first', 'time_period': '202526'}, + {'id': 'older', 'time_period': '202425'}] + assert [r['id'] for r in precedence.newest_first(releases)] == ['revised', 'first', 'older'] -- 2.54.0 From 951c666f8ac2f8e7655654546c2444761c685d65 Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 12:50:27 +0100 Subject: [PATCH 10/11] test(dbt): warn when DfE's LA averages fall behind the school results (review) Co-Authored-By: Claude Opus 5.5 --- .../tests/assert_ks4_la_averages_current.sql | 20 +++++++++++++++++++ 1 file changed, 20 insertions(+) create mode 100644 pipeline/transform/tests/assert_ks4_la_averages_current.sql diff --git a/pipeline/transform/tests/assert_ks4_la_averages_current.sql b/pipeline/transform/tests/assert_ks4_la_averages_current.sql new file mode 100644 index 0000000..ddd914c --- /dev/null +++ b/pipeline/transform/tests/assert_ks4_la_averages_current.sql @@ -0,0 +1,20 @@ +{{ config(severity='warn') }} + +-- DfE's LA averages cover the latest year with school results. They come from +-- a fixed EES data set (ks4_summary.KS4_SUMMARY_CSV_URL); when DfE publishes a +-- new year under a new data set, the school results move on and this mart does +-- not, so /api/la-averages serves an empty map and every "vs LA avg" goes. +-- A warning, not an error: it must not hold back a new year's school results. + +with schools as ( + select max(year) as year from {{ ref('stg_ees_ks4') }} where attainment_8_score is not null +), + +las as ( + select max(year) as year from {{ ref('fact_ks4_la_averages') }} +) + +select s.year as latest_school_year, l.year as latest_la_year +from schools s +cross join las l +where l.year is null or l.year < s.year -- 2.54.0 From 5f93a7ecd292587c4c3ad635e548902eb0604e95 Mon Sep 17 00:00:00 2001 From: Tudor Date: Tue, 6 Oct 2026 12:50:40 +0100 Subject: [PATCH 11/11] docs: name the benchmarks computed from our dataset (review) Co-Authored-By: Claude Opus 5.5 --- docs/ARCHITECTURE.md | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 23ae57c..3219b69 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -61,10 +61,11 @@ history from the full DataFrame and supplementary data from marts. Comparisons batch supplementary queries across selected URNs. Async routes still contain synchronous dependency calls; a fully asynchronous database layer is not present. -Benchmarks are DfE's own figures, read from marts: `fact_ks2_national_averages` -and `fact_ks4_national_averages` for England, and `fact_ks4_la_averages` (all -state-funded schools) for the search rows' "vs LA avg". The API never averages -school rows to make a benchmark. +DfE's official benchmarks live in marts: `fact_ks2_national_averages` and +`fact_ks4_national_averages` for England, and `fact_ks4_la_averages` (all +state-funded schools, per LA) for the search rows' "vs LA avg". +`data_loader.compute_benchmarks` computes further state-school benchmarks from +our own dataset; they are not DfE figures and are labelled as computed. ## Frontend boundaries -- 2.54.0