docs: design for 2023/24 KS4/KS2 data and DfE LA averages (C2, H2)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
TudorandClaude Opus 5.5 committed 2026-10-06 10:08:25 +01:00
1 parent c26b65246f
commit ac5b7ccd7f
1 file changed
+300
@@ -0,0 +1,300 @@
# 2023/24 Results and DfE LA Averages — Design
**Date:** 2026-10-06
**Status:** approved design, not yet implemented
**Scope:** `tap_uk_ees` (KS4 results, KS4 information, new LA stream),
`safe_numeric`, new `fact_ks4_la_averages` mart, annual EES DAG selector,
`/api/la-averages`, secondary search rows and map cards
**Fixes:** audit findings C2 and H2, from the 3 Oct 2026 accuracy audit
## Goal
Show every 2023/24 result and school-information figure DfE published, and
compare each state-funded secondary school with DfE's own local-authority
average.
## The problem
### C2: 2023/24 is empty
Bishop Stopford School (137086) shows Attainment 8 60.7 for 2022/23, nothing
for 2023/24 and 58.7 for 2024/25. DfE's 2023/24 figures are Attainment 8 64.1,
Progress 8 +1.02 and English and maths grade 4+ 91.7%. 2023/24 is the last year
DfE published Progress 8 (2024/25 has no KS2 baseline), so the site shows no
recent Progress 8 for any school. Site-wide, 4,170 listed schools have a DfE
2023/24 Attainment 8 and 3,384 a Progress 8.
There are three separate causes.
1. **KS4 results.** The 2024/25 release's
`202425_performance_tables_schools_final.csv` is a time series: it holds
2022/23, 2023/24 and 2024/25 under the current column names, with the right
values (Bishop Stopford 2023/24: 64.1, 1.02, 91.7). The 2023/24 release's own
`202324_performance_tables_schools_final.csv`, re-issued on 10 March 2026,
uses the older names (`t_pupils`, `avg_att8`, `avg_p8score`,
`pt_l2basics_94` …). `EESDatasetStream` reads releases in the API's order,
newest first, so the old file is read last. Its rows carry none of the
declared fields, and target-postgres upserts on the stream's primary key
(`append_only = not key_properties` in meltanolabs-target-postgres 0.8.0),
so they overwrite the good 2023/24 rows with nulls.
The two files share the keys of every "Total" row, but 34,254 of the 57,090
2023/24 sub-group rows use different labels ("Low prior" against "Low prior
attainment"). Those old-label rows sit in `raw.ees_ks4_performance` with
null measures. Nothing reads them.
2. **KS4 school information.** 2023/24 information exists only in the 2023/24
release, in `202324_information_about_schools_final.csv`, with the older
names. `EESKS4InfoStream` declares the newer ones, so prior attainment,
SEN percentages, disadvantage gaps and Progress 8 banding are null for
2023/24.
3. **KS2 school information.** The 2023/24 file
`ks2_school_information_data.csv` uses the declared names, and pupil counts
load (school 147411: 818 pupils, 112 eligible). Its percentages are written
with a sign (`ptfsm6cla1a = "34%"`). `safe_numeric` accepts only
`^-?[0-9]+(\.[0-9]+)?$`, so disadvantaged, EAL, SEN and mobility percentages
are null for every school in 2023/24.
DfE published no school-level KS2 information file for 2022/23 (the 2023/24
release's attainment file carries 2022/23 attainment rows only). Those nulls are
a gap in the source, not a defect.
### H2: "vs LA avg" uses the wrong average
`/api/la-averages` (`backend/app.py`) takes an unweighted mean of every school
in the LA with an Attainment 8 score, independent and special schools included.
Independent schools score low because DfE measures exclude IGCSEs, and special
schools score low for other reasons, so the average is too low almost
everywhere. Kensington and Chelsea's "LA avg" is 35.2: the mean of 6 state
schools (54.9) and 8 independent schools (20.4). DfE's figure is 54.5. Of the
151 LAs the audit compared, ours was lower in 147, by 7.1 points on average and
by up to 19.3, so most secondary schools look better than their area.
DfE's LA averages are already in the file the pipeline downloads for the
England averages: the "summary, all state-funded" data set
(`data-catalogue/data-set/1b649e16-01e8-435b-a814-56be2faf9054/csv`). Its
`Local authority` / `All state-funded` / Total rows match DfE's published
performance-table LA averages (RECTYPE 4) for all 152 LAs in 2024/25, with no
difference. `EESKs4NationalStream` keeps the England row and discards them.
"All state-funded" is the same population as the England benchmark the site
already shows.
Computing the average ourselves from state-funded schools, weighted by pupils,
was tested and rejected: it differs from DfE's figure by 1.1 points on average
and is never exact.
## Non-goals
- 2022/23 KS2 school information (DfE published none).
- An "excludes IGCSEs" note wherever an independent school's Attainment 8
appears. This change only stops comparing independent schools with the LA.
- Showing 2023/24 Progress 8 in the GCSE section's headline. The history section
shows it once the data loads; the 2024/25 banner stays true.
- Updating the fixed EES data-set id when DfE publishes 2025/26. It already
feeds the England averages; the LA stream shares it.
- LA comparisons for other measures or on other pages.
- Deleting the leftover old-label raw rows automatically.
## The rules
### The newest release owns every year it contains (KS4 results only)
`EESDatasetStream` gets an opt-in class attribute,
`_newest_release_owns_period: bool = False`. `EESKS4PerformanceStream` sets it
to `True`. When it is on:
- releases are processed newest first by `time_period` (from the release slug),
not in the API's order;
- the stream records each `time_period` it has emitted;
- in each older release, rows whose `time_period` a newer release already
emitted are dropped, and the stream logs how many it skipped and for which
years;
- years only an older release contains are emitted as before.
The filter is a pure function, testable without a download. Other streams keep
the current behaviour. A general rule would be wrong: the 2024/25 KS2 file holds
98,448 of the 955,956 rows the 2023/24 release has for 2023/24.
### KS4 information: old names
`EESKS4InfoStream._column_renames` maps the 2023/24 names onto the declared
fields:
| 2023/24 column | Declared field |
|---|---|
| `t_allks_pupils` | `allks_pupil_count` |
| `t_allks_boys` | `allks_boys_count` |
| `t_allks_girls` | `allks_girls_count` |
| `t_pupils` | `endks4_pupil_count` |
| `avg_ks2_scaledscore` | `ks2_scaledscore_average` |
| `pt_sen_with_ehcp` | `sen_with_ehcp_pupil_percent` |
| `pt_sen` | `sen_pupil_percent` |
| `pt_sen_no_ehcp` | `sen_no_ehcp_pupil_percent` |
| `diffn_att8` | `attainment8_diffn` |
| `diffn_p8mea` | `progress8_diffn` |
| `p8_banding` | `progress8_banding` |
Newer files contain none of the old names, so they are unaffected.
### `safe_numeric` accepts a trailing `%`
The pattern becomes `^-?[0-9]+(\.[0-9]+)?%?$` and the cast reads
`rtrim(col, '%')`. A value such as `34%` can only mean 34. Suppression codes
(`c`, `z`, `x` …) still become null.
### LA averages
A new stream, `ees_ks4_la`, reads the same CSV as `ees_ks4_national` and keeps
rows where `geographic_level = 'Local authority'`,
`establishment_type_group = 'All state-funded'`, `breakdown_topic = 'Total'` and
`breakdown = 'Total'` (case-insensitive, as the national stream compares). It
emits `time_period`, `old_la_code`, `new_la_code`, `la_name` and the 8 headline
measures the national stream emits (`_KS4_NATIONAL_COL_MAP`). Primary key:
(`time_period`, `old_la_code`). The two streams share one download-and-filter
helper.
`old_la_code` is the GIAS LA code (`local_authority_code`), so schools join on
the code, not the name. Names match today for every LA DfE publishes; DfE
publishes no figure for City of London.
### Which schools get a gap
The search row and the map card show "vs LA avg" only when the school has an
Attainment 8 score, is neither special (`isSpecialSchool`) nor independent
(`isIndependentSchool`: "independent" in the GIAS type, which covers "Other
independent school" and "Other independent special school"), and its LA has a
DfE figure for the year the endpoint serves.
## Delivery
Two PRs, as for C1.
### PR 1: pipeline
- `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py`: release
precedence, KS4 information renames, `ees_ks4_la` stream registered in
`discover_streams`, shared national/LA helper.
- `pipeline/transform/macros/safe_numeric.sql`: trailing `%`.
- `pipeline/transform/models/staging/`: source `raw.ees_ks4_la` and
`stg_ees_ks4_la` (view): `cast(old_la_code as integer) as la_code`,
`cast(time_period as integer) as year`, `la_name`, measures via
`safe_numeric`, named as in `stg_ees_ks4_national`.
- `pipeline/transform/models/marts/fact_ks4_la_averages.sql` (table): one row
per (`year`, `la_code`) with `la_name` and the columns of
`fact_ks4_national_averages`. Schema tests: unique (`year`, `la_code`);
`year`, `la_code` and `la_name` not null.
- Data tests in `pipeline/transform/tests/`:
- `assert_ks4_years_have_results`: every year in `stg_ees_ks4` has a non-null
Attainment 8 for at least 50% of its rows. DfE's files reach 82% each year;
2023/24 loads at 0% today. Pre-2019 years come from the legacy model and are
not tested.
- `assert_ks2_info_percentages_loaded`: for every year in `stg_ees_ks2` where
at least 1,000 rows have `total_pupils`, at least 90% of those rows have
`disadvantaged_pct`. DfE's files reach 96–97%. 2022/23 has no pupil counts
and is skipped.
- `assert_ks4_la_averages_cover_las`: the latest year in
`fact_ks4_la_averages` has at least 145 LAs (DfE: 152).
- `pipeline/dags/school_data_pipeline.py`: the annual EES build selects
`stg_ees_ks4_la+`.
- `docs/ARCHITECTURE.md`: LA averages come from DfE's data set.
### PR 2: backend and UI (after the EES DAG has run on PR 1)
- `backend/models.py`: `Ks4LaAverage` for `marts.fact_ks4_la_averages`.
- `backend/app.py` `/api/la-averages`: the year is the latest with any school
Attainment 8 (as now). It reads that year's rows from the mart and keys each
`attainment_8_score` by our LA name, through the `local_authority_code` →
`local_authority` pairs in the school data. The response shape is unchanged.
No rows for that year, a missing table or a query error give an empty map,
logged, never another year's figures and never a computed mean.
- `nextjs-app/lib/utils.ts`: `isIndependentSchool(school)`.
- `nextjs-app/components/SecondarySchoolRow.tsx` and
`nextjs-app/components/LeafletMapInner.tsx`: the rule in "Which schools get a
gap".
- `e2e/tests/journeys.spec.ts`: the two journeys under Testing.
## Testing
**Extractor (pytest, `pipeline/tests/`, new):**
- precedence: with the flag on, a year in a newer release is emitted once, from
the newer release; a year only an older release has is kept; with the flag
off, every row passes;
- releases arriving oldest first are still processed newest first;
- KS4 information: an old-format row yields `endks4_pupil_count`,
`ks2_scaledscore_average`, `progress8_banding` and the rest of the table;
- LA filter: a small CSV with national, regional, LA and sub-group rows yields
one row per LA and year, with the declared fields.
**dbt (local `pgserver`, as for C1):**
- unit test on `stg_ees_ks2`: `34%` → 34, `34` → 34, `c` → null;
- unit test on `stg_ees_ks4_la` or the mart: codes and years cast, measures
carried;
- the three data tests and the schema tests above;
- `pipeline/tests/test_dag_selectors.py` passes with the new selector.
**Backend (pytest):**
- the response gives the mart's figure where the fixture's plain mean differs;
- matching works by code when the mart's `la_name` differs from ours;
- a mart without the served year, and a missing table, give an empty map.
**Front end (Jest):**
- `isIndependentSchool` for "Other independent school", "Other independent
special school", "Academy converter" and null;
- `SecondarySchoolRow` and the map card show no gap for an independent school
and show one for a state school with an LA figure.
**E2E (`journeys.spec.ts`, PR 2):**
- C2: Bishop Stopford (137086) history shows 2023/24 Attainment 8 64.1 and
Progress 8 +1.02. These are final figures, so the journey stays stable.
- H2: a search returning an independent and a state secondary: the independent
row has no "vs LA avg", the state row has one. The exact gap is not asserted:
DfE's 2025/26 provisional KS4 data is due and would change it.
## Rollout and verification
1. PR 1 merges to staging. Tudor runs `school_data_annual_ees` on staging.
2. Through the staging API: Bishop Stopford 2023/24 Attainment 8 64.1 and
Progress 8 1.02; school 147411 2023/24 disadvantaged 34; the LA mart has 152
LAs for 2024/25. If a data test fails, the build stops before search sync;
investigate before going further.
3. PR 2 merges, so the post-merge E2E gate runs against loaded data.
4. Production, Tudor's decision, in order: promote PR 1, run the EES DAG on
production, promote PR 2. If PR 2 arrives first, the endpoint returns an
empty map and rows show no gap, never a wrong one.
5. Optional cleanup, in the PR 1 description for Tudor:
`delete from raw.ees_ks4_performance where time_period = '202324' and pupil_count is null`.
New-format rows always have `pupil_count` (suppressed values are `c` or `z`,
not null), so this removes exactly the old-label leftovers.
6. Validation on production against DfE's files: every listed school's 2023/24
Attainment 8, Progress 8 and English and maths figures match; the 2023/24 KS4
and KS2 information fields match; `/api/la-averages` equals DfE for all 152
LAs; independent rows show no gap. Then mark C2 and H2 resolved in the audit
report.
## Expected visible change
- Secondary history charts and tables run unbroken from 2022/23 to 2024/25, and
2023/24 Progress 8 appears for about 3,400 schools.
- 2023/24 school information (KS2 and KS4) fills in.
- "vs LA avg" drops for most state secondaries. In Kensington and Chelsea, a
school with Attainment 8 60.0 moves from +24.8 to +5.5.
- Independent schools and City of London schools show no LA gap.
## Risks
- **Data-test thresholds stop a good build.** They were set from DfE's own
files (82% and 96–97% against 50% and 90%). A failure means the data changed
shape, which is what they are for.
- **`safe_numeric` is shared by 14 models.** Values written as `n%` were null
and become numbers. No DfE column uses `%` for anything but a percentage.
The staging EES run is the check.
- **Precedence drops data an older release holds more completely.** It is
opt-in for KS4 results, where both files hold the same 57,090 2023/24 keys.
- **The fixed EES data-set id goes stale** when 2025/26 is published. The
endpoint's year check then gives no gap rather than a mismatched one.