Compare commits

...
Author SHA1 Message Date
TudorandClaude Fable 5 d5cd0abfee ci: re-run PR checks (AI review job errored without posting findings)
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m38s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 46s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 5m50s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-13 13:15:52 +01:00
TudorandClaude Fable 5 436ec6151b fix(pipeline): thread KS2 progress CI columns through the legacy union and lineage model
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m39s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 46s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 28s
The AI review gate caught that stg_ees_ks2's 7 new columns broke the
positional UNION ALL with stg_legacy_ks2 in int_ks2_with_lineage, and
that the lineage CTEs never emitted them (same class of bug fixed for
KS4 in 34a5de2). Legacy gets typed null placeholders at matching
positions; both lineage CTEs pass the columns through. 45/45 columns
verified name-identical in order across both union branches.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-13 08:42:40 +01:00
TudorandClaude Fable 5 6f925abf6b fix(pipeline): harden banding against EES sentinels; diagnostic cleanups
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m40s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 12s
PR Checks / Build Frontend (no push) (pull_request) Successful in 41s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 46s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m17s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-13 08:22:34 +01:00
TudorandClaude Fable 5 03518520f8 fix(pipeline): unknown safeguarding values parse to NULL, not "not met"
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 22:17:11 +01:00
TudorandClaude Fable 5 c2ed002118 docs(pipeline): record Task 7 report-card grade value sample + collision check
Evidence trail for the rc_* mapping in the prior commit: real value_counts()
over the 7 MI report-card columns, confirming the 5-value grade vocabulary
and that 'Achievement'/'Safeguarding standards' match by exact string only
(no legacy OEIF column accidentally consumed).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 22:13:03 +01:00
TudorandClaude Fable 5 02084e427c feat(pipeline): extract Ofsted report-card judgements (rc_* columns)
Wires the tap TODO in stg_ofsted_inspections.sql: maps the 7 confirmed
report-card MI columns (Safeguarding standards, Inclusion, Curriculum
and teaching, Achievement, Attendance and behaviour, Personal
development and wellbeing, Leadership and governance) into rc_*
fields, parsed via the new parse_report_card_grade macro against
real sampled grade values (Exceptional/Strong standard/Expected
standard/Needs attention/Urgent improvement). rc_safeguarding_met
becomes boolean from Met/Not met. rc_early_years/rc_sixth_form have
no MI column yet and are intentionally omitted from COLUMN_PRIORITY,
staying NULL.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 22:12:33 +01:00
TudorandClaude Fable 5 bee63a7836 docs(spec): 2021/22 school-level KS2 is a permanent DfE source gap, not a pipeline task
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 22:10:05 +01:00
TudorandClaude Fable 5 fc21783298 chore(pipeline): verify 2021/22 legacy KS2 archive compatibility
DfE never published school-level KS2 2021/22 data publicly (confirmed via
EES release notes and by walking the Compare School Performance download
wizard, which has no ks2 checkbox for 2021-2022, same as the COVID-cancelled
2020-2021 year). No archive exists to verify column headers against or
upload to the filebrowser; Task 6 is blocked at the source-data level.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 22:08:36 +01:00
TudorandClaude Fable 5 ccd8e73fe8 fix(pipeline): include 2015/16 national averages
Widen the year filter in stg_ees_ks2_national.sql from >= 201617 to
>= 201516 so the England national-averages line no longer starts a
year late; the catalogue CSV has a real, comparable 201516 row (2015/16
was the first year of the current expected-standard tests, so it's the
correct floor).

GPS/science/scaled-score national columns confirmed present at source
with correct mapping; prod NULLs are stale raw data, backfilled by the
next extract run. No _KS2_NATIONAL_COL_MAP change accompanies this fix.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 22:02:28 +01:00
TudorandClaude Fable 5 5f1b6adb44 fix(pipeline): align SEN column order across KS4 union branches
int_ks4_with_lineage.sql unions stg_ees_ks4 and stg_legacy_ks4 via
`select *`, which PostgreSQL aligns positionally. stg_legacy_ks4 listed
sen_support_pct before sen_ehcp_pct while stg_ees_ks4 lists sen_ehcp_pct
before sen_support_pct, swapping the two values for legacy-sourced rows
in marts.fact_ks4_performance. Reordered stg_legacy_ks4's final select
to match.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 22:00:15 +01:00
TudorandClaude Fable 5 34a5de2687 feat(pipeline): Progress 8 banding and KS4 disadvantage gaps in marts
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 21:55:34 +01:00
Tudor af43b291e7 feat(pipeline): KS2 progress confidence intervals and writing working-towards 2026-07-12 21:34:59 +01:00
Tudor 5a94f470e1 feat(pipeline): admissions preference breakdown and cross-LA demand in marts 2026-07-12 21:30:50 +01:00
Tudor 0fe1ea0d6a chore(pipeline): diagnostic for compare-screen data gaps 2026-07-12 21:25:00 +01:00
TudorandClaude Fable 5 297bdbd12e docs: compare-screen redesign spec, expert review, and data-foundation plan
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
2026-07-12 21:19:53 +01:00
tudor 58e90fef61 Merge pull request 'fix(api): blank-name GIAS sentinel codes map to empty string, not Unknown(n)' (#27) from fix/gias-blank-name-codes into main
Deploy (staging -> E2E gate -> production) / Build Backend (FastAPI) (push) Successful in 20s
Deploy (staging -> E2E gate -> production) / Build Frontend (Next.js) (push) Successful in 54s
Deploy (staging -> E2E gate -> production) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m5s
Deploy (staging -> E2E gate -> production) / Deploy to Staging (push) Successful in 1s
Deploy (staging -> E2E gate -> production) / E2E Journeys against Staging (push) Successful in 41s
Deploy (staging -> E2E gate -> production) / Promote to Production (push) Successful in 9s
Reviewed-on: #27
2026-07-09 21:21:08 +00:00
TudorandClaude Fable 5 3710529e49 fix(api): map blank-name GIAS sentinel codes to empty string, not Unknown
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m40s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 20s
PR Checks / Build Frontend (no push) (pull_request) Successful in 54s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 37s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m35s
ReligiousCharacter 99 (~4k schools) and AdmissionsPolicy 9 (~5.6k) carry a
code with a blank name in the GIAS CSV; the generator skipped them so they
hit the Unknown(<code>) path — wrongly triggering the Faith-priority tag
and polluting filters. Blank-only codes now map to "" (byte-identical to
the old name pipeline); accepted_values lists extended to match the seed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 22:04:53 +01:00
tudor 159207c6f5 Merge pull request 'fix(pipeline): add the missing backend cache-invalidation step to the data DAGs' (#26) from fix/daily-dag-cache-invalidation into main
Deploy (staging -> E2E gate -> production) / Build Backend (FastAPI) (push) Successful in 13s
Deploy (staging -> E2E gate -> production) / Build Frontend (Next.js) (push) Successful in 51s
Deploy (staging -> E2E gate -> production) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m3s
Deploy (staging -> E2E gate -> production) / Deploy to Staging (push) Successful in 1s
Deploy (staging -> E2E gate -> production) / E2E Journeys against Staging (push) Successful in 43s
Deploy (staging -> E2E gate -> production) / Promote to Production (push) Successful in 9s
Reviewed-on: #26
2026-07-09 20:41:42 +00:00
TudorandClaude Fable 5 d677b54533 fix(pipeline): cache invalidation for IDACI DAG too; curl timeouts
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m37s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 48s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 35s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m40s
Addresses AI-review findings: the annual IDACI DAG also rebuilds a mart
(fact_deprivation) and needs the reload; curl gets connect/max timeouts
so an unreachable backend fails fast instead of hanging the task.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 21:26:45 +01:00
TudorandClaude Fable 5 c353e36072 fix(pipeline): actually invalidate the backend cache after data rebuilds
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m40s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 51s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 1m20s
The daily/monthly/annual DAG docstring promised an Invalidate Cache step
that never existed — after a marts rebuild the backend kept serving its
startup-cached (possibly empty) DataFrame until a container restart.
Add a POST /api/admin/reload task at the end of each pipeline DAG,
mirroring the sitemap DAG's admin-call pattern.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 21:12:01 +01:00
tudor d9223a6d6e Merge pull request 'fix(api): legacy name-column fallback when marts predate the GIAS code migration' (#25) from fix/gias-legacy-fallback into main
Deploy (staging -> E2E gate -> production) / Build Backend (FastAPI) (push) Successful in 19s
Deploy (staging -> E2E gate -> production) / Build Frontend (Next.js) (push) Successful in 54s
Deploy (staging -> E2E gate -> production) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Deploy (staging -> E2E gate -> production) / Deploy to Staging (push) Successful in 1s
Deploy (staging -> E2E gate -> production) / E2E Journeys against Staging (push) Failing after 4m45s
Deploy (staging -> E2E gate -> production) / Promote to Production (push) Has been skipped
Reviewed-on: #25
2026-07-09 19:13:55 +00:00
TudorandClaude Fable 5 74ca76d150 fix(api): match missing-column fallbacks on the DBAPI error, not the statement
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m41s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 20s
PR Checks / Build Frontend (no push) (pull_request) Successful in 53s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m41s
str(ProgrammingError) embeds the full SQL, which contains every column
name — the substring check matched any error and could take the wrong
retry branch. Parse the missing column from exc.orig instead.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 19:29:58 +01:00
TudorandClaude Fable 5 4b75152ee0 fix(api): fall back to legacy name-column query when marts predate code migration
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m41s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m8s
Closes the deploy window flagged by CI review — the backend now works
against both the old (name) and new (code) mart schemas.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 14:48:06 +01:00
tudor 84dfc6c1bb Merge pull request 'feat: GIAS classification fields stored as codes, translated in code' (#24) from feat/gias-code-dictionaries into main
Deploy (staging -> E2E gate -> production) / Build Backend (FastAPI) (push) Successful in 19s
Deploy (staging -> E2E gate -> production) / Build Frontend (Next.js) (push) Successful in 52s
Deploy (staging -> E2E gate -> production) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m16s
Deploy (staging -> E2E gate -> production) / Deploy to Staging (push) Successful in 1s
Deploy (staging -> E2E gate -> production) / E2E Journeys against Staging (push) Failing after 4m47s
Deploy (staging -> E2E gate -> production) / Promote to Production (push) Has been skipped
Reviewed-on: #24
2026-07-09 13:32:43 +00:00
TudorandClaude Fable 5 c26755750d docs: mark GIAS code dictionaries spec implemented
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m38s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 54s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m58s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 14:12:52 +01:00
TudorandClaude Fable 5 254a19eb42 fix(pipeline): run gias_code_names seed + drift test in the daily DAG
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 14:11:31 +01:00
TudorandClaude Fable 5 4f6b2b0edc feat(pipeline): typesense sync translates GIAS codes before indexing
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 10:51:40 +01:00
TudorandClaude Fable 5 f1a013ec01 feat(api): translate GIAS codes to names at the query boundary
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 10:48:53 +01:00
TudorandClaude Fable 5 fa6c929a3a feat(pipeline): dim_school/dim_location store GIAS codes; seed drift test
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 10:45:15 +01:00
TudorandClaude Fable 5 d898e6279b feat(pipeline): ingest GIAS code columns; staging exposes codes not names
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 10:41:52 +01:00
TudorandClaude Fable 5 e188c2ff4b feat: GIAS code->name dictionaries generated from live bulk CSV
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 10:38:19 +01:00
TudorandClaude Fable 5 08bd86db05 docs: implementation plan for GIAS code dictionaries
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 10:10:26 +01:00
TudorandClaude Fable 5 1ae5762a0a docs: design spec for GIAS code dictionaries (codes in marts, names in code)
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-09 09:55:53 +01:00
tudor bc87e56545 Merge pull request 'fix(ui): shorten proposed-to-close notice copy' (#23) from fix/proposed-to-close-copy into main
Deploy (staging -> E2E gate -> production) / Build Backend (FastAPI) (push) Successful in 12s
Deploy (staging -> E2E gate -> production) / Build Frontend (Next.js) (push) Successful in 48s
Deploy (staging -> E2E gate -> production) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Deploy (staging -> E2E gate -> production) / Deploy to Staging (push) Successful in 1s
Deploy (staging -> E2E gate -> production) / E2E Journeys against Staging (push) Successful in 38s
Deploy (staging -> E2E gate -> production) / Promote to Production (push) Successful in 9s
Reviewed-on: #23
2026-07-08 21:52:56 +00:00
TudorandClaude Fable 5 4522cbf645 fix(ui): shorten proposed-to-close notice copy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m38s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 41s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 27s
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 22:51:32 +01:00
tudor 7370712888 Merge pull request 'feat: include and mark 'Open, but proposed to close' schools' (#22) from feat/proposed-to-close-schools into main
Deploy (staging -> E2E gate -> production) / Build Backend (FastAPI) (push) Successful in 20s
Deploy (staging -> E2E gate -> production) / Build Frontend (Next.js) (push) Successful in 55s
Deploy (staging -> E2E gate -> production) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m10s
Deploy (staging -> E2E gate -> production) / Deploy to Staging (push) Successful in 1s
Deploy (staging -> E2E gate -> production) / E2E Journeys against Staging (push) Successful in 41s
Deploy (staging -> E2E gate -> production) / Promote to Production (push) Successful in 10s
Reviewed-on: #22
2026-07-08 21:23:23 +00:00
TudorandClaude Fable 5 45ab479062 feat(ui): mark proposed-to-close schools in listings and detail pages
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 9m46s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 21s
PR Checks / Build Frontend (no push) (pull_request) Successful in 50s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 35s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m1s
Amber tag in listing rows (option A) and a slim notice strip under the
detail-page header (option E): proposed for closure, formal process not
necessarily started, check with the local authority before applying.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 22:05:55 +01:00
TudorandClaude Fable 5 6f602f4a9e feat(api): expose GIAS establishment status on school payloads
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 22:05:55 +01:00
TudorandClaude Fable 5 de81e9cdbd feat(pipeline): include 'Open, but proposed to close' schools in dims
These schools are still operating and publish results; they drop out
automatically when GIAS flips them to Closed since marts fully rebuild
each run.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-08 21:54:08 +01:00
tudor 45c68b60b4 Merge pull request 'feat: drive sixth-form separation from GIAS OfficialSixthForm flag' (#21) from feat/gias-sixth-form-flag into main
Deploy (staging -> E2E gate -> production) / Build Backend (FastAPI) (push) Successful in 20s
Deploy (staging -> E2E gate -> production) / Build Frontend (Next.js) (push) Successful in 50s
Deploy (staging -> E2E gate -> production) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m23s
Deploy (staging -> E2E gate -> production) / Deploy to Staging (push) Successful in 0s
Deploy (staging -> E2E gate -> production) / E2E Journeys against Staging (push) Successful in 44s
Deploy (staging -> E2E gate -> production) / Promote to Production (push) Successful in 9s
Reviewed-on: #21
2026-07-07 13:48:12 +00:00
52 changed files with 3796 additions and 91 deletions
+1
View File
@@ -604,6 +604,7 @@ async def get_school_details(request: Request, urn: int):
"religious_denomination": latest.get("religious_denomination", ""), "religious_denomination": latest.get("religious_denomination", ""),
"age_range": latest.get("age_range", ""), "age_range": latest.get("age_range", ""),
"has_sixth_form": latest.get("has_sixth_form"), "has_sixth_form": latest.get("has_sixth_form"),
"status": latest.get("status"),
"latitude": latest.get("latitude"), "latitude": latest.get("latitude"),
"longitude": latest.get("longitude"), "longitude": latest.get("longitude"),
"phase": latest.get("phase"), "phase": latest.get("phase"),
+107 -15
View File
@@ -4,6 +4,7 @@ Provides efficient queries with caching.
""" """
import logging import logging
import re
import pandas as pd import pandas as pd
import numpy as np import numpy as np
@@ -21,6 +22,38 @@ from .models import (
FactDeprivation, FactFinance, FactPupilCharacteristics, FactDeprivation, FactFinance, FactPupilCharacteristics,
) )
from .schemas import SCHOOL_TYPE_MAP from .schemas import SCHOOL_TYPE_MAP
from .gias_codes import (
ADMISSIONS_POLICY,
ESTABLISHMENT_STATUS,
PHASE_OF_EDUCATION,
RELIGIOUS_CHARACTER,
SCHOOL_TYPE,
translate,
)
# mart code column -> (API name column, dictionary)
_GIAS_CODE_COLUMNS = {
"phase_code": ("phase", PHASE_OF_EDUCATION),
"school_type_code": ("school_type", SCHOOL_TYPE),
"status_code": ("status", ESTABLISHMENT_STATUS),
"religious_character_code": ("religious_denomination", RELIGIOUS_CHARACTER),
"admissions_policy_code": ("admissions_policy", ADMISSIONS_POLICY),
}
def translate_gias_code_columns(df: pd.DataFrame) -> pd.DataFrame:
"""Map GIAS code columns to today's name columns (API contract).
Runs immediately after pd.read_sql so every downstream consumer —
filters, PHASE_GROUPS, payloads, /api/filters — keeps seeing names.
DataFrames without the code columns (old schema, test fixtures) pass
through unchanged.
"""
for code_col, (name_col, mapping) in _GIAS_CODE_COLUMNS.items():
if code_col in df.columns:
df[name_col] = df[code_col].map(lambda c: translate(c, mapping))
return df
_postcode_cache: Dict[str, Tuple[float, float]] = {} _postcode_cache: Dict[str, Tuple[float, float]] = {}
_typesense_client = None _typesense_client = None
@@ -121,15 +154,16 @@ _MAIN_QUERY = text("""
SELECT SELECT
s.urn, s.urn,
s.school_name, s.school_name,
s.phase, s.phase_code,
s.school_type, s.school_type_code,
s.academy_trust_name AS trust_name, s.academy_trust_name AS trust_name,
s.academy_trust_uid AS trust_uid, s.academy_trust_uid AS trust_uid,
s.religious_character AS religious_denomination, s.religious_character_code,
s.gender, s.gender,
s.age_range, s.age_range,
s.has_sixth_form, s.has_sixth_form,
s.admissions_policy, s.status_code,
s.admissions_policy_code,
s.capacity, s.capacity,
s.total_pupils AS gias_total_pupils, s.total_pupils AS gias_total_pupils,
s.headteacher_name, s.headteacher_name,
@@ -229,25 +263,81 @@ assert "NULL AS has_sixth_form" in str(_MAIN_QUERY_NO_SIXTH_FORM), (
"expected replacement of 's.has_sixth_form,' to have taken effect" "expected replacement of 's.has_sixth_form,' to have taken effect"
) )
# Fallback used when marts.dim_school predates the GIAS code-dictionary
# migration (i.e. the nightly dbt pipeline hasn't rebuilt the mart yet on
# this DB, so it still has the old name columns instead of *_code columns).
_MAIN_QUERY_LEGACY_NAMES = str(_MAIN_QUERY)
_LEGACY_NAME_REPLACEMENTS = [
("s.phase_code,", "s.phase,"),
("s.school_type_code,", "s.school_type,"),
(
"s.religious_character_code,",
"s.religious_character AS religious_denomination,",
),
("s.status_code,", "s.status,"),
("s.admissions_policy_code,", "s.admissions_policy,"),
]
for _old, _new in _LEGACY_NAME_REPLACEMENTS:
assert _old in _MAIN_QUERY_LEGACY_NAMES, (
f"expected {_old!r} to be present in _MAIN_QUERY before replacement"
)
_MAIN_QUERY_LEGACY_NAMES = _MAIN_QUERY_LEGACY_NAMES.replace(_old, _new)
_MAIN_QUERY_LEGACY_NAMES = text(_MAIN_QUERY_LEGACY_NAMES)
_GIAS_CODE_COLUMN_NAMES = (
"phase_code",
"school_type_code",
"religious_character_code",
"status_code",
"admissions_policy_code",
)
_MISSING_COLUMN_RE = re.compile(r'column "?(?:s\.)?(\w+)"? does not exist')
def _missing_column_name(exc: Exception) -> Optional[str]:
"""Name of the missing column from a psycopg2 UndefinedColumn error.
Inspects exc.orig (the DBAPI error), whose message names only the
offending column — str(exc) also embeds the full SQL statement, which
contains every column name and therefore must not be matched against.
"""
orig = getattr(exc, "orig", None)
match = _MISSING_COLUMN_RE.search(str(orig) if orig is not None else str(exc))
return match.group(1) if match else None
def load_school_data_as_dataframe() -> pd.DataFrame: def load_school_data_as_dataframe() -> pd.DataFrame:
"""Load all school + KS2 data as a pandas DataFrame.""" """Load all school + KS2 data as a pandas DataFrame."""
try: try:
df = pd.read_sql(_MAIN_QUERY, engine) df = pd.read_sql(_MAIN_QUERY, engine)
except sqlalchemy.exc.ProgrammingError as exc: except sqlalchemy.exc.ProgrammingError as exc:
if "has_sixth_form" not in str(exc): missing = _missing_column_name(exc)
if missing in _GIAS_CODE_COLUMN_NAMES:
logging.getLogger(__name__).warning(
"marts predate the GIAS code migration — falling back to "
"legacy name-column query: %s",
exc,
)
try:
df = pd.read_sql(_MAIN_QUERY_LEGACY_NAMES, engine)
except Exception as exc2:
print(f"Warning: Could not load school data from marts: {exc2}")
return pd.DataFrame()
elif missing == "has_sixth_form":
logging.getLogger(__name__).warning(
"marts.dim_school is missing has_sixth_form (pipeline hasn't "
"rebuilt the mart yet on this DB) — retrying without it: %s",
exc,
)
try:
df = pd.read_sql(_MAIN_QUERY_NO_SIXTH_FORM, engine)
except Exception as exc2:
print(f"Warning: Could not load school data from marts: {exc2}")
return pd.DataFrame()
else:
print(f"Warning: Could not load school data from marts: {exc}") print(f"Warning: Could not load school data from marts: {exc}")
return pd.DataFrame() return pd.DataFrame()
logging.getLogger(__name__).warning(
"marts.dim_school is missing has_sixth_form (pipeline hasn't "
"rebuilt the mart yet on this DB) — retrying without it: %s",
exc,
)
try:
df = pd.read_sql(_MAIN_QUERY_NO_SIXTH_FORM, engine)
except Exception as exc2:
print(f"Warning: Could not load school data from marts: {exc2}")
return pd.DataFrame()
except Exception as exc: except Exception as exc:
print(f"Warning: Could not load school data from marts: {exc}") print(f"Warning: Could not load school data from marts: {exc}")
return pd.DataFrame() return pd.DataFrame()
@@ -255,6 +345,8 @@ def load_school_data_as_dataframe() -> pd.DataFrame:
if df.empty: if df.empty:
return df return df
df = translate_gias_code_columns(df)
# Build address string # Build address string
df["address"] = df.apply( df["address"] = df.apply(
lambda r: ", ".join( lambda r: ", ".join(
+155
View File
@@ -0,0 +1,155 @@
"""GIAS code -> name dictionaries.
GENERATED by pipeline/scripts/generate_gias_codes.py from the GIAS bulk CSV
— do not edit by hand; rerun the script when the dbt drift test warns.
The canonical file is backend/gias_codes.py; pipeline/scripts/gias_codes.py
must be byte-identical (enforced by backend/tests/test_gias_codes.py).
"""
from __future__ import annotations
import logging
import math
logger = logging.getLogger(__name__)
SCHOOL_TYPE: dict[int, str] = {
1: "Community school",
2: "Voluntary aided school",
3: "Voluntary controlled school",
5: "Foundation school",
6: "City technology college",
7: "Community special school",
8: "Non-maintained special school",
10: "Other independent special school",
11: "Other independent school",
12: "Foundation special school",
14: "Pupil referral unit",
15: "Local authority nursery school",
18: "Further education",
24: "Secure units",
25: "Offshore schools",
26: "Service children's education",
27: "Miscellaneous",
28: "Academy sponsor led",
29: "Higher education institutions",
30: "Welsh establishment",
31: "Sixth form centres",
32: "Special post 16 institution",
33: "Academy special sponsor led",
34: "Academy converter",
35: "Free schools",
36: "Free schools special",
37: "British schools overseas",
38: "Free schools alternative provision",
39: "Free schools 16 to 19",
40: "University technical college",
41: "Studio schools",
42: "Academy alternative provision converter",
43: "Academy alternative provision sponsor led",
44: "Academy special converter",
45: "Academy 16-19 converter",
46: "Academy 16 to 19 sponsor led",
49: "Online provider",
56: "Institution funded by other government department",
57: "Academy secure 16 to 19",
}
ESTABLISHMENT_STATUS: dict[int, str] = {
1: "Open",
2: "Closed",
3: "Open, but proposed to close",
4: "Proposed to open",
}
PHASE_OF_EDUCATION: dict[int, str] = {
0: "Not applicable",
1: "Nursery",
2: "Primary",
3: "Middle deemed primary",
4: "Secondary",
5: "Middle deemed secondary",
6: "16 plus",
7: "All-through",
}
OFFICIAL_SIXTH_FORM: dict[int, str] = {
0: "Not applicable",
1: "Has a sixth form",
2: "Does not have a sixth form",
9: "",
}
RELIGIOUS_CHARACTER: dict[int, str] = {
0: "Does not apply",
2: "Church of England",
3: "Roman Catholic",
4: "Methodist",
5: "Jewish",
6: "None",
7: "Muslim",
8: "Seventh Day Adventist",
9: "Church of England/Methodist",
10: "Methodist/Church of England",
11: "Church of England/Roman Catholic",
12: "Church of England/United Reformed Church",
13: "Roman Catholic/Church of England",
14: "Quaker",
15: "Christian",
16: "United Reformed Church",
17: "Congregational Church",
18: "Free Church",
19: "Church of England/Free Church",
20: "Church of England/Christian",
21: "Sikh",
22: "Greek Orthodox",
24: "Buddhist",
25: "Hindu",
26: "Moravian",
28: "Inter- / non- denominational",
29: "Multi-faith",
30: "Church of England/Methodist/United Reform Church/Baptist",
31: "Anglican",
32: "Anglican/Christian",
33: "Anglican/Evangelical",
34: "Anglican/Church of England",
35: "Catholic",
36: "Charadi Jewish",
37: "Christian/Evangelical",
38: "Christian Science",
39: "Christian/Methodist",
40: "Christian/non-denominational",
41: "Church of England/Evangelical",
42: "Islam",
43: "Orthodox Jewish",
44: "Plymouth Brethren Christian Church",
45: "Protestant",
46: "Protestant/Evangelical",
47: "Reformed Baptist",
48: "Roman Catholic/Anglican",
49: "Sunni Deobandi",
99: "",
}
ADMISSIONS_POLICY: dict[int, str] = {
0: "Not applicable",
2: "Selective",
4: "Non-selective",
9: "",
}
def translate(code, mapping: dict[int, str]) -> str | None:
"""Translate a GIAS code to its display name.
None/NaN -> None (column absent or suppressed). Unknown codes degrade to
"Unknown (<code>)" with a warning so a new DfE value never blanks the UI.
"""
if code is None or (isinstance(code, float) and math.isnan(code)):
return None
code = int(code)
if code not in mapping:
logger.warning("Unknown GIAS code %s (not in dictionary)", code)
return f"Unknown ({code})"
return mapping[code]
+5 -5
View File
@@ -17,11 +17,11 @@ class DimSchool(Base):
urn = Column(Integer, primary_key=True) urn = Column(Integer, primary_key=True)
school_name = Column(String(255), nullable=False) school_name = Column(String(255), nullable=False)
phase = Column(String(100)) phase_code = Column(Integer)
school_type = Column(String(100)) school_type_code = Column(Integer)
academy_trust_name = Column(String(255)) academy_trust_name = Column(String(255))
academy_trust_uid = Column(String(20)) academy_trust_uid = Column(String(20))
religious_character = Column(String(100)) religious_character_code = Column(Integer)
gender = Column(String(20)) gender = Column(String(20))
age_range = Column(String(20)) age_range = Column(String(20))
has_sixth_form = Column(Boolean) has_sixth_form = Column(Boolean)
@@ -30,9 +30,9 @@ class DimSchool(Base):
headteacher_name = Column(String(200)) headteacher_name = Column(String(200))
website = Column(String(255)) website = Column(String(255))
telephone = Column(String(30)) telephone = Column(String(30))
status = Column(String(50)) status_code = Column(Integer)
nursery_provision = Column(Boolean) nursery_provision = Column(Boolean)
admissions_policy = Column(String(50)) admissions_policy_code = Column(Integer)
# Denormalised Ofsted summary (updated by monthly pipeline) # Denormalised Ofsted summary (updated by monthly pipeline)
ofsted_grade = Column(Integer) ofsted_grade = Column(Integer)
ofsted_date = Column(Date) ofsted_date = Column(Date)
+1
View File
@@ -544,6 +544,7 @@ SCHOOL_COLUMNS = [
"religious_denomination", "religious_denomination",
"age_range", "age_range",
"has_sixth_form", "has_sixth_form",
"status",
"gender", "gender",
"admissions_policy", "admissions_policy",
"ofsted_grade", "ofsted_grade",
+95
View File
@@ -0,0 +1,95 @@
"""Tests for the GIAS code->name dictionaries (spec 2026-07-09).
The dictionaries are generated from the live GIAS bulk CSV by
pipeline/scripts/generate_gias_codes.py — these tests assert the module's
contract, key sentinel values the marts/UI depend on, and that the pipeline
copy has not drifted from the canonical backend module.
"""
import math
from pathlib import Path
from backend.gias_codes import (
ADMISSIONS_POLICY,
ESTABLISHMENT_STATUS,
OFFICIAL_SIXTH_FORM,
PHASE_OF_EDUCATION,
RELIGIOUS_CHARACTER,
SCHOOL_TYPE,
translate,
)
REPO = Path(__file__).resolve().parents[2]
def test_translate_known_code():
open_code = next(c for c, n in ESTABLISHMENT_STATUS.items() if n == "Open")
assert translate(open_code, ESTABLISHMENT_STATUS) == "Open"
def test_translate_unknown_code_degrades_gracefully():
assert translate(9999, ESTABLISHMENT_STATUS) == "Unknown (9999)"
def test_translate_none_and_nan_return_none():
assert translate(None, ESTABLISHMENT_STATUS) is None
assert translate(float("nan"), ESTABLISHMENT_STATUS) is None
def test_translate_accepts_float_codes():
# pd.read_sql yields float columns when NULLs are present
open_code = next(c for c, n in ESTABLISHMENT_STATUS.items() if n == "Open")
assert translate(float(open_code), ESTABLISHMENT_STATUS) == "Open"
def test_sentinel_names_present():
"""Names the marts/UI compare against must exist verbatim."""
assert "Open" in ESTABLISHMENT_STATUS.values()
assert "Open, but proposed to close" in ESTABLISHMENT_STATUS.values()
assert "Has a sixth form" in OFFICIAL_SIXTH_FORM.values()
assert "Primary" in PHASE_OF_EDUCATION.values()
assert "Secondary" in PHASE_OF_EDUCATION.values()
assert "Does not apply" in RELIGIOUS_CHARACTER.values()
assert all(len(d) > 0 for d in (
SCHOOL_TYPE, ESTABLISHMENT_STATUS, PHASE_OF_EDUCATION,
OFFICIAL_SIXTH_FORM, RELIGIOUS_CHARACTER, ADMISSIONS_POLICY,
))
def test_pipeline_copy_is_identical():
canonical = (REPO / "backend" / "gias_codes.py").read_text()
copy = (REPO / "pipeline" / "scripts" / "gias_codes.py").read_text()
assert canonical == copy, (
"pipeline/scripts/gias_codes.py has drifted from backend/gias_codes.py — "
"regenerate with pipeline/scripts/generate_gias_codes.py and copy the file"
)
def test_seed_matches_dictionaries():
import csv
fields = {
"school_type": SCHOOL_TYPE,
"establishment_status": ESTABLISHMENT_STATUS,
"phase_of_education": PHASE_OF_EDUCATION,
"official_sixth_form": OFFICIAL_SIXTH_FORM,
"religious_character": RELIGIOUS_CHARACTER,
"admissions_policy": ADMISSIONS_POLICY,
}
seed_path = REPO / "pipeline" / "transform" / "seeds" / "gias_code_names.csv"
seed: dict[str, dict[int, str]] = {k: {} for k in fields}
with open(seed_path, newline="") as fh:
for row in csv.DictReader(fh):
seed[row["field"]][int(row["code"])] = row["name"]
assert seed == fields
def test_blank_name_sentinel_codes_map_to_empty_string():
"""GIAS carries codes whose (name) column is blank — e.g. ReligiousCharacter
99 (~4k schools) and AdmissionsPolicy 9 (~5.6k schools). The old name
pipeline served these as empty strings; the dictionaries must reproduce
that ("" is falsy, so UI tag heuristics stay silent) rather than letting
them hit the "Unknown (<code>)" path meant for genuinely new codes."""
assert RELIGIOUS_CHARACTER[99] == ""
assert ADMISSIONS_POLICY[9] == ""
assert translate(99, RELIGIOUS_CHARACTER) == ""
assert translate(9, ADMISSIONS_POLICY) == ""
+132
View File
@@ -0,0 +1,132 @@
"""API-boundary translation: marts now carry GIAS codes; the DataFrame the
rest of the backend sees must carry today's name strings."""
import numpy as np
import pandas as pd
from backend.data_loader import _missing_column_name, translate_gias_code_columns
from backend.gias_codes import ESTABLISHMENT_STATUS, PHASE_OF_EDUCATION
def _code_for(mapping, name):
return next(c for c, n in mapping.items() if n == name)
def test_codes_become_todays_names():
df = pd.DataFrame([{
"urn": 1,
"phase_code": float(_code_for(PHASE_OF_EDUCATION, "Primary")),
"school_type_code": np.nan,
"status_code": float(_code_for(ESTABLISHMENT_STATUS, "Open, but proposed to close")),
"religious_character_code": np.nan,
"admissions_policy_code": np.nan,
}])
out = translate_gias_code_columns(df)
row = out.iloc[0]
assert row["phase"] == "Primary"
assert row["status"] == "Open, but proposed to close"
assert row["school_type"] is None
assert row["religious_denomination"] is None
assert row["admissions_policy"] is None
def test_unknown_code_degrades_not_blanks():
df = pd.DataFrame([{"urn": 1, "phase_code": 9999.0}])
out = translate_gias_code_columns(df)
assert out.iloc[0]["phase"] == "Unknown (9999)"
def test_missing_code_columns_are_a_noop():
"""Old-schema DataFrames (tests, pre-pipeline DBs) pass through untouched."""
df = pd.DataFrame([{"urn": 1, "phase": "Primary", "status": "Open"}])
out = translate_gias_code_columns(df)
assert out.iloc[0]["phase"] == "Primary"
assert out.iloc[0]["status"] == "Open"
def _fake_exc(orig_message):
"""A stand-in for sqlalchemy.exc.ProgrammingError: str(exc) embeds the
full SQL statement (deliberately containing every column name below, to
prove the matcher doesn't fall back to it), while .orig carries the real
DBAPI error message naming only the offending column."""
exc = Exception(
"SELECT s.phase_code, s.school_type_code, s.religious_character_code, "
"s.status_code, s.admissions_policy_code, s.has_sixth_form FROM ... "
f"[SQL: ...] (Background on this error at: https://...)"
)
exc.orig = Exception(orig_message) if orig_message is not None else None
return exc
def test_missing_column_name_quoted():
assert _missing_column_name(_fake_exc('column "phase_code" does not exist')) == "phase_code"
def test_missing_column_name_unquoted():
assert _missing_column_name(_fake_exc("column phase_code does not exist")) == "phase_code"
def test_missing_column_name_table_prefixed():
assert (
_missing_column_name(_fake_exc("column s.has_sixth_form does not exist"))
== "has_sixth_form"
)
def test_missing_column_name_no_match_returns_none():
assert _missing_column_name(_fake_exc("relation \"marts.dim_school\" does not exist")) is None
def test_load_school_data_survives_premigration_marts(monkeypatch):
"""Real prod state until the nightly pipeline first rebuilds the mart with
the GIAS code columns: marts.dim_school still has the old name columns
(phase, school_type, religious_character, status, admissions_policy)
instead of the new *_code columns. The first query raises UndefinedColumn
on s.phase_code; load_school_data_as_dataframe must retry with the
legacy name-column query rather than swallow the error and return (and
then have load_school_data cache) an empty DataFrame."""
import sqlalchemy.exc
from backend import data_loader
data_loader._df_cache = None
data_loader._df_latest_cache = None
good_df = pd.DataFrame(
[
{
"urn": 1,
"school_name": "Legacy School",
"phase": "Primary",
"school_type": "Academy",
"status": "Open",
}
]
)
calls = []
def fake_read_sql(query, con):
calls.append(query)
if len(calls) == 1:
raise sqlalchemy.exc.ProgrammingError(
statement=str(data_loader._MAIN_QUERY),
params=None,
orig=Exception(
"(psycopg2.errors.UndefinedColumn) column s.phase_code "
"does not exist\nLINE 5: s.phase_code,"
),
)
return good_df.copy()
monkeypatch.setattr(data_loader.pd, "read_sql", fake_read_sql)
try:
df = data_loader.load_school_data_as_dataframe()
finally:
data_loader._df_cache = None
data_loader._df_latest_cache = None
assert len(calls) == 2, "must retry with the legacy name-column query variant"
assert calls[1] is data_loader._MAIN_QUERY_LEGACY_NAMES
assert not df.empty
assert df["phase"].iloc[0] == "Primary"
assert df["status"].iloc[0] == "Open"
+70
View File
@@ -0,0 +1,70 @@
"""Tests for GIAS establishment status exposure.
"Open, but proposed to close" schools are now kept by the dims; the API must
surface `status` on list items and school_info so the UI can render the
proposed-to-close marker (listing tag) and notice strip (detail page).
"""
import numpy as np
import pandas as pd
import pytest
from fastapi.testclient import TestClient
PROPOSED = "Open, but proposed to close"
def _schools_df() -> pd.DataFrame:
base = {
"local_authority": "Testshire",
"school_type": "Academy",
"phase": "Secondary",
"address": "1 Test Street",
"town": "Testtown",
"postcode": "TS1 1AA",
"religious_denomination": None,
"gender": "Mixed",
"age_range": "11-16",
"admissions_policy": None,
"has_sixth_form": False,
"ofsted_grade": np.nan,
"ofsted_date": None,
"ofsted_framework": None,
"latitude": 51.5,
"longitude": -0.1,
"year": 202425,
"total_pupils": 800,
"rwm_expected_pct": np.nan,
"attainment_8_score": 48.0,
}
return pd.DataFrame(
[
{**base, "urn": 200001, "school_name": "Alpha Academy",
"status": "Open"},
{**base, "urn": 200002, "school_name": "Sarson High School",
"status": PROPOSED},
]
)
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_latest_school_data", _schools_df)
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
return TestClient(app_module.app, raise_server_exceptions=False)
def test_list_payload_includes_status(client):
resp = client.get("/api/schools")
assert resp.status_code == 200, resp.text
by_urn = {s["urn"]: s for s in resp.json()["schools"]}
assert by_urn[200001]["status"] == "Open"
assert by_urn[200002]["status"] == PROPOSED
def test_detail_payload_includes_status(client):
resp = client.get("/api/schools/200002")
assert resp.status_code == 200, resp.text
assert resp.json()["school_info"]["status"] == PROPOSED
+7 -3
View File
@@ -148,10 +148,14 @@ def test_load_school_data_survives_missing_has_sixth_form_column(monkeypatch):
def fake_read_sql(query, con): def fake_read_sql(query, con):
calls.append(query) calls.append(query)
if len(calls) == 1: if len(calls) == 1:
# The statement text still contains phase_code, school_type_code,
# etc. (it's the full _MAIN_QUERY SELECT list) — that's exactly
# the collision this test guards against: matching must be done
# against exc.orig (the DBAPI error), not str(exc)/the statement.
raise sqlalchemy.exc.ProgrammingError( raise sqlalchemy.exc.ProgrammingError(
"SELECT ...", statement=str(data_loader._MAIN_QUERY),
None, params=None,
Exception( orig=Exception(
"(psycopg2.errors.UndefinedColumn) column s.has_sixth_form " "(psycopg2.errors.UndefinedColumn) column s.has_sixth_form "
"does not exist" "does not exist"
), ),
@@ -0,0 +1,799 @@
# GIAS Code Dictionaries Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Store the six GIAS classification fields as official DfE integer codes in the marts and translate code → name in application code, leaving the API contract (name strings) unchanged.
**Architecture:** A generation script downloads the public GIAS bulk CSV and emits the dictionaries (Python dicts + a dbt seed) from real data. The tap ingests the `(code)` columns, staging casts them, `dim_school`/`dim_location` keep only codes, and translation happens in exactly two places: `backend/data_loader.py` right after `pd.read_sql`, and `pipeline/scripts/sync_typesense.py` before indexing. A dbt seed test warns when DfE adds/renames a value; a parity test keeps the backend and pipeline dictionary copies identical.
**Tech Stack:** Singer SDK tap, dbt (Postgres), FastAPI + pandas, Typesense sync script, pytest.
**Spec:** `docs/superpowers/specs/2026-07-09-gias-code-dictionaries-design.md`
## Global Constraints
- **Numeric code values are never assumed.** Every literal code used in SQL or yml (status filter, sixth-form derivation, phase cascade) must be verified against `pipeline/transform/seeds/gias_code_names.csv` generated in Task 1 from the live CSV. The literals written in this plan are best-current-knowledge and each carries a verification step.
- **Names served by the API must stay byte-identical** to today's strings (e.g. `Does not apply`, `Open, but proposed to close`) — UI heuristics compare exact strings.
- The `(name)` columns stay declared in the tap and present in raw; staging stops exposing them.
- `dim_school` and `dim_location` status filters must stay identical (API inner-joins them).
- Backend tests run via: `uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests -v` (no local pytest exists).
- dbt cannot run locally — dbt changes are verified statically (grep / yaml parse) + CI.
- Never push to `main`. Work on branch `feat/gias-code-dictionaries` (branch off `docs/gias-code-dictionaries` so the spec is included, or off `main` if that has merged).
- Commits end with: `Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>`
- Deploy runbook (accepted window, spec §7): merge → deploy → trigger `school_data_daily` immediately. No code-level fallback for the old-schema window.
---
### Task 1: Dictionary generation script, canonical module, pipeline copy, seed
**Files:**
- Create: `pipeline/scripts/generate_gias_codes.py`
- Create: `backend/gias_codes.py` (content generated by the script)
- Create: `pipeline/scripts/gias_codes.py` (byte-identical copy)
- Create: `pipeline/transform/seeds/gias_code_names.csv` (generated)
- Test: `backend/tests/test_gias_codes.py`
**Interfaces:**
- Produces: `backend/gias_codes.py` exporting `SCHOOL_TYPE`, `ESTABLISHMENT_STATUS`, `PHASE_OF_EDUCATION`, `OFFICIAL_SIXTH_FORM`, `RELIGIOUS_CHARACTER`, `ADMISSIONS_POLICY` (each `dict[int, str]`) and `translate(code, mapping) -> str | None`. Task 4 imports these; Task 5 imports the pipeline copy; Task 3 reads code literals from the seed CSV.
- [ ] **Step 1: Write the failing tests**
Create `backend/tests/test_gias_codes.py`:
```python
"""Tests for the GIAS code->name dictionaries (spec 2026-07-09).
The dictionaries are generated from the live GIAS bulk CSV by
pipeline/scripts/generate_gias_codes.py — these tests assert the module's
contract, key sentinel values the marts/UI depend on, and that the pipeline
copy has not drifted from the canonical backend module.
"""
import math
from pathlib import Path
from backend.gias_codes import (
ADMISSIONS_POLICY,
ESTABLISHMENT_STATUS,
OFFICIAL_SIXTH_FORM,
PHASE_OF_EDUCATION,
RELIGIOUS_CHARACTER,
SCHOOL_TYPE,
translate,
)
REPO = Path(__file__).resolve().parents[2]
def test_translate_known_code():
open_code = next(c for c, n in ESTABLISHMENT_STATUS.items() if n == "Open")
assert translate(open_code, ESTABLISHMENT_STATUS) == "Open"
def test_translate_unknown_code_degrades_gracefully():
assert translate(9999, ESTABLISHMENT_STATUS) == "Unknown (9999)"
def test_translate_none_and_nan_return_none():
assert translate(None, ESTABLISHMENT_STATUS) is None
assert translate(float("nan"), ESTABLISHMENT_STATUS) is None
def test_translate_accepts_float_codes():
# pd.read_sql yields float columns when NULLs are present
open_code = next(c for c, n in ESTABLISHMENT_STATUS.items() if n == "Open")
assert translate(float(open_code), ESTABLISHMENT_STATUS) == "Open"
def test_sentinel_names_present():
"""Names the marts/UI compare against must exist verbatim."""
assert "Open" in ESTABLISHMENT_STATUS.values()
assert "Open, but proposed to close" in ESTABLISHMENT_STATUS.values()
assert "Has a sixth form" in OFFICIAL_SIXTH_FORM.values()
assert "Primary" in PHASE_OF_EDUCATION.values()
assert "Secondary" in PHASE_OF_EDUCATION.values()
assert "Does not apply" in RELIGIOUS_CHARACTER.values()
assert all(len(d) > 0 for d in (
SCHOOL_TYPE, ESTABLISHMENT_STATUS, PHASE_OF_EDUCATION,
OFFICIAL_SIXTH_FORM, RELIGIOUS_CHARACTER, ADMISSIONS_POLICY,
))
def test_pipeline_copy_is_identical():
canonical = (REPO / "backend" / "gias_codes.py").read_text()
copy = (REPO / "pipeline" / "scripts" / "gias_codes.py").read_text()
assert canonical == copy, (
"pipeline/scripts/gias_codes.py has drifted from backend/gias_codes.py — "
"regenerate with pipeline/scripts/generate_gias_codes.py and copy the file"
)
def test_seed_matches_dictionaries():
import csv
fields = {
"school_type": SCHOOL_TYPE,
"establishment_status": ESTABLISHMENT_STATUS,
"phase_of_education": PHASE_OF_EDUCATION,
"official_sixth_form": OFFICIAL_SIXTH_FORM,
"religious_character": RELIGIOUS_CHARACTER,
"admissions_policy": ADMISSIONS_POLICY,
}
seed_path = REPO / "pipeline" / "transform" / "seeds" / "gias_code_names.csv"
seed: dict[str, dict[int, str]] = {k: {} for k in fields}
with open(seed_path, newline="") as fh:
for row in csv.DictReader(fh):
seed[row["field"]][int(row["code"])] = row["name"]
assert seed == fields
```
- [ ] **Step 2: Run tests to verify they fail**
Run: `cd /Users/tudor/projects/school_compare && uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests/test_gias_codes.py -v`
Expected: FAIL at import — `ModuleNotFoundError: No module named 'backend.gias_codes'`.
- [ ] **Step 3: Write the generation script**
Create `pipeline/scripts/generate_gias_codes.py`:
```python
"""Generate GIAS code->name dictionaries from the live bulk CSV.
Writes:
- backend/gias_codes.py (canonical Python module)
- pipeline/scripts/gias_codes.py (byte-identical copy)
- pipeline/transform/seeds/gias_code_names.csv (dbt seed for drift test)
Run from the repo root whenever the dbt drift test warns that DfE
added/renamed a value: python pipeline/scripts/generate_gias_codes.py
"""
from __future__ import annotations
import io
import sys
from datetime import date, timedelta
from pathlib import Path
import pandas as pd
import requests
GIAS_URL = (
"https://ea-edubase-api-prod.azurewebsites.net"
"/edubase/downloads/public/edubasealldata{date}.csv"
)
# (CSV code column, CSV name column, python dict name, seed field key)
FIELDS = [
("TypeOfEstablishment (code)", "TypeOfEstablishment (name)", "SCHOOL_TYPE", "school_type"),
("EstablishmentStatus (code)", "EstablishmentStatus (name)", "ESTABLISHMENT_STATUS", "establishment_status"),
("PhaseOfEducation (code)", "PhaseOfEducation (name)", "PHASE_OF_EDUCATION", "phase_of_education"),
("OfficialSixthForm (code)", "OfficialSixthForm (name)", "OFFICIAL_SIXTH_FORM", "official_sixth_form"),
("ReligiousCharacter (code)", "ReligiousCharacter (name)", "RELIGIOUS_CHARACTER", "religious_character"),
("AdmissionsPolicy (code)", "AdmissionsPolicy (name)", "ADMISSIONS_POLICY", "admissions_policy"),
]
MODULE_HEADER = '''"""GIAS code -> name dictionaries.
GENERATED by pipeline/scripts/generate_gias_codes.py from the GIAS bulk CSV
— do not edit by hand; rerun the script when the dbt drift test warns.
The canonical file is backend/gias_codes.py; pipeline/scripts/gias_codes.py
must be byte-identical (enforced by backend/tests/test_gias_codes.py).
"""
from __future__ import annotations
import logging
import math
logger = logging.getLogger(__name__)
'''
MODULE_FOOTER = '''
def translate(code, mapping: dict[int, str]) -> str | None:
"""Translate a GIAS code to its display name.
None/NaN -> None (column absent or suppressed). Unknown codes degrade to
"Unknown (<code>)" with a warning so a new DfE value never blanks the UI.
"""
if code is None or (isinstance(code, float) and math.isnan(code)):
return None
code = int(code)
if code not in mapping:
logger.warning("Unknown GIAS code %s (not in dictionary)", code)
return f"Unknown ({code})"
return mapping[code]
'''
def download_csv() -> pd.DataFrame:
for day in (date.today(), date.today() - timedelta(days=1)):
url = GIAS_URL.format(date=day.strftime("%Y%m%d"))
print(f"Downloading {url}")
resp = requests.get(url, timeout=300)
if resp.status_code == 404:
continue
resp.raise_for_status()
return pd.read_csv(
io.StringIO(resp.content.decode("latin-1")),
dtype=str, keep_default_na=False,
)
sys.exit("GIAS CSV not available for today or yesterday")
def main() -> None:
repo = Path(__file__).resolve().parents[2]
df = download_csv()
module_parts = [MODULE_HEADER]
seed_rows: list[tuple[str, int, str]] = []
for code_col, name_col, dict_name, field_key in FIELDS:
pairs = (
df[[code_col, name_col]]
.loc[lambda d: (d[code_col] != "") & (d[name_col] != "")]
.drop_duplicates()
)
mapping = sorted((int(c), n) for c, n in pairs.itertuples(index=False))
dupes = len(mapping) - len({c for c, _ in mapping})
if dupes:
sys.exit(f"{code_col}: {dupes} codes map to multiple names — investigate before generating")
lines = [f"{dict_name}: dict[int, str] = {{"]
for code, name in mapping:
escaped = name.replace('"', '\\"')
lines.append(f' {code}: "{escaped}",')
lines.append("}\n")
module_parts.append("\n".join(lines))
seed_rows += [(field_key, code, name) for code, name in mapping]
module = "\n".join(module_parts) + MODULE_FOOTER
(repo / "backend" / "gias_codes.py").write_text(module)
(repo / "pipeline" / "scripts" / "gias_codes.py").write_text(module)
seed_path = repo / "pipeline" / "transform" / "seeds" / "gias_code_names.csv"
with open(seed_path, "w", newline="") as fh:
import csv
w = csv.writer(fh)
w.writerow(["field", "code", "name"])
w.writerows(seed_rows)
print(f"Wrote backend/gias_codes.py, pipeline/scripts/gias_codes.py, {seed_path.name}")
print("\nKey codes for the dbt work (Task 3):")
for field in ("establishment_status", "phase_of_education", "official_sixth_form"):
print(f" {field}:")
for f, code, name in seed_rows:
if f == field:
print(f" {code} = {name}")
if __name__ == "__main__":
main()
```
- [ ] **Step 4: Run the generator**
Run: `cd /Users/tudor/projects/school_compare && uv run --with pandas --with requests python pipeline/scripts/generate_gias_codes.py`
Expected: downloads the CSV (~100MB, may take a minute), writes the three files, and prints the status/phase/sixth-form code tables. **Record the printed code tables — Task 3 needs them.** If the download fails twice, report BLOCKED (no network or GIAS outage) rather than inventing dictionary content.
- [ ] **Step 5: Run the tests again**
Run: `cd /Users/tudor/projects/school_compare && uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests/test_gias_codes.py -v`
Expected: 7 passed. If `test_sentinel_names_present` fails, the GIAS vocabulary differs from expectations — inspect the generated module and report DONE_WITH_CONCERNS naming the differing value; do not edit the generated names.
- [ ] **Step 6: Commit**
```bash
git add pipeline/scripts/generate_gias_codes.py backend/gias_codes.py pipeline/scripts/gias_codes.py pipeline/transform/seeds/gias_code_names.csv backend/tests/test_gias_codes.py
git commit -m "feat: GIAS code->name dictionaries generated from live bulk CSV
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"
```
---
### Task 2: Tap ingests the (code) columns; staging exposes codes, drops names
**Files:**
- Modify: `pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/tap.py` (Singer schema)
- Modify: `pipeline/transform/models/staging/stg_gias_establishments.sql`
**Interfaces:**
- Produces: staging columns `school_type_code`, `status_code`, `phase_code`, `official_sixth_form_code`, `religious_character_code`, `admissions_policy_code` (all int) consumed by Task 3. Staging **stops exposing** `school_type`, `status`, `phase`, `official_sixth_form`, `religious_character`, `admissions_policy` (names stay in raw only).
- [ ] **Step 1: Add the six (code) properties to the Singer schema**
In `tap.py`, `GIASEstablishmentsStream.schema`, add each `(code)` property directly above its existing `(name)` sibling:
```python
th.Property("TypeOfEstablishment (code)", th.StringType),
th.Property("PhaseOfEducation (code)", th.StringType),
th.Property("EstablishmentStatus (code)", th.StringType),
th.Property("Gender (name)", ...) # existing line — for placement reference only
th.Property("ReligiousCharacter (code)", th.StringType),
th.Property("AdmissionsPolicy (code)", th.StringType),
th.Property("OfficialSixthForm (code)", th.StringType),
```
(The exact insertion order doesn't matter — the schema is a dict — but keep each `(code)` adjacent to its `(name)` for readability. Do NOT remove any `(name)` property.)
- [ ] **Step 2: Rewrite the six columns in staging**
In `stg_gias_establishments.sql` `renamed` CTE, replace:
```sql
"TypeOfEstablishment (name)" as school_type,
"PhaseOfEducation (name)" as phase,
nullif(trim("OfficialSixthForm (name)"), '') as official_sixth_form,
"ReligiousCharacter (name)" as religious_character,
"AdmissionsPolicy (name)" as admissions_policy,
"EstablishmentStatus (name)" as status,
```
with:
```sql
cast(nullif(trim("TypeOfEstablishment (code)"), '') as integer) as school_type_code,
cast(nullif(trim("PhaseOfEducation (code)"), '') as integer) as phase_code,
cast(nullif(trim("OfficialSixthForm (code)"), '') as integer) as official_sixth_form_code,
cast(nullif(trim("ReligiousCharacter (code)"), '') as integer) as religious_character_code,
cast(nullif(trim("AdmissionsPolicy (code)"), '') as integer) as admissions_policy_code,
cast(nullif(trim("EstablishmentStatus (code)"), '') as integer) as status_code,
```
(The name lines are scattered through the CTE — replace each in place; the six name aliases must no longer appear in the model.)
- [ ] **Step 3: Verify statically**
Run:
```bash
cd /Users/tudor/projects/school_compare && \
python3 -c "import ast; ast.parse(open('pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/tap.py').read()); print('tap OK')" && \
grep -c "(code)" pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/tap.py && \
grep -E "as (school_type|status|phase|official_sixth_form|religious_character|admissions_policy)," pipeline/transform/models/staging/stg_gias_establishments.sql; echo "name-alias grep exit=$? (want 1 = none found)"
```
Expected: `tap OK`, code-column count `6`, and the final grep finds nothing (exit 1).
- [ ] **Step 4: Commit**
```bash
git add pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/tap.py pipeline/transform/models/staging/stg_gias_establishments.sql
git commit -m "feat(pipeline): ingest GIAS code columns; staging exposes codes not names
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"
```
---
### Task 3: Marts store codes; dbt tests + drift test
**Files:**
- Modify: `pipeline/transform/models/marts/dim_school.sql`
- Modify: `pipeline/transform/models/marts/dim_location.sql`
- Modify: `pipeline/transform/models/marts/_marts_schema.yml`
- Create: `pipeline/transform/tests/assert_gias_code_names_match_seed.sql`
**Interfaces:**
- Consumes: staging code columns from Task 2; code literals from `pipeline/transform/seeds/gias_code_names.csv` (Task 1).
- Produces: `dim_school` columns `school_type_code`, `status_code`, `phase_code`, `religious_character_code`, `admissions_policy_code` (int) replacing their string columns; `has_sixth_form` unchanged (bool). Task 4's `_MAIN_QUERY` selects these.
**Before writing SQL: open `pipeline/transform/seeds/gias_code_names.csv` and confirm the literals below.** Best-current-knowledge values (VERIFY EACH):
`establishment_status`: 1 = Open, 3 = "Open, but proposed to close" (2 = Closed, 4 = Proposed to open).
`phase_of_education`: 0 = Not applicable, 2 = Primary, 4 = Secondary, 7 = All-through.
`official_sixth_form`: 1 = Has a sixth form, 2 = Does not have a sixth form, 0 = Not applicable.
If any differ, use the seed's values everywhere below and say so in your report.
- [ ] **Step 1: Rewrite dim_school.sql derivations in code space**
Replace the phase cascade block (`case ... end as phase,`) with:
```sql
-- Phase in GIAS code space (see seeds/gias_code_names.csv):
-- 2 = Primary, 4 = Secondary, 7 = All-through, 0 = Not applicable.
case
-- 1. Trust GIAS phase when it's a real value (0 = the catch-all "Not Applicable")
when s.phase_code is not null and s.phase_code != 0
then s.phase_code
-- 2. Infer from statutory age range (independent schools still publish these)
when s.statutory_high_age is not null and s.statutory_high_age <= 11 then 2
when s.statutory_low_age is not null and s.statutory_low_age >= 11 then 4
when s.statutory_low_age is not null and s.statutory_high_age is not null
and s.statutory_low_age < 11 and s.statutory_high_age > 11 then 7
-- 3. Fallback: infer from school name (covers independents with missing ages)
when s.school_name ilike '%primary%'
or s.school_name ilike '%infant%'
or s.school_name ilike '%junior%'
or s.school_name ilike '%preparatory%'
or s.school_name ilike '% prep school%'
or s.school_name ilike '% prep %'
then 2
when s.school_name ilike '%secondary%'
or s.school_name ilike '%high school%'
or s.school_name ilike '%grammar%'
or s.school_name ilike '%senior school%'
or s.school_name ilike '%upper school%'
then 4
-- 4. Give up — null renders no phase pill
else null
end as phase_code,
```
Replace `s.school_type,` with `s.school_type_code,`; `s.religious_character,` with `s.religious_character_code,`; `s.admissions_policy,` with `s.admissions_policy_code,`; `s.status,` with `s.status_code,`.
Replace the has_sixth_form case with:
```sql
-- GIAS OfficialSixthForm in code space: 1 = has, 2 = does not, 0 = N/A.
-- Null (rare, new establishments) falls back to the statutory age range.
case
when s.official_sixth_form_code = 1 then true
when s.official_sixth_form_code in (0, 2) then false
else coalesce(s.statutory_high_age >= 18, false)
end as has_sixth_form,
```
Replace the status filter with:
```sql
-- 1 = Open; 3 = Open, but proposed to close (still operating; drops out when
-- GIAS flips to Closed — marts fully rebuild each run).
where s.status_code in (1, 3)
```
- [ ] **Step 2: Same filter in dim_location.sql**
Replace its `where s.status in ('Open', 'Open, but proposed to close')` (and the comment above it) with:
```sql
-- Must match dim_school's status filter exactly (the API inner-joins the two).
where s.status_code in (1, 3)
```
- [ ] **Step 3: Update _marts_schema.yml**
Under `dim_school` columns: rename `phase``phase_code` (keep the warn-severity not_null, reword description to mention codes); replace the `status` accepted_values block with:
```yaml
- name: status_code
description: GIAS EstablishmentStatus code (1 = Open, 3 = Open but proposed to close)
tests:
- accepted_values:
values: [1, 3]
```
Add warn-severity accepted_values for the other codes, values copied from the seed (school_type/religious/admissions lists are long — paste the full code list from `gias_code_names.csv` for each):
```yaml
- name: school_type_code
tests:
- accepted_values:
severity: warn
values: [<all school_type codes from the seed>]
- name: religious_character_code
tests:
- accepted_values:
severity: warn
values: [<all religious_character codes from the seed>]
- name: admissions_policy_code
tests:
- accepted_values:
severity: warn
values: [<all admissions_policy codes from the seed>]
```
(`<...>` here means: paste the actual comma-separated integers from the seed file — the lists exist by the time this task runs. Leaving a literal `<...>` in the yml is a task failure.)
`has_sixth_form` tests stay unchanged.
- [ ] **Step 4: Write the drift test**
Create `pipeline/transform/tests/assert_gias_code_names_match_seed.sql`:
```sql
-- Warn when the live GIAS CSV carries a (code, name) pair we don't have in
-- the dictionary seed — i.e. DfE added or renamed a value. Fix by rerunning
-- pipeline/scripts/generate_gias_codes.py and committing the regenerated
-- dictionaries + seed together.
{{ config(severity='warn') }}
with raw_pairs as (
{% for field_key, code_col, name_col in [
('school_type', 'TypeOfEstablishment (code)', 'TypeOfEstablishment (name)'),
('establishment_status', 'EstablishmentStatus (code)', 'EstablishmentStatus (name)'),
('phase_of_education', 'PhaseOfEducation (code)', 'PhaseOfEducation (name)'),
('official_sixth_form', 'OfficialSixthForm (code)', 'OfficialSixthForm (name)'),
('religious_character', 'ReligiousCharacter (code)', 'ReligiousCharacter (name)'),
('admissions_policy', 'AdmissionsPolicy (code)', 'AdmissionsPolicy (name)')
] %}
select distinct
'{{ field_key }}' as field,
cast(nullif(trim("{{ code_col }}"), '') as integer) as code,
nullif(trim("{{ name_col }}"), '') as name
from {{ source('raw', 'gias_establishments') }}
where nullif(trim("{{ code_col }}"), '') is not null
and nullif(trim("{{ name_col }}"), '') is not null
{% if not loop.last %}union all{% endif %}
{% endfor %}
)
select r.*
from raw_pairs r
left join {{ ref('gias_code_names') }} s
on s.field = r.field
and s.code = r.code
and s.name = r.name
where s.field is null
```
- [ ] **Step 5: Verify statically**
Run:
```bash
cd /Users/tudor/projects/school_compare && \
uv run --with pyyaml python -c "import yaml; yaml.safe_load(open('pipeline/transform/models/marts/_marts_schema.yml')); print('yml OK')" && \
grep -c "_code" pipeline/transform/models/marts/dim_school.sql && \
grep -n "status_code in (1, 3)" pipeline/transform/models/marts/dim_school.sql pipeline/transform/models/marts/dim_location.sql && \
grep -rn "s\.status\b\|s\.phase\b\|s\.school_type\b\|s\.religious_character\b\|s\.admissions_policy\b\|official_sixth_form\b" pipeline/transform/models/marts/dim_school.sql | grep -v "_code"; echo "stale-name grep exit=$? (want 1)"
```
Expected: `yml OK`, both filters matched, and no stale name-column references (final grep exits 1).
- [ ] **Step 6: Commit**
```bash
git add pipeline/transform/models/marts/dim_school.sql pipeline/transform/models/marts/dim_location.sql pipeline/transform/models/marts/_marts_schema.yml pipeline/transform/tests/assert_gias_code_names_match_seed.sql
git commit -m "feat(pipeline): dim_school/dim_location store GIAS codes; seed drift test
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"
```
---
### Task 4: Backend translates at the API boundary
**Files:**
- Modify: `backend/models.py` (DimSchool columns)
- Modify: `backend/data_loader.py` (`_MAIN_QUERY` + translation)
- Test: `backend/tests/test_gias_translation.py` (new)
**Interfaces:**
- Consumes: `backend/gias_codes.py` dictionaries + `translate` (Task 1); mart code columns (Task 3).
- Produces: `translate_gias_code_columns(df) -> df` in `backend/data_loader.py`; after `load_school_data_as_dataframe()` the DataFrame carries today's name columns (`phase`, `school_type`, `status`, `religious_denomination`, `admissions_policy`) — every downstream consumer unchanged.
- [ ] **Step 1: Write the failing tests**
Create `backend/tests/test_gias_translation.py`:
```python
"""API-boundary translation: marts now carry GIAS codes; the DataFrame the
rest of the backend sees must carry today's name strings."""
import numpy as np
import pandas as pd
from backend.data_loader import translate_gias_code_columns
from backend.gias_codes import ESTABLISHMENT_STATUS, PHASE_OF_EDUCATION
def _code_for(mapping, name):
return next(c for c, n in mapping.items() if n == name)
def test_codes_become_todays_names():
df = pd.DataFrame([{
"urn": 1,
"phase_code": float(_code_for(PHASE_OF_EDUCATION, "Primary")),
"school_type_code": np.nan,
"status_code": float(_code_for(ESTABLISHMENT_STATUS, "Open, but proposed to close")),
"religious_character_code": np.nan,
"admissions_policy_code": np.nan,
}])
out = translate_gias_code_columns(df)
row = out.iloc[0]
assert row["phase"] == "Primary"
assert row["status"] == "Open, but proposed to close"
assert row["school_type"] is None
assert row["religious_denomination"] is None
assert row["admissions_policy"] is None
def test_unknown_code_degrades_not_blanks():
df = pd.DataFrame([{"urn": 1, "phase_code": 9999.0}])
out = translate_gias_code_columns(df)
assert out.iloc[0]["phase"] == "Unknown (9999)"
def test_missing_code_columns_are_a_noop():
"""Old-schema DataFrames (tests, pre-pipeline DBs) pass through untouched."""
df = pd.DataFrame([{"urn": 1, "phase": "Primary", "status": "Open"}])
out = translate_gias_code_columns(df)
assert out.iloc[0]["phase"] == "Primary"
assert out.iloc[0]["status"] == "Open"
```
- [ ] **Step 2: Run to verify failure**
Run: `cd /Users/tudor/projects/school_compare && uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests/test_gias_translation.py -v`
Expected: FAIL — `ImportError: cannot import name 'translate_gias_code_columns'`.
- [ ] **Step 3: Implement translation in data_loader.py**
Add near the top of `backend/data_loader.py` (after existing imports):
```python
from .gias_codes import (
ADMISSIONS_POLICY,
ESTABLISHMENT_STATUS,
PHASE_OF_EDUCATION,
RELIGIOUS_CHARACTER,
SCHOOL_TYPE,
translate,
)
# mart code column -> (API name column, dictionary)
_GIAS_CODE_COLUMNS = {
"phase_code": ("phase", PHASE_OF_EDUCATION),
"school_type_code": ("school_type", SCHOOL_TYPE),
"status_code": ("status", ESTABLISHMENT_STATUS),
"religious_character_code": ("religious_denomination", RELIGIOUS_CHARACTER),
"admissions_policy_code": ("admissions_policy", ADMISSIONS_POLICY),
}
def translate_gias_code_columns(df: pd.DataFrame) -> pd.DataFrame:
"""Map GIAS code columns to today's name columns (API contract).
Runs immediately after pd.read_sql so every downstream consumer —
filters, PHASE_GROUPS, payloads, /api/filters — keeps seeing names.
DataFrames without the code columns (old schema, test fixtures) pass
through unchanged.
"""
for code_col, (name_col, mapping) in _GIAS_CODE_COLUMNS.items():
if code_col in df.columns:
df[name_col] = df[code_col].map(lambda c: translate(c, mapping))
return df
```
- [ ] **Step 4: Switch `_MAIN_QUERY` to code columns and call the translation**
In `_MAIN_QUERY` replace:
`s.phase,``s.phase_code,` · `s.school_type,``s.school_type_code,` · `s.religious_character AS religious_denomination,``s.religious_character_code,` · `s.admissions_policy,``s.admissions_policy_code,` · `s.status,``s.status_code,`
In `load_school_data_as_dataframe()`, insert the call immediately after the empty-check and **before** the existing `normalize_school_type` line:
```python
if df.empty:
return df
df = translate_gias_code_columns(df)
# Build address string
...
# Normalize school type (existing line — now normalises the translated name)
df["school_type"] = df["school_type"].apply(normalize_school_type)
```
- [ ] **Step 5: Update DimSchool in models.py**
Replace `phase = Column(String(100))`, `school_type = Column(String(100))`, `religious_character = Column(String(100))`, `admissions_policy = Column(String(50))`, `status = Column(String(50))` with:
```python
phase_code = Column(Integer)
school_type_code = Column(Integer)
religious_character_code = Column(Integer)
admissions_policy_code = Column(Integer)
status_code = Column(Integer)
```
Then check nothing else references the removed attributes:
```bash
grep -rn "\.phase\b\|\.school_type\b\|\.religious_character\b\|\.admissions_policy\b\|\.status\b" backend/*.py | grep -i "dimschool\|DimSchool"
```
Expected: no hits (the backend reads via `_MAIN_QUERY`, not ORM attributes). If there are hits, update them to the `_code` columns + translation and note it in your report.
- [ ] **Step 6: Run the new tests and the whole backend suite**
Run: `cd /Users/tudor/projects/school_compare && uv run --with-requirements requirements.txt --with pytest --with "httpx==0.27.0" python -m pytest backend/tests -v`
Expected: all pass — 3 new + all pre-existing (their fixtures carry name columns; translation is a no-op on them).
- [ ] **Step 7: Commit**
```bash
git add backend/models.py backend/data_loader.py backend/tests/test_gias_translation.py
git commit -m "feat(api): translate GIAS codes to names at the query boundary
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"
```
---
### Task 5: Typesense sync translates before indexing
**Files:**
- Modify: `pipeline/scripts/sync_typesense.py`
**Interfaces:**
- Consumes: `pipeline/scripts/gias_codes.py` (Task 1), mart code columns (Task 3).
- Produces: identical Typesense documents to today (facet values are names).
- [ ] **Step 1: Switch the SELECT and translate**
In `sync_typesense.py`: add at the top (the DAG runs `python scripts/sync_typesense.py`, so `scripts/` is `sys.path[0]` and a plain import works):
```python
from gias_codes import PHASE_OF_EDUCATION, RELIGIOUS_CHARACTER, SCHOOL_TYPE, translate
```
In the SQL, replace `s.phase,``s.phase_code,`, `s.school_type,``s.school_type_code,`, `s.religious_character,``s.religious_character_code,`.
In the document builder, replace:
```python
"phase": row["phase"] or "",
"school_type": row["school_type"] or "",
```
with:
```python
"phase": translate(row["phase_code"], PHASE_OF_EDUCATION) or "",
"school_type": translate(row["school_type_code"], SCHOOL_TYPE) or "",
```
and:
```python
if row.get("religious_character"):
doc["religious_character"] = row["religious_character"]
```
with:
```python
religious_character = translate(row.get("religious_character_code"), RELIGIOUS_CHARACTER)
if religious_character:
doc["religious_character"] = religious_character
```
- [ ] **Step 2: Verify statically**
Run:
```bash
cd /Users/tudor/projects/school_compare && \
python3 -c "import ast; ast.parse(open('pipeline/scripts/sync_typesense.py').read()); print('sync OK')" && \
grep -n "row\[\"phase\"\]\|row\[\"school_type\"\]\|row\[\"religious_character\"\]" pipeline/scripts/sync_typesense.py; echo "stale grep exit=$? (want 1)"
```
Expected: `sync OK`, no stale name-column row accesses.
- [ ] **Step 3: Commit**
```bash
git add pipeline/scripts/sync_typesense.py
git commit -m "feat(pipeline): typesense sync translates GIAS codes before indexing
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"
```
---
### Task 6: Spec status, PR, deploy runbook
**Files:**
- Modify: `docs/superpowers/specs/2026-07-09-gias-code-dictionaries-design.md` (status line)
- [ ] **Step 1: Mark the spec implemented**
Change `**Status:** Approved design` to `**Status:** Implemented 2026-07-09 — see docs/superpowers/plans/2026-07-09-gias-code-dictionaries.md`.
- [ ] **Step 2: Commit and push**
```bash
git add docs/superpowers/specs/2026-07-09-gias-code-dictionaries-design.md
git commit -m "docs: mark GIAS code dictionaries spec implemented
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>"
git push -u origin feat/gias-code-dictionaries
```
- [ ] **Step 3: Open the PR (Gitea API via git credential fill — token-header auth 401s)**
Title: `feat: GIAS classification fields stored as codes, translated in code`
Body must include: (1) API contract unchanged — names still served, translation at the query boundary; (2) the **deploy runbook: merge → deploy → trigger `school_data_daily` immediately** (accepted empty-API window until the marts rebuild — spec §7); (3) dictionary maintenance loop (dbt drift test warns → rerun `generate_gias_codes.py` → commit regenerated files); (4) no frontend/e2e changes. End with the standard generation footer.
- [ ] **Step 4: Watch CI**
All PR checks must pass. Do not merge — merging triggers the deploy window; the human runs the runbook.
@@ -0,0 +1,591 @@
# Compare-Screen Data Foundation (Pipeline PR) Implementation Plan
> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking.
**Goal:** Land every pipeline/dbt change the compare-screen redesign needs (spec §5 + §8 of `docs/superpowers/specs/2026-07-11-compare-screen-redesign-design.md`): promote raw-but-unstored fields to marts, close the national-averages gaps, and wire the Ofsted report-card columns.
**Architecture:** Meltano Singer taps load `raw.*` tables; dbt builds `staging``marts` (read-only for the backend). All changes here are additive columns/rows — no breaking changes to existing marts. The full `dbt build` runs on the server via the Airflow DAGs; locally we gate with `dbt parse` (no DB needed) plus network-only diagnostic scripts.
**Tech Stack:** Python (Singer SDK taps), dbt-postgres ~1.10 (invoked as `python -m dbt.cli.main`), Meltano, PostgreSQL.
## Global Constraints
- **No new external sources** (spec §5): only fields already in the `raw` schema or in files the taps already download. The one sanctioned tap change is the Ofsted MI report-card columns (spec §5, §8.4) and the legacy-KS2 year addition (same DfE performance-tables source).
- **Additive only:** never rename or drop existing mart columns; the backend maps them 1:1 in `backend/models.py`.
- **Never push to `main`.** Branch: `feat/compare-data-foundation`; PR checks must pass.
- Backend `models.py` changes belong to the follow-up backend PR, not this one.
- dbt invocation is always `python -m dbt.cli.main` (a bare `dbt` resolves to the wrong binary — see `pipeline/dags/school_data_pipeline.py:27`).
- EES suppression codes `z`/`c`/`x` must go through the `safe_numeric` macro.
- Computed benchmarks (FSM/EAL/SEN medians, disadvantaged national average) are **backend work** (spec §5) — explicitly out of scope here.
---
### Task 0: Create the branch
**Files:** none
- [ ] **Step 1:** `git checkout main && git pull && git checkout -b feat/compare-data-foundation`
---
### Task 1: Diagnostics — pin the three unknowns
The spec flags three facts we must confirm from the actual files before wiring code: (a) why `gps_expected_pct`/`science_expected_pct` are NULL in `marts.fact_ks2_national_averages` despite being mapped end-to-end; (b) what the KS2 attainment long file calls its subjects/years for 2021/22 and 2022/23 (subject-level 2022/23 is NULL in prod; school-level 2021/22 is absent); (c) the exact report-card column headers in the current Ofsted MI CSV.
**Files:**
- Create: `pipeline/scripts/diagnose_compare_gaps.py`
**Interfaces:**
- Produces: a printed findings report; Tasks 5, 6, 7 consume the confirmed column/label names. Precedent: `pipeline/scripts/diagnose_ees_ks4.py`.
- [ ] **Step 1: Write the diagnostic script**
```python
"""Diagnose the three data gaps blocking the compare-screen redesign.
Run from repo root (network access required, no DB needed):
python pipeline/scripts/diagnose_compare_gaps.py
"""
import io
import re
import sys
import zipfile
import pandas as pd
import requests
sys.path.insert(0, "pipeline/plugins/extractors/tap-uk-ees")
sys.path.insert(0, "pipeline/plugins/extractors/tap-uk-ofsted")
from tap_uk_ees.tap import ( # noqa: E402
_KS2_NATIONAL_COL_MAP,
_KS2_NATIONAL_CSV_URL,
download_release_zip,
get_all_releases,
)
from tap_uk_ofsted.tap import discover_csv_url # noqa: E402
TIMEOUT = 120
def check_national_gps_science():
print("\n=== (a) National catalogue CSV: GPS/science columns ===")
resp = requests.get(_KS2_NATIONAL_CSV_URL, timeout=TIMEOUT)
resp.raise_for_status()
df = pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False)
df.columns = [c.strip().lower() for c in df.columns]
for csv_col in ("pt_gps_exp", "pt_scita_exp", "avg_readscore", "avg_matscore", "avg_gpsscore"):
status = "PRESENT" if csv_col in df.columns else "MISSING"
print(f" {csv_col}: {status}")
gps_like = [c for c in df.columns if "gps" in c or "scita" in c or "sci" in c]
print(f" all gps/science-ish columns: {gps_like}")
nat = df[df.get("geographic_level", "").str.strip().str.lower() == "national"]
print(f" national rows time_periods: {sorted(nat['time_period'].unique())}")
# Sample the values our map would read for the latest year
latest = nat[nat["time_period"] == nat["time_period"].max()]
for csv_col, field in _KS2_NATIONAL_COL_MAP.items():
val = latest.iloc[0].get(csv_col, "<col missing>") if len(latest) else "<no row>"
print(f" {field} <- {csv_col} = {val!r}")
def check_ks2_attainment_years_subjects():
print("\n=== (b) EES KS2 attainment: years & subject labels ===")
releases = get_all_releases("key-stage-2-attainment")
print(f" releases found: {[r['time_period'] for r in releases]}")
for release in releases:
zf = download_release_zip(release["id"])
name = next((n for n in zf.namelist()
if "ks2_school_attainment_data" in n and n.endswith(".csv")), None)
if not name:
print(f" {release['time_period']}: NO school attainment CSV in ZIP")
continue
with zf.open(name) as f:
df = pd.read_csv(f, dtype=str, keep_default_na=False, nrows=200000)
years = sorted(df["time_period"].unique())
subjects = sorted(df["subject"].unique())
print(f" release {release['time_period']}: time_periods={years}")
print(f" subjects={subjects}")
def check_ofsted_report_card_columns():
print("\n=== (c) Ofsted MI CSV: report-card columns ===")
url = discover_csv_url()
print(f" MI file: {url}")
resp = requests.get(url, timeout=TIMEOUT)
resp.raise_for_status()
df = pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False, nrows=5)
rc_like = [c for c in df.columns
if re.search(r"report card|inclusion|curriculum|achievement|safeguard|well.?being|governance", c, re.I)]
print(f" candidate report-card columns ({len(rc_like)}):")
for c in rc_like:
print(f" - {c!r}")
if __name__ == "__main__":
check_national_gps_science()
check_ks2_attainment_years_subjects()
check_ofsted_report_card_columns()
```
Note: if `_KS2_NATIONAL_CSV_URL` is named differently in `tap_uk_ees/tap.py` (it is defined near the `_KS2_NATIONAL_COL_MAP` around line ~490), import whatever constant holds the catalogue CSV URL.
- [ ] **Step 2: Run it and record findings**
Run: `python pipeline/scripts/diagnose_compare_gaps.py 2>&1 | tee /tmp/compare-gaps-findings.txt`
Expected: three sections printed. Paste the findings as a comment block at the bottom of the script (so they're committed evidence), e.g. `# FINDINGS 2026-07-12: pt_gps_exp MISSING (actual col: ...), 202122 present in release X, rc columns: [...]`.
- [ ] **Step 3: Commit**
```bash
git add pipeline/scripts/diagnose_compare_gaps.py
git commit -m "chore(pipeline): diagnostic for compare-screen data gaps"
```
---
### Task 2: Admissions preference detail → mart
Staging already extracts `second_preference_offers`, `third_preference_offers`, `total_offers` (`stg_ees_admissions.sql:26-29`) — the mart drops them. The cross-LA fields are declared in the tap (`all_applications_from_another_LA`, `offers_to_applicants_from_another_LA`) but not selected in staging.
**Files:**
- Modify: `pipeline/transform/models/staging/stg_ees_admissions.sql` (after line 33, in `renamed`)
- Modify: `pipeline/transform/models/marts/fact_admissions.sql`
- Modify: `pipeline/transform/models/marts/_marts_schema.yml` (fact_admissions block, ~line 120)
**Interfaces:**
- Produces mart columns: `total_offers int`, `second_preference_offers int`, `third_preference_offers int`, `cross_la_applications int`, `cross_la_offers int`. The backend PR will map these in `FactAdmissions`.
- [ ] **Step 1: Add cross-LA columns to staging**
In `stg_ees_admissions.sql`, after the `first_preference_applications` line (line 33):
```sql
-- Cross-borough demand: applications naming this school from families
-- living in another local authority, and offers made to them.
{{ safe_numeric('"all_applications_from_another_LA"') }}::integer as cross_la_applications,
{{ safe_numeric('"offers_to_applicants_from_another_LA"') }}::integer as cross_la_offers,
```
(Quote the identifiers — the tap emits them with mixed case, same trap as `FSM_eligible_percent`, see the header comment in that file. If `dbt parse` or the DAG run later shows the raw columns are lower-cased in Postgres, drop the double quotes.)
- [ ] **Step 2: Pass everything through the mart**
Replace the full select list in `fact_admissions.sql`:
```sql
-- Mart: School admissions — one row per URN per year
select
urn,
year,
school_phase,
places_offered,
total_offers,
total_applications,
first_preference_applications,
first_preference_offers,
second_preference_offers,
third_preference_offers,
cross_la_applications,
cross_la_offers,
first_preference_offer_pct,
oversubscription_ratio,
oversubscribed,
admissions_policy
from {{ ref('stg_ees_admissions') }}
```
- [ ] **Step 3: Add schema tests**
In `_marts_schema.yml` under `fact_admissions.columns`, append:
```yaml
- name: second_preference_offers
- name: third_preference_offers
- name: cross_la_applications
- name: cross_la_offers
- name: total_offers
```
- [ ] **Step 4: Parse gate**
Run: `cd pipeline/transform && python -m dbt.cli.main parse --profiles-dir .`
Expected: `Done.` with no compilation errors.
- [ ] **Step 5: Commit**
```bash
git add pipeline/transform/models/staging/stg_ees_admissions.sql pipeline/transform/models/marts/fact_admissions.sql pipeline/transform/models/marts/_marts_schema.yml
git commit -m "feat(pipeline): admissions preference breakdown and cross-LA demand in marts"
```
---
### Task 3: KS2 progress confidence intervals + writing working-towards
The tap already emits `progress_measure_lower_conf_interval`, `progress_measure_upper_conf_interval`, `working_towards_expected_standard_pupil_percent` (tap.py:203-206). The staging pivot drops them. These power the CI-based Above/Average/Below progress chips (spec §8, first-review item on statistical honesty).
**Files:**
- Modify: `pipeline/transform/models/staging/stg_ees_ks2.sql` (inside the `pivoted` CTE, next to each subject's `progress_measure_score` case, lines ~41/55/72, and in the final select ~lines 145-152)
- Modify: `pipeline/transform/models/marts/fact_ks2_performance.sql`
- Modify: `pipeline/transform/models/marts/_marts_schema.yml` (fact_ks2_performance block, ~line 82)
**Interfaces:**
- Produces mart columns: `reading_progress_lower_ci`, `reading_progress_upper_ci`, `writing_progress_lower_ci`, `writing_progress_upper_ci`, `maths_progress_lower_ci`, `maths_progress_upper_ci` (float), `writing_working_towards_pct` (float).
- [ ] **Step 1: Add pivot cases in staging**
After the `reading_progress` case (line ~41), add:
```sql
max(case when subject = 'Reading'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as reading_progress_lower_ci,
max(case when subject = 'Reading'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as reading_progress_upper_ci,
```
After the `writing_progress` case (line ~55), add:
```sql
max(case when subject = 'Writing'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as writing_progress_lower_ci,
max(case when subject = 'Writing'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as writing_progress_upper_ci,
max(case when subject = 'Writing'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('working_towards_expected_standard_pupil_percent') }} end) as writing_working_towards_pct,
```
After the `maths_progress` case (line ~72), add:
```sql
max(case when subject = 'Maths'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as maths_progress_lower_ci,
max(case when subject = 'Maths'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as maths_progress_upper_ci,
```
Then add the seven new columns to the model's final select (next to the existing `p.reading_progress` / `p.writing_progress` / `p.maths_progress` lines ~145-152):
```sql
p.reading_progress_lower_ci,
p.reading_progress_upper_ci,
p.writing_progress_lower_ci,
p.writing_progress_upper_ci,
p.writing_working_towards_pct,
p.maths_progress_lower_ci,
p.maths_progress_upper_ci,
```
- [ ] **Step 2: Pass through the mart**
In `fact_ks2_performance.sql`, add the same seven column names to the select list immediately after the existing `maths_progress` line (this mart selects staging columns by name; match the file's existing alias style — if columns are selected bare, add them bare).
- [ ] **Step 3: Schema tests**
In `_marts_schema.yml` under `fact_ks2_performance.columns`, append the seven names (no tests beyond presence — values are legitimately NULL for 2023/24+ since progress measures ended with 2022/23, spec §4.3):
```yaml
- name: reading_progress_lower_ci
- name: reading_progress_upper_ci
- name: writing_progress_lower_ci
- name: writing_progress_upper_ci
- name: writing_working_towards_pct
- name: maths_progress_lower_ci
- name: maths_progress_upper_ci
```
- [ ] **Step 4: Parse gate**
Run: `cd pipeline/transform && python -m dbt.cli.main parse --profiles-dir .`
Expected: `Done.`
- [ ] **Step 5: Commit**
```bash
git add pipeline/transform/models/staging/stg_ees_ks2.sql pipeline/transform/models/marts/fact_ks2_performance.sql pipeline/transform/models/marts/_marts_schema.yml
git commit -m "feat(pipeline): KS2 progress confidence intervals and writing working-towards"
```
---
### Task 4: KS4 — Progress 8 banding and disadvantage gaps
The tap's `ees_ks4_info` stream already declares `progress8_banding` (DfE's own "well above average … well below average" label — the ready-made secondary chip), `attainment8_diffn` and `progress8_diffn` (tap.py:338-340). Wire them through staging into the mart.
**Files:**
- Modify: `pipeline/transform/models/staging/stg_ees_ks4.sql` (the CTE that reads `ees_ks4_info` — the same one that already surfaces `sen_pct`; add three columns to its select and to the final joined select)
- Modify: `pipeline/transform/models/marts/fact_ks4_performance.sql` (add after `progress_8_upper_ci`)
- Modify: `pipeline/transform/models/marts/_marts_schema.yml` (fact_ks4_performance block, ~line 93)
**Interfaces:**
- Produces mart columns: `progress_8_banding text`, `attainment_8_disadvantage_gap float`, `progress_8_disadvantage_gap float`.
- [ ] **Step 1: Staging — select from the info source**
In the info CTE of `stg_ees_ks4.sql` add:
```sql
nullif(trim(progress8_banding), '') as progress_8_banding,
{{ safe_numeric('attainment8_diffn') }} as attainment_8_disadvantage_gap,
{{ safe_numeric('progress8_diffn') }} as progress_8_disadvantage_gap,
```
and add the three names to the model's final select (aliased the same way the CTE's other columns are).
- [ ] **Step 2: Mart passthrough**
In `fact_ks4_performance.sql`, after the `progress_8_upper_ci,` line:
```sql
progress_8_banding,
attainment_8_disadvantage_gap,
progress_8_disadvantage_gap,
```
- [ ] **Step 3: Schema tests** — append the three names under `fact_ks4_performance.columns`, plus an accepted-values guard that tolerates NULL:
```yaml
- name: progress_8_banding
tests:
- accepted_values:
values: ['Well above average', 'Above average', 'Average', 'Below average', 'Well below average']
config:
where: "progress_8_banding is not null"
- name: attainment_8_disadvantage_gap
- name: progress_8_disadvantage_gap
```
(If the DAG run later shows different capitalisation in the data, fix the accepted values to match the data, not vice versa.)
- [ ] **Step 4: Parse gate**`cd pipeline/transform && python -m dbt.cli.main parse --profiles-dir .``Done.`
- [ ] **Step 5: Commit**
```bash
git add pipeline/transform/models/staging/stg_ees_ks4.sql pipeline/transform/models/marts/fact_ks4_performance.sql pipeline/transform/models/marts/_marts_schema.yml
git commit -m "feat(pipeline): Progress 8 banding and KS4 disadvantage gaps in marts"
```
---
### Task 5: National averages — 2015/16 row and GPS/science/scaled-score fix
Two changes. (1) `stg_ees_ks2_national.sql:34` filters `>= 201617`, which is exactly why the England line starts a year late (2015/16 RWM = 53% exists in the catalogue). (2) GPS/science expected are NULL in prod despite full end-to-end mapping — Task 1's findings say whether the catalogue CSV column names differ from `_KS2_NATIONAL_COL_MAP` (`pt_gps_exp`, `pt_scita_exp`) or whether values are suppressed at source.
**Files:**
- Modify: `pipeline/transform/models/staging/stg_ees_ks2_national.sql:34`
- Modify (conditional on Task 1 findings): `pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py` (`_KS2_NATIONAL_COL_MAP`)
**Interfaces:**
- Produces: a 201516 row in `marts.fact_ks2_national_averages`; non-NULL `gps_expected_pct`, `science_expected_pct`, `reading_avg_score`, `maths_avg_score`, `gps_avg_score` for years the DfE publishes them. Backend/frontend consume via `/api/national-averages` unchanged (additive year + newly non-NULL fields).
- [ ] **Step 1: Widen the year filter**
In `stg_ees_ks2_national.sql`, change line 34:
```sql
and cast(trim(time_period) as integer) >= 201516
```
(2015/16 was the first year of the current expected-standard tests; nothing earlier is comparable, so keep a floor.)
- [ ] **Step 2: Fix the column map per Task 1 findings**
If Task 1 reported the actual CSV column names for GPS/science/scaled scores differ, update `_KS2_NATIONAL_COL_MAP` in `tap.py` accordingly, e.g. (illustrative — use the diagnosed names):
```python
_KS2_NATIONAL_COL_MAP = {
# ... existing entries ...
"pt_gps_exp": "gps_expected_pct", # replace key with diagnosed name
"pt_scita_exp": "science_expected_pct", # replace key with diagnosed name
}
```
If Task 1 showed the columns are present but suppressed (`x`) at national level for all years, instead delete the two entries from the map, delete the corresponding lines from `stg_ees_ks2_national.sql` and `fact_ks2_national_averages.sql`, and record in the PR description that GPS/science England ticks stay "not in dataset" (the mockups already carry that caveat).
- [ ] **Step 3: Parse gate**`cd pipeline/transform && python -m dbt.cli.main parse --profiles-dir .``Done.`
- [ ] **Step 4: Commit**
```bash
git add pipeline/transform/models/staging/stg_ees_ks2_national.sql pipeline/plugins/extractors/tap-uk-ees/tap_uk_ees/tap.py
git commit -m "fix(pipeline): include 2015/16 national averages; fix GPS/science national mapping"
```
---
### Task 6: Legacy KS2 — load the 2021/22 school-level year
School-level 2021/22 exists in DfE performance-tables archives (same source as the four legacy years already loaded) but in neither our legacy config (stops at 201819, `pipeline/meltano.yml:33-37`) nor EES (starts 2022/23) — unless Task 1's finding (b) showed an EES release carrying 202122, in which case skip this task and note why in the PR.
The legacy URLs point at the self-hosted filebrowser (`10.0.1.224:8081`) — **the 2021/22 DfE archive must be uploaded there first; this is the one human dependency in this plan.**
**Files:**
- Modify: `pipeline/meltano.yml` (legacy_ks2_urls block, line ~33)
**Interfaces:**
- Produces: `raw.legacy_ks2` rows with `year = '202122'`, flowing through `stg_legacy_ks2``fact_ks2_performance` unchanged (the stream maps old column names already; 2021/22 CSVs use the same `PTRWM_EXP`-style headers as 2018/19).
- [ ] **Step 1: Verify the 2021/22 CSV headers match `_LEGACY_KS2_COLUMN_MAP`**
Download the DfE 2021/22 KS2 revised archive (gov.uk "Compare School Performance data download": 2021-2022 all-schools ZIP), then:
Run: `python -c "import zipfile,io,pandas as pd; zf=zipfile.ZipFile('/path/to/2021-2022.zip'); n=[x for x in zf.namelist() if 'ks2final' in x.lower() and x.endswith('.csv')][0]; df=pd.read_csv(zf.open(n), dtype=str, nrows=5); import sys; sys.path.insert(0,'pipeline/plugins/extractors/tap-uk-ees'); from tap_uk_ees.tap import _LEGACY_KS2_COLUMN_MAP as m; missing=[c for c in m if c not in df.columns]; print('missing legacy columns:', missing)"`
Expected: `missing legacy columns: []` (progress columns `READPROG` etc. may legitimately be missing/blank in 2021/22 — acceptable, they load as NULL).
- [ ] **Step 2: Upload the archive to the filebrowser and add the config entry**
In `pipeline/meltano.yml` under `legacy_ks2_urls`, add (with the real share URL from the filebrowser upload):
```yaml
"202122": "http://10.0.1.224:8081/filebrowser/api/public/dl/<SHARE_ID>?inline=true"
```
- [ ] **Step 3: Commit**
```bash
git add pipeline/meltano.yml
git commit -m "feat(pipeline): load 2021/22 school-level KS2 from legacy performance tables"
```
- [ ] **Step 4 (only if Task 1(b) showed 2022/23 subject labels differ):** widen the subject matchers in `stg_ees_ks2.sql` the same way GPS already is (`subject ilike '%grammar%' or subject = 'GPS'`), e.g. `subject in ('Reading', 'reading')` → use the diagnosed labels. Parse-gate and commit as `fix(pipeline): match 2022/23 KS2 subject labels`.
---
### Task 7: Ofsted report-card columns (rc_*)
Resolves the tap TODO (`stg_ofsted_inspections.sql:37`). The marts/backed columns already exist as stubs; this wires real values. Uses Task 1(c)'s confirmed MI column names — the candidates below follow the MI file's existing naming style and must be corrected against the diagnostic output.
**Files:**
- Modify: `pipeline/plugins/extractors/tap-uk-ofsted/tap_uk_ofsted/tap.py` (COLUMN_PRIORITY ~line 19-72, schema ~line 100-114)
- Create: `pipeline/transform/macros/parse_report_card_grade.sql`
- Modify: `pipeline/transform/models/staging/stg_ofsted_inspections.sql:36-46`
**Interfaces:**
- Produces mart columns (already declared in `fact_ofsted_inspection`): `rc_safeguarding_met boolean`, and `rc_inclusion``rc_sixth_form` as integers on the 5-point scale `1=Exceptional, 2=Strong standard, 3=Expected standard, 4=Needs attention/Attention needed, 5=Urgent improvement`. The backend translates codes to labels (same pattern as `gias_codes.py`), verifying wording against Ofsted's published toolkit (spec §8.4).
- [ ] **Step 1: Add tap column mappings**
In `COLUMN_PRIORITY` add (replace candidate strings with Task 1(c)'s exact headers — keep them as priority lists so older files degrade to blank):
```python
"rc_safeguarding_met": ["Report card safeguarding", "Safeguarding"],
"rc_inclusion": ["Report card inclusion", "Inclusion"],
"rc_curriculum_teaching": ["Report card curriculum and teaching", "Curriculum and teaching"],
"rc_achievement": ["Report card achievement", "Achievement"],
"rc_attendance_behaviour": ["Report card attendance and behaviour", "Attendance and behaviour"],
"rc_personal_development": ["Report card personal development and well-being", "Personal development and well-being"],
"rc_leadership_governance": ["Report card leadership and governance", "Leadership and governance"],
"rc_early_years": ["Report card early years", "Early years"],
"rc_sixth_form": ["Report card sixth form", "Sixth form"],
```
And in the stream schema (next to `report_url`, ~line 114):
```python
th.Property("rc_safeguarding_met", th.StringType),
th.Property("rc_inclusion", th.StringType),
th.Property("rc_curriculum_teaching", th.StringType),
th.Property("rc_achievement", th.StringType),
th.Property("rc_attendance_behaviour", th.StringType),
th.Property("rc_personal_development", th.StringType),
th.Property("rc_leadership_governance", th.StringType),
th.Property("rc_early_years", th.StringType),
th.Property("rc_sixth_form", th.StringType),
```
- [ ] **Step 2: Write the grade-parsing macro**
`pipeline/transform/macros/parse_report_card_grade.sql`:
```sql
{% macro parse_report_card_grade(column_name) %}
case lower(trim(nullif({{ column_name }}, 'NULL')))
when 'exceptional' then 1
when 'strong standard' then 2
when 'expected standard' then 3
when 'needs attention' then 4
when 'attention needed' then 4
when 'urgent improvement' then 5
end
{% endmacro %}
```
- [ ] **Step 3: Wire staging**
Replace `stg_ofsted_inspections.sql` lines 36-46 (the NULL stubs) with:
```sql
-- Report Card fields (post-Nov 2025 framework), 5-point scale:
-- 1 Exceptional · 2 Strong standard · 3 Expected standard
-- · 4 Needs attention · 5 Urgent improvement
(lower(trim(nullif(rc_safeguarding_met, 'NULL'))) = 'met') as rc_safeguarding_met,
{{ parse_report_card_grade('rc_inclusion') }}::integer as rc_inclusion,
{{ parse_report_card_grade('rc_curriculum_teaching') }}::integer as rc_curriculum_teaching,
{{ parse_report_card_grade('rc_achievement') }}::integer as rc_achievement,
{{ parse_report_card_grade('rc_attendance_behaviour') }}::integer as rc_attendance_behaviour,
{{ parse_report_card_grade('rc_personal_development') }}::integer as rc_personal_development,
{{ parse_report_card_grade('rc_leadership_governance') }}::integer as rc_leadership_governance,
{{ parse_report_card_grade('rc_early_years') }}::integer as rc_early_years,
{{ parse_report_card_grade('rc_sixth_form') }}::integer as rc_sixth_form,
```
Note `rc_safeguarding_met` becomes boolean (NULL when blank) — matching `fact_ofsted_inspection`'s `rc_safeguarding_met` Boolean column. If `fact_ofsted_inspection.sql` casts these columns, align its casts too (inspect that model; it currently passes the text stubs through).
- [ ] **Step 4: Parse gate + tap smoke test**
Run: `cd pipeline/transform && python -m dbt.cli.main parse --profiles-dir .``Done.`
Run: `python -c "import sys; sys.path.insert(0,'pipeline/plugins/extractors/tap-uk-ofsted'); from tap_uk_ofsted.tap import COLUMN_PRIORITY; assert 'rc_inclusion' in COLUMN_PRIORITY; print('ok')"``ok`
- [ ] **Step 5: Commit**
```bash
git add pipeline/plugins/extractors/tap-uk-ofsted/tap_uk_ofsted/tap.py pipeline/transform/macros/parse_report_card_grade.sql pipeline/transform/models/staging/stg_ofsted_inspections.sql
git commit -m "feat(pipeline): extract Ofsted report-card judgements (rc_* columns)"
```
---
### Task 8: PR + post-merge verification
**Files:** none new
- [ ] **Step 1: Push and open the PR** (Gitea — use the git credential helper + basic-auth API pattern; token-header auth 401s):
```bash
git push -u origin feat/compare-data-foundation
# then create the PR via the Gitea API with basic auth from `git credential fill`
```
PR body: link spec §5/§8, list the new mart columns, note the Task 6 human dependency (filebrowser upload) and the Task 1 findings file.
- [ ] **Step 2: After merge, verify the DAG run picked everything up**
The daily/monthly DAGs rebuild the affected models (`pipeline/dags/school_data_pipeline.py`). Spot-check via the public API (production after promotion, staging first at stx.schoolcompare.co.uk — note external /api is broken at the staging proxy, so check staging from the host):
```bash
# 2015/16 national row exists
curl -sL "https://www.schoolcompare.co.uk/api/national-averages" | python3 -c "import json,sys; d=json.load(sys.stdin); assert any(r['year']==201516 and r['primary'] for r in d['by_year']), '2015/16 missing'; print('201516 ok')"
# 2021/22 school rows exist (Barclay)
curl -sL "https://www.schoolcompare.co.uk/api/schools/138690" | python3 -c "import json,sys; d=json.load(sys.stdin); ys=[r['year'] for r in d['yearly_data']]; assert 202122 in [int(y) for y in ys], ys; print('202122 ok')"
```
(The admissions/CI/KS4/rc_* columns aren't API-visible until the backend PR maps them — verify those directly in Postgres from the pipeline host: `select count(*) from marts.fact_admissions where second_preference_offers is not null;` etc.)
- [ ] **Step 3: Update the spec** — tick off the §5 promotions this PR delivered (edit the spec's promotion list to note "landed in PR #NN") and commit to main via a docs PR or alongside the backend PR.
---
## Out of scope (next plans)
1. **Backend PR:** map new columns in `backend/models.py`, extend `/api/compare` with supplementary blocks + `national_averages`, computed benchmarks (FSM/EAL/SEN/size medians, disadvantaged national average), CI-based progress banding, report-card label translation (verify against Ofsted toolkit), Ofsted provider-page URLs, graded-vs-ungraded surfacing.
2. **Frontend PR:** rebuild `/compare` per the mockups + e2e journeys (promotion gate).
3. **Separate bug fix:** third school's series not rendering on the current production chart.
4. **Post-v1 (spec):** census ethnicity/young-carer promotion, IDACI display, attendance section, gender-split/absence tier-2 measures.
5. **Already in marts, no work needed:** KS4 EBacc entry/APS, grade 5+ English & maths, Progress 8 CIs — `fact_ks4_performance` carries them today; only the backend needs to expose them.
@@ -0,0 +1,182 @@
# GIAS Code Dictionaries — Codes in Marts, Names in Code
**Date:** 2026-07-09
**Status:** Implemented 2026-07-09 — see docs/superpowers/plans/2026-07-09-gias-code-dictionaries.md
## Goal
Six GIAS classification fields are stored in the marts as repeated name
strings. Replace them with the official DfE integer codes and translate
code → name in application code. After this change the marts carry only
codes for:
| GIAS field | Today (marts, string) | After (marts, int) |
|---|---|---|
| `TypeOfEstablishment (name)` | `dim_school.school_type` | `school_type_code` |
| `EstablishmentStatus (name)` | `dim_school.status` | `status_code` |
| `PhaseOfEducation (name)` | `dim_school.phase` | `phase_code` |
| `OfficialSixthForm (name)` | (already reduced to `has_sixth_form` bool) | `official_sixth_form_code` in staging only; mart keeps the bool |
| `ReligiousCharacter (name)` | `dim_school.religious_character` | `religious_character_code` |
| `AdmissionsPolicy (name)` | `dim_school.admissions_policy` | `admissions_policy_code` |
Motivation: smaller marts and stable enum values for filtering. (Honest
sizing note: at ~25k open schools the raw performance win is modest; the
durable benefits are storage, DfE-governed vocabulary, and filter values
that can't drift with GIAS renames.)
## Decisions (made during brainstorming)
1. **GIAS native codes**, not custom enums. The GIAS bulk CSV publishes an
official `X (code)` column beside every `X (name)` column. We ingest the
DfE's own codes; no invented mapping to maintain.
2. **Translation lives in the backend at the API boundary.** The API keeps
serving today's name strings; the frontend, e2e journeys, and API
consumers are untouched.
## Design
### 1. Tap (Singer schema)
Add the six `(code)` columns to `GIASEstablishmentsStream.schema` in
`pipeline/plugins/extractors/tap-uk-gias/tap_uk_gias/tap.py`:
```
"TypeOfEstablishment (code)", "EstablishmentStatus (code)",
"PhaseOfEducation (code)", "OfficialSixthForm (code)",
"ReligiousCharacter (code)", "AdmissionsPolicy (code)"
```
The `(name)` columns **stay declared** — raw keeps both so we can detect
dictionary drift (§4) and regenerate dictionaries from live data.
### 2. Staging (`stg_gias_establishments.sql`)
- Add int casts: `school_type_code`, `status_code`, `phase_code`,
`official_sixth_form_code`, `religious_character_code`,
`admissions_policy_code` (all `cast(nullif(trim(...), '') as integer)`).
- Remove the corresponding name columns from the staging select
(`school_type`, `status`, `phase`, `official_sixth_form`,
`religious_character`, `admissions_policy`). Names live only in raw.
### 3. Marts
**`dim_school`** stores codes only:
- `school_type_code`, `status_code`, `phase_code`,
`religious_character_code`, `admissions_policy_code` replace their
string columns.
- Status filter becomes `where status_code in (<open>, <proposed-to-close>)`.
The numeric values are read from live raw data at implementation time
(`select distinct "EstablishmentStatus (code)", "EstablishmentStatus (name)"`),
never assumed from memory. Same filter in `dim_location`.
- `has_sixth_form` derives from `official_sixth_form_code`
(`<has-code>` → true, `<does-not>/<not-applicable>` → false, null →
`statutory_high_age >= 18` fallback). The `lower(trim(...))` string guard
becomes obsolete and is removed.
- `phase_code` derivation keeps today's cascade but emits codes:
1. GIAS `phase_code` when it is a real value (not the not-applicable code);
2. statutory-age inference emits the matching GIAS code
(Primary / Secondary / All-through — numeric values confirmed from
live data at implementation);
3. school-name heuristics (unchanged — they match `school_name`, which is
not one of the six fields) emit the same codes;
4. else null.
- dbt schema tests: `accepted_values` (severity **warn**) on every code
column, values taken from the dictionary; `not_null` warn on `phase_code`
(mirrors today's phase test); `has_sixth_form` tests unchanged.
**`dim_location`**: only the status filter changes (must stay byte-identical
to `dim_school`'s — the API inner-joins the two).
### 4. Dictionaries
**Canonical module: `backend/gias_codes.py`**
```python
ESTABLISHMENT_STATUS: dict[int, str]
SCHOOL_TYPE: dict[int, str]
PHASE_OF_EDUCATION: dict[int, str]
OFFICIAL_SIXTH_FORM: dict[int, str]
RELIGIOUS_CHARACTER: dict[int, str]
ADMISSIONS_POLICY: dict[int, str]
def translate(code: int | None, mapping: dict[int, str]) -> str | None:
"""None -> None; unknown code -> 'Unknown (<code>)' + warning log."""
```
- Contents are generated from live raw data
(`SELECT DISTINCT code, name FROM raw.gias_establishments ...` per field)
and sanity-checked against the DfE GIAS registers. Names must be
byte-identical to what the API serves today.
- Unknown codes never blank the UI: `translate` returns `"Unknown (<code>)"`
and logs, so a new DfE value degrades gracefully.
**Pipeline copy: `pipeline/scripts/gias_codes.py`**
The app and pipeline Docker images have disjoint build contexts
(`Dockerfile` copies `backend/`; `pipeline/Dockerfile` copies `pipeline/`),
so the Typesense sync cannot import the backend module. It gets a
byte-identical copy, and a backend unit test asserts
`backend/gias_codes.py` and `pipeline/scripts/gias_codes.py` have identical
content — drift fails CI. (Deliberately chosen over codegen: six dicts do
not justify build machinery.)
**Seed for drift detection: `pipeline/transform/seeds/gias_code_names.csv`**
Columns `field,code,name` mirroring the dictionary. A dbt test (severity
warn) compares live raw `(code, name)` pairs against the seed; when DfE adds
or renames a value the nightly run warns, prompting a dictionary + seed
update in one PR.
### 5. Backend translation (API contract unchanged)
- `_MAIN_QUERY` selects the code columns instead of the name columns.
- `load_school_data_as_dataframe()` translates immediately after
`pd.read_sql`, writing today's column names:
```python
df["phase"] = df["phase_code"].map(...)
df["school_type"] = df["school_type_code"].map(...) # then normalize_school_type as today
df["status"] = df["status_code"].map(...)
df["religious_denomination"] = df["religious_character_code"].map(...)
df["admissions_policy"] = df["admissions_policy_code"].map(...)
```
Everything downstream — `PHASE_GROUPS`, filters, payload builders,
`/api/filters`, frontend, e2e — sees exactly today's strings. No frontend
changes.
- `backend/models.py` `DimSchool`: string columns replaced by
`*_code = Column(Integer)`.
### 6. Typesense sync
`pipeline/scripts/sync_typesense.py` selects `phase`, `school_type`,
`religious_character` today. It switches to the code columns and translates
via `pipeline/scripts/gias_codes.py` before indexing, so facet values in
search are unchanged.
### 7. Rollout
- No DB migration: marts are full-rebuild tables.
- Deploy window: until the first post-merge pipeline run, the old marts
still carry string columns while the new backend queries code columns, so
the backend's query fails and it serves empty data (the one-column retry
built for `has_sixth_form` doesn't generalise to six columns, and a full
old-schema fallback query isn't worth it). **Decision: accept the window
and close it operationally — the runbook is merge → deploy → trigger
`school_data_daily` immediately.** The DAG's final step already calls
`/api/admin/reload`, so the backend recovers without a restart.
- Tests: backend unit tests for `translate()` (known / unknown / None),
payload tests asserting names still served, the file-parity test, dbt
schema/seed tests. Frontend: no changes; existing Jest suite is the
regression net.
## Out of scope
- Recoding other string columns (`gender`, `urban_rural`,
`nursery_provision`, `local_authority_name` …) — same pattern can follow
later if this proves out.
- Collapsing academy subtypes (today's `normalize_school_type`) — kept
as-is, applied after translation.
- Serving codes through the API — the contract deliberately keeps names.
@@ -0,0 +1,177 @@
# Compare Screen Redesign — Expert Data Review
**Date:** 2026-07-11
**Reviewer:** subagent briefed as an English education-standards / DfE-Ofsted data expert
**Subject:** desktop + mobile compare mockups and the redesign spec
(`2026-07-11-compare-screen-redesign-design.md`)
**Status:** first-pass must-fixes applied 2026-07-12; second-pass
findings (below) applied 2026-07-12 — mockups + spec §4/§8 updated
## Must-fix
1. **COVID gap is wrong and drops a real results year.** KS2 tests were
cancelled 2019/20 and 2020/21 only; they resumed in 2021/22 with
published school-level results (England RWM ≈ 59%). The mockup charts
omit 2021/22 entirely and the tooltip claims no tests were held
2019/202021/22. Fix: add 2021/22 to axis and all series; shrink the
gap band; optionally annotate 2021/22 with DfE's post-pandemic
comparability caution.
2. **Report-card at-a-glance summary miscounts areas.** Detail list has
4 Strong / 2 Expected / 1 Attention needed + Safeguarding met, but
the summary says "3 areas Expected standard" — it counts safeguarding
as a graded area. Safeguarding is a separate binary judgement and
must be excluded from rating counts.
3. **"Where the offers went" derivation is unsound.** Places 1st-pref
offers ≠ "second or third choices": the residual can include 4th6th
preference offers (pan-London scheme) and LA-allocated children who
didn't choose the school; and offers don't necessarily equal PAN.
Use the real 2nd/3rd-preference fields being promoted from
`raw.ees_admissions`; until then drop the row.
4. **Ofsted timeline in the copy is wrong.** Overall grades were
abolished September 2024, not November 2025; Sept 2024Nov 2025
inspections kept the four key judgements without an overall grade
(ungraded inspections carried grades forward). Neither mockup shows
the interim regime, which will dominate real comparisons. Fix copy
and add an interim example.
5. **Barclay's "published an overall grade only — no area-by-area
detail" misdescribes inspections.** No inspection type does that; a
2021 graded inspection necessarily had subgrades — the gap is in our
dataset. If it was an ungraded (s8) inspection, "Outstanding" is a
carried-forward grade and should say so. Fix: "We don't hold
area-by-area detail for this inspection", and distinguish graded vs
ungraded in the data model.
## Should-fix
6. Writing is teacher assessment, not a test — "national tests and
teacher assessments"; note TA caveat on the Writing strip.
7. Verify renewed-framework wording against Ofsted's final toolkit:
likely "Needs attention" (not "Attention needed") and "Personal
development and well-being" (which otherwise collides with the
identically-named legacy judgement). Pin every label to the
published toolkit.
8. "Expected standard" now means two things on one page (Ofsted area
rating vs KS2 measure) — disambiguate in tooltips.
9. Disadvantaged row: DfE definition includes looked-after / previously
looked-after children, not just FSM6; benchmark labels inconsistent
across desktop/mobile; subgroup percentages need cohort sizes or a
volatility threshold before chips are attached.
10. "Trend, last 7 years" spans ten years; sparklines render the COVID
gap as equal spacing (the exact defect the audit criticises) and
"Improved: 52% → 87%" endpoint-cherry-picks a volatile series.
11. At-a-glance "Getting a place" uses different metrics per school
(Barclay is also oversubscribed on total preferences but shows a
green chip). Standardise on first-preference success %. Explain the
equal-preference rule; condition "living close by matters" on the
school's actual oversubscription criteria.
12. "457 applications for 180 places" = total preferences at any rank,
not head-to-head applicants; lead with first preferences vs places.
Add offers-vs-final-intake (waiting lists/appeals) caveat.
13. Elmhurst's subgrade list is likely missing Early years provision
(school has a nursery) — possible pipeline gap.
14. "Ofsted rating" label is obsolete post-Sept-2024 — use "Latest
Ofsted inspection"; check whether Oct 2021 is the latest inspection
or merely the latest graded one.
15. SEN: "EHCP plans" is redundant; 28% SEN support often indicates
resourced provision — add a note; England SEN-support ≈ 14%, not 13%.
## Nice-to-have
16. Consistent labelling of official DfE vs dataset-computed benchmarks
(and medians shouldn't be called averages inconsistently).
17. England 2015/16 RWM (53%) exists in DfE publications — the null is
a dataset gap; source it or the England line looks broken.
18. "1 in 4 first choices missed out" — actually more than 1 in 4.
19. "1,273 of 1,260 places (full)" is over capacity; capacity figures
are often stale — say "at or above capacity".
20. State the actual suppression rule (DfE: ≤5 pupils suppressed,
small numbers rounded) instead of "a handful".
21. Spec §4.3 progress chips can't exist for displayed years: KS2
progress ended with 2022/23 (no KS1 baseline) and returns
~2027/28 with the reception baseline. Make explicit in the spec.
IDACI (spec §4.5) is absent from mockups; if shipped, caveat it
describes pupils' neighbourhoods, not the school.
22. Tooltips should give the official term "first preference" alongside
the plain-English "first choice".
## Overall assessment (verbatim gist)
The bones are genuinely good by education-data standards —
England-average anchoring, explicit non-comparability messaging across
Ofsted regimes, refusal to synthesise an overall grade, time-true
x-axis, neutral FSM/EAL framing — better than most commercial
school-comparison sites. But items 15 are outright factual errors or
misdescriptions that a well-informed parent or Ofsted would catch;
the admissions section needs the most conceptual work (equal
preference, preferences-vs-applicants, offers-vs-intake). Fix 15
before user testing; the rest fold into the planned PRs.
---
# Second-pass review (2026-07-12)
Same reviewer, after the must-fixes and the new three-tier metric
exposure model were applied.
## Verification of first-pass must-fixes
- **1 (COVID/2021/22): resolved.** Time-true axis, band covers only the
cancelled years, England 58.7% consistent with official figures,
dataset gaps break lines honestly; reading/maths England series all
match published figures; RWM ≤ min(subject) checks pass.
- **2 (report-card count): resolved** — safeguarding excluded, spec §8.2.
- **3 (offers derivation): resolved** — row removed, spec §8.3 bans it.
- **4 (Ofsted timeline): resolved on desktop; mobile omits the interim
regime clause** (see finding 6).
- **5 (Barclay explanation): resolved.**
## New findings
1. **Should-fix — scaled-score strip domain contradicts caption.**
Caption says "scaled scores run 80120", strips render 100120;
truncated domain exaggerates small gaps and below-100 averages
would fall off the edge. Render 80120, or caption the 100120
window honestly and define below-100 behaviour.
2. **Should-fix — scaled-score England ticks (106/105/105) unsourced.**
Plausible but hand-entered; verify against DfE 2024/25 tables and
add loading official England scaled scores to the pipeline list
(absent from §8.1/§8.6).
3. **Should-fix — "Writing" listed under "Higher standard" in the
picker.** Writing TA outcome is "greater depth" (GDS), never
"higher standard". Label "Writing — greater depth (teacher
assessment)"; tooltip the combined higher-standard composition.
4. Nice — "grammar & punctuation" summary line drops "spelling" (GPS).
5. Nice — science is teacher-assessed (no KS2 test since 2009) and
coarse; tooltip it like writing; reconsider its tier-2 slot.
6. **Should-fix — mobile Ofsted copy skips the interim regime**
(Sept 2024Nov 2025) that desktop explains. One clause fixes it.
7. **Should-fix — benchmark provenance still inconsistent** (EAL
tooltip unsourced; FSM/disadvantaged chips vs tooltips use three
vocabularies; header note says all England averages are official).
Adopt one house style: official = "England average", computed =
"benchmark / typical state school (our dataset)". Also tighten EAL
definition to census wording ("first language known or believed to
be other than English").
8. Nice — "community primaries" distance note attached to an academy
(Elmhurst); say "non-faith primaries" or condition on policy field.
9. Nice — "Improving since 2022" → "since 2022/23".
10. Nice — England chart tooltips show decimals; §7 mandates whole
percents.
## Residual gaps not covered by spec §8
11. Spec promises IDACI-in-words, Attendance section, and tier-2
gender/absence that the mockups never show — mark post-v1 or
demonstrate, so implementation scope is unambiguous.
12. Add official England scaled-score averages to the pipeline task
list.
13. Add the writing/greater-depth terminology rule to §8.7.
## Verdict
All must-fixes genuinely resolved; the tier model is conceptually
sound ("no measure is lost", honest dataset-gap breaks, grouped
picker). Remaining issues are contained: one internal contradiction
(80120 vs 100120), one provenance inconsistency, one terminology
error (writing/GDS). With findings 13 and 67 addressed, the data
framing is fit to put in front of parents.
@@ -0,0 +1,324 @@
# Compare Screen Redesign — Audit & Design
**Date:** 2026-07-11
**Status:** Draft — awaiting review
**Scope:** `/compare` page (nextjs-app), `/api/compare` endpoint (backend)
## 1. Audit of the current screen
The current compare page (`nextjs-app/components/ComparisonView.tsx`) is a
single-metric analyst tool: a `<select>` with ~40 KS2/GCSE metrics, one
line chart over time, and a year-by-year table — all for the one selected
metric. Observed on production with 3 primary schools:
**What works**
- URL-shareable state (`?urns=…&metric=…`), native share sheet.
- Phase tabs (primary/secondary) with sensible auto-detection.
- Colour-coded school cards tied to chart series.
- Metric descriptions from `/api/metrics` (single source of truth).
**What doesn't**
1. **Performance-only.** The database already holds Ofsted inspections,
admissions/oversubscription history, pupil characteristics (FSM/EAL),
SEN, deprivation (IDACI), finance, capacity, faith, gender, trust —
none of it reaches the compare screen. `/api/compare` returns only
`yearly_data` + minimal `school_info`, while `/api/schools/{urn}`
already returns all supplementary blocks.
2. **One metric at a time.** A parent must know which of ~40 metrics
matters, select each in turn, and hold results in their head. There is
no side-by-side overview and no way to see two dimensions at once.
3. **No benchmarks.** Numbers float without anchors: is 79% RWM good?
The DB has official national averages (`fact_ks2_national_averages`)
but the page never shows them.
4. **Domain jargon untranslated.** "GPS Expected %", "Progress scores",
"RWM Combined" assume DfE literacy. The only plain-English help is one
note for progress scores.
5. **Raw numbers, no judgement support.** 87.0% vs 92.0% vs 79.0% — the
page never says "all three are well above the England average of 62%",
which is the fact a parent actually needs.
6. **Bugs/paper cuts observed:** the third school's series did not render
on the production chart despite table data (worth a separate fix);
the COVID gap (2018/19 → 2022/23) renders as equal spacing with no
annotation; table shows "87.0%" precision that implies false accuracy.
## 2. Data inventory (available vs shown)
| Domain | Source table | On detail page | On compare |
|---|---|---|---|
| KS2 attainment/progress | fact_ks2_performance | yes | **yes** (only thing shown) |
| National averages | fact_ks2_national_averages | partial | no |
| Ofsted (latest + subgrades + report-card fields) | fact_ofsted_inspection, dim_school | yes | no |
| Admissions & oversubscription (multi-year) | fact_admissions | yes | no |
| Pupil characteristics (FSM, EAL, gender split) | fact_pupil_characteristics | yes | no |
| Context (SEN, disadvantaged, stability, absence) | fact_ks2_performance | via metric picker | buried in picker |
| Deprivation (IDACI) | fact_deprivation | yes | no |
| Finance (per-pupil spend) | fact_finance | yes | no |
| School facts (capacity, faith, ages, trust, nursery, gender) | dim_school | yes | no |
| Location/distance | dim_location | map | no |
## 3. Design goals
1. **Answer parent questions, in order:** Is it a good school (Ofsted)?
Do children do well there (academics vs England)? Will my child get a
place (admissions)? What is the school like (size, community, faith)?
2. **Every number gets an anchor** — the England average, rendered as a
consistent visual tick, plus a plain-English chip
(Above / Close to / Below England average).
3. **Plain English first, jargon on demand.** Labels are questions or
sentences ("Children reaching the expected standard in reading,
writing and maths"), codes/acronyms live in tooltips.
4. **Scan whole-picture first, drill down second.** The single-metric
trend explorer survives, demoted to an "Explore trends" section at the
bottom rather than being the entire page.
## 4. Proposed structure
Columns = schools (max 4 visible on desktop, horizontal scroll beyond),
rows = dimensions. Sticky compact school header keeps column identity
while scrolling. Sections, in order:
1. **At a glance** — verdict row per school: Ofsted badge, headline
attainment vs England (dot strip + chip), oversubscription chip,
size, distance (when a location is set).
2. **Ofsted inspection** — must handle all three inspection regimes,
which will coexist in comparisons for years:
- **Legacy graded (pre-Sept 2024):** overall grade badge
(Outstanding/Good/Requires improvement/Inadequate). Subgrades,
where published, are rendered in the **same area-by-rating chip
list UX as report cards** (one row per judgement area, rating as
a chip) — one visual grammar for inspection detail across both
regimes. Where our dataset has no subgrades for an inspection,
say so honestly ("We don't hold area-by-area detail for this
inspection") and point to the school's Ofsted page — never claim
the inspection itself published no detail (graded inspections
always have subgrades; if it was ungraded, the grade is
carried forward and must be labelled as such).
- **Interim ungraded (Sept 2024 Nov 2025):** parsed outcome
("remains Good") shown as the effective grade, marked as such.
- **Renewed framework report card (from Nov 2025):** no overall
grade exists. Render the report card as an area-by-rating list
using Ofsted's 5-point scale (Exceptional / Strong standard /
Expected standard / Attention needed / Urgent improvement) across
the evaluation areas we model (`rc_inclusion`,
`rc_curriculum_teaching`, `rc_achievement`,
`rc_attendance_behaviour`, `rc_personal_development`,
`rc_leadership_governance`, `rc_early_years`, `rc_sixth_form`)
plus the separate safeguarding met/not-met flag. **At-a-glance
summary rule:** never an unlabelled colour strip — summarise by
counting areas per rating, best first ("5 areas Strong standard ·
3 areas Expected standard"), and always name any area rated
Attention needed or Urgent improvement explicitly (never fold
problems into a count), plus "Safeguarding not met" whenever that
flag is false. When everything is Expected standard or better,
add the reassurance line "No areas need attention".
When a comparison mixes regimes, show a one-line comparability note
("Ofsted changed how it reports in Nov 2025 — a report card and an
older overall grade aren't directly comparable"). Never derive a
fake overall grade from report-card areas.
3. **Academics (KS2)** — one dot-strip row per headline measure (RWM
expected, RWM higher, reading/writing/maths expected), each with the
England-average tick and per-school dots; copy must say "tests and
teacher assessments" (writing is TA, not a test). Progress scores
translated to Above/Average/Below chips (CI-based) — **but note KS2
progress measures ended with 2022/23** (no KS1 baseline afterwards)
and return only when the reception-baseline cohort reaches Y6
(~2027/28), so progress chips apply to historical years in the
trends explorer, not the headline view. Sparkline per school over
the full published period, with an honest gap for the cancelled
test years (2019/202020/21). Disadvantaged-pupils row under an
"Equity" subheading, always with cohort size shown and DfE's full
definition (FSM6 **or** looked-after/previously looked-after).
4. **Getting a place** — oversubscription ratio as plain sentence
("184 applications for 80 places"), first-preference success %, trend
vs last year, admissions policy.
5. **Who goes there** — pupils on roll (vs capacity), boys/girls, FSM %,
EAL %, SEN support %, faith, ages, nursery, trust. *Post-v1:* IDACI
decile in words (needs a coverage check of `fact_deprivation` and
the neighbourhood-not-school caveat, §8.7).
6. **Attendance***post-v1.* The KS2 test-day absence fields are the
only per-school absence data we hold; they're near-zero for most
schools and easy to misread as general attendance. Ship only if a
general-absence source lands.
7. **Explore trends** (existing feature, collapsed) — metric picker +
multi-year line chart + table, with an added England-average
reference line and a COVID-gap annotation.
**Metric exposure model (three tiers).** No measure from the current
page is lost; they surface at three levels of prominence:
- **Tier 1 — headline strips (always visible):** RWM expected,
reading/writing/maths expected, RWM higher standard.
- **Tier 2 — "More measures" expansion inside Academics:** GPS and
science expected % (science labelled teacher-assessed), average
scaled scores (reading/maths/GPS, same dot-strip grammar showing
the 100120 window of the 80120 scale, widening below 100, with
the England tick) — one tap/click away, same visual language.
*Post-v1:* gender split and absence (see §4.6).
- **Tier 3 — Explore trends:** the full grouped catalogue (the
current page's ~40 metrics, including equity and school-context
measures, and the GCSE set for secondary phase) drives the
year-by-year chart and table via the grouped metric picker.
The tier assignment is a content decision per phase (secondary:
Attainment 8, Progress 8 banding, grade 5+ English & maths as tier 1;
EBacc and subject entries as tier 2).
Finance (per-pupil spend) is deliberately deferred: low parent value,
risk of misreading. Revisit later.
**Mobile (design target — mobile first):** the desktop grid is the
adaptation, not the other way round. On mobile the layout goes
*measure-first*: each row is one measure with all schools listed under
it (colour dot + short name + value + chip), so comparison never
requires horizontal swiping between school cards. A sticky horizontal
school-chip bar keeps identity and add/remove available while
scrolling. Dot strips already read measure-first and carry over
unchanged. The trend chart scrolls horizontally inside its container.
## 5. Data strategy — existing dataset only
Constraint (agreed 2026-07-11): use only data already in marts plus
fields already present in the `raw` schema extracts we pull today.
No new external sources.
**Gaps in the mockup, resolved within this constraint:**
| Mockup element | Resolution |
|---|---|
| England average for disadvantaged pupils | Compute from our own data: `stg_ees_ks2` already pivots the Disadvantaged breakdown per school; aggregate it (weighted by eligible pupils) into `fact_ks2_national_averages` or compute in the API. Label it "England average (state schools)". |
| England context for FSM / EAL / SEN chips | Compute dataset-wide medians per phase, same pattern as `/api/national-averages` does for KS4. |
| "Much larger than average" size label | Dataset median pupils-on-roll per phase. |
| Ofsted link | We don't have deep links to the latest report, so always link to the school's Ofsted provider page, `https://reports.ofsted.gov.uk/provider/21/{urn}`, derived from URN (label it "the school's Ofsted page", not "the report"). |
**Raw fields we already pull but don't store — promote to marts (one
dbt/pipeline PR, no tap changes):**
- `raw.ees_admissions`: 2nd/3rd preference applications and offers,
total-preference counts, cross-LA applications and offers → richer
"Getting a place" (e.g. "offers reached 2nd-choice families",
competition from outside the borough).
- `raw.ees_ks2_attainment`: progress-measure confidence intervals and
"working towards" % → lets the Above/Average/Below progress chips be
statistically honest (band by CI overlap with 0, mirroring DfE
methodology) instead of thresholding the point estimate.
- `raw.ees_ks4_performance` / `ees_ks4_info`: `progress8_banding`
(DfE's own plain-English "well above average … well below average"
label — exactly the chip we want for secondary), EBacc entry/APS,
grade-5+ English & maths, `attainment8_diffn`/`progress8_diffn`
(disadvantage gaps) → the secondary-phase version of the Academics
section.
- `raw.ees_census`: young-carer % and the ethnicity breakdown →
optional "Who goes there" enrichment; hold for a later iteration
(presentation needs care), but the data requires no new extract.
- `raw.ofsted_inspections` / tap-uk-ofsted: the `rc_*` report-card
columns exist in staging/marts but are stubbed `null` — the tap has a
TODO to map the report-card column names from the Ofsted MI file
(same monthly extract we already download; inspections from Nov 2025
onward carry them). This is the one promotion that needs a small tap
schema addition, and it's a prerequisite for the new-framework Ofsted
display above.
Explicitly out (not in any current extract): school-level phonics,
workforce/teacher data, per-school attendance beyond the KS2 test-day
absence fields, Ofsted report-card documents themselves.
## 6. API changes
Extend `GET /api/compare` response per URN with the same supplementary
blocks the detail endpoint already builds (`get_supplementary_data`):
`ofsted`, `census`, `admissions` (+ `admissions_history`), `deprivation`,
plus a top-level `national_averages` block for the latest year. Reuse the
existing function; no new tables. Response stays backward-compatible
(additive fields only). Add derived helper fields server-side or compute
chips client-side from `national_averages` (client-side preferred — no
schema churn).
## 7. Accessibility & comprehension devices
- Verdict chips are text + colour + position (never colour alone).
- Every acronym has a tooltip using existing `MetricTooltip`.
- "How to read this" one-liner at the top of each section.
- Chart palette: coral `#e07256`, teal `#00949b`, purple `#8664c9`
(validated: lightness band, chroma, CVD separation, contrast — the
current `--chart-2/-4` tokens fail chroma/contrast checks and should
be nudged to these).
- Numbers rounded to whole percents; England tick labelled on first use.
## 8. Expert-review requirements
An adversarial review by an education-data expert (full findings in
`2026-07-11-compare-screen-expert-review.md`) was applied to the
mockups on 2026-07-12. The following are binding requirements for
implementation, beyond what the mockups can show:
1. **Chart truthfulness:** KS2 tests were cancelled 2019/202020/21
only. **2021/22 school-level figures are a permanent source gap**
DfE stated it would not publish KS2 2021/22 in performance tables
(verified 2026-07-12 against EES, the CSP download service, and
DfE release notes; see `# TASK 6 VERIFICATION` in
`pipeline/scripts/diagnose_compare_gaps.py`). The chart's England-
only 2021/22 point with broken school lines is therefore the
correct permanent rendering; copy should say "DfE didn't publish
school-level figures for 2021/22", not "not in our dataset yet".
The 2015/16 national figure and the GPS/science/scaled-score
England averages ARE loadable (mapping already correct; refreshed
raw extract backfills them). Never render missing years as if time
were continuous.
2. **Report-card summaries** count graded areas only — safeguarding is
a separate binary flag, never included in rating counts.
3. **Admissions:** use the real preference-breakdown fields from
`raw.ees_admissions`; never derive "lower-preference offers" as
places first-preference offers. Frame total applications as
"named on N forms" (any rank), lead with first-preference success,
and standardise at-a-glance chips on that one metric. Explain the
equal-preference rule; caveat offers vs final intake (waiting
lists/appeals); condition "distance decides" on the school's actual
oversubscription criteria where we have the admissions-policy field.
4. **Ofsted:** overall grades ended September 2024 (report cards from
November 2025); the interim regime must be renderable. Distinguish
graded (s5) vs ungraded (s8) inspections and surface carried-forward
grades as such; "we don't hold the detail" is a statement about our
dataset, never about the inspection. Verify every scale/area label
against Ofsted's final published toolkit before launch (e.g. "Needs
attention" vs "Attention needed"; "Personal development and
well-being" vs the identically-named legacy judgement). Check
whether a school's latest inspection is merely its latest *graded*
one. Confirm Early years provision subgrades flow through the
pipeline for schools with nurseries.
5. **Subgroup honesty:** disadvantaged-pupil percentages carry cohort
sizes and follow the DfE suppression rule (≤5 pupils suppressed);
state the rule verbatim in the footer.
6. **Benchmark provenance:** official DfE figures and
dataset-computed benchmarks must be labelled distinctly and
consistently everywhere (a computed median is a "benchmark",
not an "England average").
7. **Copy details:** "Latest Ofsted inspection" (not "Ofsted rating");
"EHC plans"; SEN-support benchmark ≈14%; high SEN share may
indicate resourced provision (say so neutrally); "at or above
capacity" rather than "full" (capacity data is often stale);
disambiguate Ofsted's "Expected standard" from the KS2 measure;
give official terms ("first preference") alongside plain English.
Writing has no "higher standard" — its TA outcome is "greater
depth (GDS)"; never list writing under a higher-standard group.
Science and writing are teacher-assessed and must be labelled as
such (no KS2 science test since 2009). House style for benchmark
provenance: official DfE figures say "England average"; computed
figures say "state-school average (computed from our dataset)" —
applied to every chip, tooltip, header note and section intro.
EAL uses the census wording: first language known or believed to
be other than English. If IDACI ships, caveat that it describes
pupils' home neighbourhoods, not the school.
## 9. Rollout
1. **PR 1 (backend):** extend `/api/compare` + tests.
2. **PR 2 (frontend):** new compare layout behind the existing route;
e2e journey updated in the same PR (promotion gate).
3. **Fix separately:** missing third series on the current chart.
## 10. Open questions for review
- Max schools: keep 10 in API but cap visible columns at 4 with scroll?
- Should distance-from-home appear when the user searched by postcode
(data exists via `dim_location`)?
- Keep finance out of v1? (Recommended: yes, out.)
@@ -44,3 +44,24 @@ describe('SecondarySchoolRow sixth-form tag', () => {
expect(screen.queryByText('Sixth form')).not.toBeInTheDocument(); expect(screen.queryByText('Sixth form')).not.toBeInTheDocument();
}); });
}); });
describe('SecondarySchoolRow proposed-to-close tag', () => {
it('shows the tag when GIAS status is "Open, but proposed to close"', () => {
render(
<SecondarySchoolRow
school={{ ...base, status: 'Open, but proposed to close' }}
/>,
);
expect(screen.getByText(/Proposed to close/)).toBeInTheDocument();
});
it('hides the tag for a plain open school', () => {
render(<SecondarySchoolRow school={{ ...base, status: 'Open' }} />);
expect(screen.queryByText(/Proposed to close/)).not.toBeInTheDocument();
});
it('hides the tag when status is missing', () => {
render(<SecondarySchoolRow school={base} />);
expect(screen.queryByText(/Proposed to close/)).not.toBeInTheDocument();
});
});
+11
View File
@@ -212,3 +212,14 @@ describe('computeYBounds', () => {
expect(computeYBounds([], 'progress')).toEqual({}); expect(computeYBounds([], 'progress')).toEqual({});
}); });
}); });
describe('isProposedToClose', () => {
const { isProposedToClose } = require('@/lib/utils');
it('is true only for the exact GIAS proposed-to-close status', () => {
expect(isProposedToClose({ status: 'Open, but proposed to close' })).toBe(true);
expect(isProposedToClose({ status: 'Open' })).toBe(false);
expect(isProposedToClose({ status: null })).toBe(false);
expect(isProposedToClose({})).toBe(false);
});
});
@@ -1543,3 +1543,18 @@
.historyDisclosure[open] > .historyToggle::before { .historyDisclosure[open] > .historyToggle::before {
transform: rotate(90deg); transform: rotate(90deg);
} }
/* GIAS "Open, but proposed to close" notice strip */
.closingStrip {
background: #fdf6e3;
border-left: 4px solid #e2c96f;
border-radius: 0 6px 6px 0;
padding: 0.55rem 0.9rem;
margin: 0.5rem 0;
font-size: 0.88rem;
color: #6e5a00;
max-width: 68ch;
}
.closingStrip strong {
color: #8a6200;
}
+7 -1
View File
@@ -18,7 +18,7 @@ import type {
SchoolDeprivation, SchoolFinance, NationalAverages, SchoolDeprivation, SchoolFinance, NationalAverages,
} from '@/lib/types'; } from '@/lib/types';
import { import {
formatPercentage, formatProgress, formatAcademicYear, formatPercentage, formatProgress, formatAcademicYear, isProposedToClose,
} from '@/lib/utils'; } from '@/lib/utils';
import { DeltaChip } from './DeltaChip'; import { DeltaChip } from './DeltaChip';
@@ -313,6 +313,12 @@ export function SchoolDetailView({
<span className={styles.metaItem}>{schoolInfo.gender}&apos;s school</span> <span className={styles.metaItem}>{schoolInfo.gender}&apos;s school</span>
)} )}
</div> </div>
{isProposedToClose(schoolInfo) && (
<div className={styles.closingStrip} role="note">
<strong> Proposed to close</strong> this school is proposed for closure,
check with the local authority before applying.
</div>
)}
{schoolInfo.address && ( {schoolInfo.address && (
<p className={styles.address}> <p className={styles.address}>
{schoolInfo.address}{schoolInfo.postcode && `, ${schoolInfo.postcode}`} {schoolInfo.address}{schoolInfo.postcode && `, ${schoolInfo.postcode}`}
@@ -254,3 +254,10 @@
justify-content: center; justify-content: center;
} }
} }
/* GIAS "Open, but proposed to close" marker */
.attrClosing {
background: #fdf6e3;
color: #8a6200;
border: 1px solid #e2c96f;
}
+4 -1
View File
@@ -9,7 +9,7 @@
*/ */
import type { School } from '@/lib/types'; import type { School } from '@/lib/types';
import { formatPercentage, calculateTrend, getPhaseStyle, schoolUrl, buildOfstedListBadge, formatAgeRange } from '@/lib/utils'; import { formatPercentage, calculateTrend, getPhaseStyle, schoolUrl, buildOfstedListBadge, formatAgeRange, isProposedToClose } from '@/lib/utils';
import styles from './SchoolRow.module.css'; import styles from './SchoolRow.module.css';
interface SchoolRowProps { interface SchoolRowProps {
@@ -78,6 +78,9 @@ export function SchoolRow({
{school.age_range && <span className={styles.attr}>{formatAgeRange(school.age_range)}</span>} {school.age_range && <span className={styles.attr}>{formatAgeRange(school.age_range)}</span>}
{showDenomination && <span className={styles.attr}>{school.religious_denomination}</span>} {showDenomination && <span className={styles.attr}>{school.religious_denomination}</span>}
{showGender && <span className={styles.attr}>{school.gender}</span>} {showGender && <span className={styles.attr}>{school.gender}</span>}
{isProposedToClose(school) && (
<span className={`${styles.attr} ${styles.attrClosing}`}> Proposed to close</span>
)}
</div> </div>
{/* Line 3: Key stats */} {/* Line 3: Key stats */}
@@ -1099,3 +1099,18 @@
padding: 0.75rem; padding: 0.75rem;
} }
} }
/* GIAS "Open, but proposed to close" notice strip */
.closingStrip {
background: #fdf6e3;
border-left: 4px solid #e2c96f;
border-radius: 0 6px 6px 0;
padding: 0.55rem 0.9rem;
margin: 0.5rem 0;
font-size: 0.88rem;
color: #6e5a00;
max-width: 68ch;
}
.closingStrip strong {
color: #8a6200;
}
@@ -23,7 +23,7 @@ import type {
SchoolAdmissions, SenDetail, Phonics, SchoolAdmissions, SenDetail, Phonics,
SchoolDeprivation, SchoolFinance, NationalAverages, SchoolDeprivation, SchoolFinance, NationalAverages,
} from '@/lib/types'; } from '@/lib/types';
import { formatPercentage, formatProgress, formatAcademicYear, formatAgeRange } from '@/lib/utils'; import { formatPercentage, formatProgress, formatAcademicYear, formatAgeRange, isProposedToClose } from '@/lib/utils';
import { DeltaChip } from './DeltaChip'; import { DeltaChip } from './DeltaChip';
import { track, getNavigationSource } from '@/lib/analytics'; import { track, getNavigationSource } from '@/lib/analytics';
import styles from './SecondarySchoolDetailView.module.css'; import styles from './SecondarySchoolDetailView.module.css';
@@ -237,6 +237,12 @@ export function SecondarySchoolDetailView({
</span> </span>
)} )}
</div> </div>
{isProposedToClose(schoolInfo) && (
<div className={styles.closingStrip} role="note">
<strong> Proposed to close</strong> this school is proposed for closure,
check with the local authority before applying.
</div>
)}
{schoolInfo.address && ( {schoolInfo.address && (
<p className={styles.address}> <p className={styles.address}>
{schoolInfo.address}{schoolInfo.postcode && `, ${schoolInfo.postcode}`} {schoolInfo.address}{schoolInfo.postcode && `, ${schoolInfo.postcode}`}
@@ -266,3 +266,9 @@
justify-content: center; justify-content: center;
} }
} }
.closingTag {
background: #fdf6e3;
color: #8a6200;
border: 1px solid #e2c96f;
}
+4 -1
View File
@@ -11,7 +11,7 @@
'use client'; 'use client';
import type { School } from '@/lib/types'; import type { School } from '@/lib/types';
import { buildOfstedListBadge, getPhaseStyle, schoolUrl, formatAgeRange } from '@/lib/utils'; import { buildOfstedListBadge, getPhaseStyle, schoolUrl, formatAgeRange, isProposedToClose } from '@/lib/utils';
import styles from './SecondarySchoolRow.module.css'; import styles from './SecondarySchoolRow.module.css';
function detectAdmissionsTag(school: School): string | null { function detectAdmissionsTag(school: School): string | null {
@@ -97,6 +97,9 @@ export function SecondarySchoolRow({
{admissionsTag} {admissionsTag}
</span> </span>
)} )}
{isProposedToClose(school) && (
<span className={`${styles.provisionTag} ${styles.closingTag}`}> Proposed to close</span>
)}
</div> </div>
{/* Line 3: KS4 stats */} {/* Line 3: KS4 stats */}
+1
View File
@@ -18,6 +18,7 @@ export interface School {
religious_denomination: string | null; religious_denomination: string | null;
age_range: string | null; age_range: string | null;
has_sixth_form?: boolean | null; has_sixth_form?: boolean | null;
status?: string | null; // GIAS establishment status ("Open" / "Open, but proposed to close")
// Address // Address
address1: string | null; address1: string | null;
+15
View File
@@ -718,3 +718,18 @@ export function buildOfstedListBadge(school: {
return { label: 'Not yet inspected', cssClass: 'ofstedPending' }; return { label: 'Not yet inspected', cssClass: 'ofstedPending' };
} }
// ============================================================================
// Establishment status
// ============================================================================
export const PROPOSED_TO_CLOSE_STATUS = 'Open, but proposed to close';
/**
* GIAS lists some operating schools as "Open, but proposed to close".
* They remain open (and may stay open if the proposal is withdrawn), but the
* UI marks them so families check with the local authority before applying.
*/
export function isProposedToClose(school: { status?: string | null }): boolean {
return school.status === PROPOSED_TO_CLOSE_STATUS;
}
+50 -5
View File
@@ -38,6 +38,31 @@ default_args = {
"retry_delay": timedelta(minutes=5), "retry_delay": timedelta(minutes=5),
} }
# The backend caches the marts DataFrame at startup; after any rebuild the
# cache must be invalidated or the API serves stale (or empty) data until the
# container restarts.
INVALIDATE_CACHE_CMD = """
set -e
BACKEND_URL="${BACKEND_URL:-http://backend:80}"
ADMIN_KEY="${ADMIN_API_KEY:-changeme}"
echo "Calling $BACKEND_URL/api/admin/reload ..."
response=$(curl -s -o /tmp/reload_response.json -w "%{http_code}" \\
--connect-timeout 10 --max-time 120 \\
-X POST "$BACKEND_URL/api/admin/reload" \\
-H "X-API-Key: $ADMIN_KEY" \\
-H "Content-Type: application/json")
echo "HTTP status: $response"
cat /tmp/reload_response.json
if [ "$response" != "200" ]; then
echo "ERROR: backend cache reload failed (HTTP $response)"
exit 1
fi
"""
# ── Daily DAG (GIAS + downstream) ────────────────────────────────────── # ── Daily DAG (GIAS + downstream) ──────────────────────────────────────
@@ -83,7 +108,7 @@ print(f'Validation passed: {{count}} GIAS rows')
dbt_build = BashOperator( dbt_build = BashOperator(
task_id="dbt_build", task_id="dbt_build",
bash_command=f"cd {PIPELINE_DIR}/transform && {DBT_BIN} build --profiles-dir . --target production --select stg_gias_establishments+ stg_gias_links+ --exclude int_ks2_with_lineage+ int_ks4_with_lineage+", bash_command=f"cd {PIPELINE_DIR}/transform && {DBT_BIN} build --profiles-dir . --target production --select stg_gias_establishments+ stg_gias_links+ gias_code_names+ --exclude int_ks2_with_lineage+ int_ks4_with_lineage+",
) )
sync_typesense = BashOperator( sync_typesense = BashOperator(
@@ -91,7 +116,12 @@ print(f'Validation passed: {{count}} GIAS rows')
bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py", bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py",
) )
extract_group >> validate_raw >> dbt_build >> sync_typesense invalidate_cache = BashOperator(
task_id="invalidate_cache",
bash_command=INVALIDATE_CACHE_CMD,
)
extract_group >> validate_raw >> dbt_build >> sync_typesense >> invalidate_cache
# ── Monthly DAG (Ofsted) ─────────────────────────────────────────────── # ── Monthly DAG (Ofsted) ───────────────────────────────────────────────
@@ -121,7 +151,12 @@ with DAG(
bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py", bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py",
) )
extract_ofsted >> dbt_build_ofsted >> sync_typesense_ofsted invalidate_cache_ofsted = BashOperator(
task_id="invalidate_cache",
bash_command=INVALIDATE_CACHE_CMD,
)
extract_ofsted >> dbt_build_ofsted >> sync_typesense_ofsted >> invalidate_cache_ofsted
# ── Annual DAG (EES: KS2, KS4, Census, Admissions) ─────────────────── # ── Annual DAG (EES: KS2, KS4, Census, Admissions) ───────────────────
@@ -153,7 +188,12 @@ with DAG(
bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py", bash_command=f"cd {PIPELINE_DIR} && python scripts/sync_typesense.py",
) )
extract_ees_group >> dbt_build_ees >> sync_typesense_ees invalidate_cache_ees = BashOperator(
task_id="invalidate_cache",
bash_command=INVALIDATE_CACHE_CMD,
)
extract_ees_group >> dbt_build_ees >> sync_typesense_ees >> invalidate_cache_ees
# ── Annual DAG (IDACI Deprivation) ──────────────────────────────────── # ── Annual DAG (IDACI Deprivation) ────────────────────────────────────
@@ -178,4 +218,9 @@ with DAG(
bash_command=f"cd {PIPELINE_DIR}/transform && {DBT_BIN} build --profiles-dir . --target production --select stg_idaci+ fact_deprivation+", bash_command=f"cd {PIPELINE_DIR}/transform && {DBT_BIN} build --profiles-dir . --target production --select stg_idaci+ fact_deprivation+",
) )
extract_idaci >> dbt_build_idaci invalidate_cache_idaci = BashOperator(
task_id="invalidate_cache",
bash_command=INVALIDATE_CACHE_CMD,
)
extract_idaci >> dbt_build_idaci >> invalidate_cache_idaci
@@ -31,16 +31,22 @@ class GIASEstablishmentsStream(Stream):
schema = th.PropertiesList( schema = th.PropertiesList(
th.Property("URN", th.IntegerType, required=True), th.Property("URN", th.IntegerType, required=True),
th.Property("EstablishmentName", th.StringType), th.Property("EstablishmentName", th.StringType),
th.Property("TypeOfEstablishment (code)", th.StringType),
th.Property("TypeOfEstablishment (name)", th.StringType), th.Property("TypeOfEstablishment (name)", th.StringType),
th.Property("PhaseOfEducation (code)", th.StringType),
th.Property("PhaseOfEducation (name)", th.StringType), th.Property("PhaseOfEducation (name)", th.StringType),
th.Property("OfficialSixthForm (code)", th.StringType),
th.Property("OfficialSixthForm (name)", th.StringType), th.Property("OfficialSixthForm (name)", th.StringType),
th.Property("LA (code)", th.StringType), th.Property("LA (code)", th.StringType),
th.Property("LA (name)", th.StringType), th.Property("LA (name)", th.StringType),
th.Property("EstablishmentNumber", th.StringType), th.Property("EstablishmentNumber", th.StringType),
th.Property("EstablishmentStatus (code)", th.StringType),
th.Property("EstablishmentStatus (name)", th.StringType), th.Property("EstablishmentStatus (name)", th.StringType),
th.Property("Postcode", th.StringType), th.Property("Postcode", th.StringType),
th.Property("Gender (name)", th.StringType), th.Property("Gender (name)", th.StringType),
th.Property("ReligiousCharacter (code)", th.StringType),
th.Property("ReligiousCharacter (name)", th.StringType), th.Property("ReligiousCharacter (name)", th.StringType),
th.Property("AdmissionsPolicy (code)", th.StringType),
th.Property("AdmissionsPolicy (name)", th.StringType), th.Property("AdmissionsPolicy (name)", th.StringType),
th.Property("SchoolCapacity", th.StringType), th.Property("SchoolCapacity", th.StringType),
th.Property("NumberOfPupils", th.StringType), th.Property("NumberOfPupils", th.StringType),
@@ -68,6 +68,19 @@ COLUMN_PRIORITY = {
"ungraded_inspection_date": [ "ungraded_inspection_date": [
"Date of latest ungraded inspection", "Date of latest ungraded inspection",
], ],
# Report Card fields (post-Nov 2025 framework). Confirmed verbatim MI
# headers per diagnose_compare_gaps.py's Task 1(c) findings. No MI column
# currently exists for early-years or sixth-form report-card grades, so
# those two fields are deliberately omitted here (see schema below) --
# they stay absent from every record, same as the existing `report_url`
# pattern for fields with no COLUMN_PRIORITY entry.
"rc_safeguarding_met": ["Safeguarding standards"],
"rc_inclusion": ["Inclusion"],
"rc_curriculum_teaching": ["Curriculum and teaching"],
"rc_achievement": ["Achievement"],
"rc_attendance_behaviour": ["Attendance and behaviour"],
"rc_personal_development": ["Personal development and wellbeing"],
"rc_leadership_governance": ["Leadership and governance"],
} }
@@ -111,6 +124,17 @@ class OfstedInspectionsStream(Stream):
th.Property("sixth_form_provision", th.StringType), th.Property("sixth_form_provision", th.StringType),
th.Property("ungraded_outcome", th.StringType), th.Property("ungraded_outcome", th.StringType),
th.Property("ungraded_inspection_date", th.StringType), th.Property("ungraded_inspection_date", th.StringType),
th.Property("rc_safeguarding_met", th.StringType),
th.Property("rc_inclusion", th.StringType),
th.Property("rc_curriculum_teaching", th.StringType),
th.Property("rc_achievement", th.StringType),
th.Property("rc_attendance_behaviour", th.StringType),
th.Property("rc_personal_development", th.StringType),
th.Property("rc_leadership_governance", th.StringType),
# No MI column exists for these yet; declared for forward
# compatibility with the mart schema, always emitted as absent/NULL.
th.Property("rc_early_years", th.StringType),
th.Property("rc_sixth_form", th.StringType),
th.Property("report_url", th.StringType), th.Property("report_url", th.StringType),
).to_dict() ).to_dict()
+305
View File
@@ -0,0 +1,305 @@
"""Diagnose the three data gaps blocking the compare-screen redesign.
Run from repo root (network access required, no DB needed):
uv run --with singer-sdk --with pandas --with requests \
python pipeline/scripts/diagnose_compare_gaps.py
(singer_sdk is a transitive import of tap_uk_ees.tap / tap_uk_ofsted.tap and
is not part of the repo's default environment, hence the `uv run --with`.)
"""
import io
import re
import sys
import pandas as pd
import requests
sys.path.insert(0, "pipeline/plugins/extractors/tap-uk-ees")
sys.path.insert(0, "pipeline/plugins/extractors/tap-uk-ofsted")
from tap_uk_ees.tap import ( # noqa: E402
_KS2_NATIONAL_COL_MAP,
_KS2_NATIONAL_CSV_URL,
download_release_zip,
get_all_releases,
)
from tap_uk_ofsted.tap import discover_csv_url # noqa: E402
TIMEOUT = 120
def check_national_gps_science():
print("\n=== (a) National catalogue CSV: GPS/science columns ===")
resp = requests.get(_KS2_NATIONAL_CSV_URL, timeout=TIMEOUT)
resp.raise_for_status()
df = pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False)
df.columns = [c.strip().lower() for c in df.columns]
for csv_col in ("pt_gps_exp", "pt_scita_exp", "avg_readscore", "avg_matscore", "avg_gpsscore"):
status = "PRESENT" if csv_col in df.columns else "MISSING"
print(f" {csv_col}: {status}")
gps_like = [c for c in df.columns if "gps" in c or "scita" in c or "sci" in c]
print(f" all gps/science-ish columns: {gps_like}")
if "geographic_level" in df.columns:
nat = df[df["geographic_level"].str.strip().str.lower() == "national"]
else:
print(" geographic_level column missing — cannot isolate national rows")
return
print(f" national rows time_periods: {sorted(nat['time_period'].unique())}")
# Sample the values our map would read for the latest year
latest = nat[nat["time_period"] == nat["time_period"].max()]
for csv_col, field in _KS2_NATIONAL_COL_MAP.items():
val = latest.iloc[0].get(csv_col, "<col missing>") if len(latest) else "<no row>"
print(f" {field} <- {csv_col} = {val!r}")
def check_ks2_attainment_years_subjects():
print("\n=== (b) EES KS2 attainment: years & subject labels ===")
releases = get_all_releases("key-stage-2-attainment")
print(f" releases found: {[r['time_period'] for r in releases]}")
for release in releases:
try:
zf = download_release_zip(release["id"])
except Exception as e:
print(f" {release['time_period']}: DOWNLOAD FAILED: {e}")
continue
name = next((n for n in zf.namelist()
if "ks2_school_attainment_data" in n and n.endswith(".csv")), None)
if not name:
print(f" {release['time_period']}: NO school attainment CSV in ZIP")
print(f" all CSVs in zip: {[n for n in zf.namelist() if n.endswith('.csv')]}")
continue
with zf.open(name) as f:
df = pd.read_csv(f, dtype=str, keep_default_na=False, nrows=200000)
years = sorted(df["time_period"].unique())
subjects = sorted(df["subject"].unique())
print(f" release {release['time_period']}: time_periods={years}")
print(f" subjects={subjects}")
def check_ofsted_report_card_columns():
print("\n=== (c) Ofsted MI CSV: report-card columns ===")
url = discover_csv_url()
print(f" MI file: {url}")
if url is None or not url.lower().endswith(".csv"):
print(f" URL is not a CSV (likely ODS) — stopping this section. url={url!r}")
return
resp = requests.get(url, timeout=TIMEOUT)
resp.raise_for_status()
df = pd.read_csv(io.BytesIO(resp.content), dtype=str, keep_default_na=False, nrows=5)
rc_like = [c for c in df.columns
if re.search(r"report card|inclusion|curriculum|achievement|safeguard|well.?being|governance", c, re.I)]
print(f" candidate report-card columns ({len(rc_like)}):")
for c in rc_like:
print(f" - {c!r}")
print(f" all columns ({len(df.columns)}):")
for c in df.columns:
print(f" - {c!r}")
if __name__ == "__main__":
check_national_gps_science()
check_ks2_attainment_years_subjects()
check_ofsted_report_card_columns()
# FINDINGS 2026-07-12: run via
# uv run --with singer-sdk --with pandas --with requests \
# python pipeline/scripts/diagnose_compare_gaps.py
#
# (a) National catalogue CSV (GPS/science) — NOT a source-data problem.
# pt_gps_exp, pt_scita_exp, avg_readscore, avg_matscore, avg_gpsscore are
# all PRESENT in the catalogue CSV and hold real numeric values for the
# latest national row (time_period 202425: pt_gps_exp='72.6' ->
# gps_expected_pct; pt_scita_exp='81.6' -> science_expected_pct).
# national time_periods present: 201516, 201617, 201718, 201819, 201920,
# 202021, 202122, 202223, 202324, 202425 (COVID years 201920/202021 are
# present as rows but suppressed with 'x' per the module docstring, not
# absent). So _KS2_NATIONAL_COL_MAP is correct and the extractor's own
# read of the source is fine end-to-end -- the NULLs in
# marts.fact_ks2_national_averages are NOT caused by a missing/renamed
# source column. The gap must be introduced downstream of the tap
# (staging/mart SQL, a stale/incomplete load, or a dbt model not
# selecting these two columns) -- Task 5/6 should look at the dbt
# staging model for ees_ks2_national and the mart definition, not the
# tap/column-map.
#
# (b) EES KS2 attainment (school-level, "key-stage-2-attainment" publication)
# releases found (via get_all_releases): [None, '202425', '202324',
# '202223', '202122']. The `None` entry is the *current/latest* release
# (its slug doesn't parse to a 6-digit time_period by _slug_to_time_period,
# but the CSV inside carries time_period='202425' -- same data as the
# 202425-labelled release).
#
# Only two of the four releases contain a school-level attainment CSV
# matching "ks2_school_attainment_data*.csv":
# - release None (latest): HAS IT -> time_periods=['202425']
# subjects=['Grammar, punctuation and spelling', 'Maths', 'Reading',
# 'Reading, writing and maths', 'Science', 'Writing']
# - release 202324: HAS IT -> time_periods=['202324']
# subjects= same 6 labels as above
# - release 202223: NO school attainment CSV in ZIP. This
# release's ZIP instead contains only LA/regional/national/MAT-level
# files (e.g. ks2_regional_and_local_authority_*, ks2_multi_academy
# _trusts_*, ks2_national_*); no data/*school*attainment*.csv file
# exists at all in this release's package. This CONFIRMS the
# "subject-level 2022/23 is NULL in prod" symptom: the source
# release literally does not publish a school-level attainment file
# for 202223 under this filename pattern -- it's not a tap bug.
# - release 202122: NO school attainment CSV in ZIP. Same
# situation: ZIP has only LA/regional/national-level files (e.g.
# ks2_regional_and_local_authority_2016_to_2022_revised.csv,
# ks2_national_school_characteristics_2016_to_2022_revised.csv);
# no school-level attainment CSV present. This CONFIRMS "school-level
# 2021/22 is absent" -- again a genuine source-data absence, not an
# extractor bug.
# Implication for Tasks 5/6/7: 202122 and 202223 school-level attainment
# cannot be backfilled from the "key-stage-2-attainment" EES publication
# via this filename pattern -- those two years must either be sourced
# from a different EES dataset/file (e.g. one of the *_school_location_
# and_pupil_characteristics or *_school_type_and_pupil_characteristics
# files present in those ZIPs, which may carry school-level rows under a
# different filename), left NULL with an explicit "source unavailable"
# note, or backfilled from the legacy DfE "Compare School Performance"
# wide-format CSVs referenced elsewhere in tap.py. Subject labels to use
# when a source *is* found for 202324/202425:
# 'Grammar, punctuation and spelling', 'Maths', 'Reading',
# 'Reading, writing and maths', 'Science', 'Writing'
# (Reading, writing and maths spans reading+writing+maths combined --
# this is the RWM row.)
#
# (c) Ofsted MI CSV (report-card columns) — confirmed PRESENT.
# discover_csv_url() resolved to (as at run time, latest inspections
# 31 May 2026):
# https://assets.publishing.service.gov.uk/media/6a27c45be13080622db38815/
# Management_information_-_state-funded_schools_-_latest_inspections_as_at_31_May_2026.csv
# This is a real .csv (not .ods) so section (c) ran to completion.
# Exact report-card column headers (7 grade columns + their paired date
# columns, all present verbatim, case/spacing exactly as below):
# 'Safeguarding standards' / 'Safeguarding standards - date of grade'
# 'Inclusion' / 'Inclusion - date of grade'
# 'Curriculum and teaching' / 'Curriculum and teaching - date of grade'
# 'Achievement' / 'Achievement - date of grade'
# 'Attendance and behaviour' / 'Attendance and behaviour - date of grade'
# 'Personal development and wellbeing' / 'Personal development and wellbeing - date of grade'
# 'Leadership and governance' / 'Leadership and governance - date of grade'
# Plus a related pass/fail-style field:
# 'Latest OEIF safeguarding is effective?' (note: double space in the
# header, verbatim from source -- preserve exactly when mapping)
# These are the new-style "report card" single-word-area grades
# (introduced alongside the "Attendance and behaviour" split from
# "Personal development"); they coexist in the same CSV with the legacy
# 5-judgement OEIF columns ('Latest OEIF overall effectiveness',
# 'Latest OEIF quality of education', 'Latest OEIF behaviour and
# attitudes', 'Latest OEIF personal development', 'Latest OEIF
# effectiveness of leadership and management'). Task 7 should map the 7
# report-card columns above (grade + date pairs, 6 of them, plus the
# safeguarding-effective flag) rather than inventing new column names.
# TASK 6 VERIFICATION 2026-07-12: 2021/22 legacy KS2 school-level archive
#
# RESULT: BLOCKED at the source-data level. School-level KS2 attainment for
# academic year 2021/22 was never published anywhere publicly by DfE -- not
# in EES (confirmed by Task 1's finding (b) above), not in the legacy
# "Compare School Performance" download wizard, and not as a standalone
# performance-tables archive/ODS on assets.publishing.service.gov.uk. This
# is a deliberate DfE decision, not a gap in our extraction logic.
#
# Confirming quote (Key stage 2 attainment 2021/22 release notes, via
# https://explore-education-statistics.service.gov.uk/find-statistics/
# key-stage-2-attainment/2021-22):
# "We will not publish key stage 2 data for academic year 2021/22 in
# performance tables (also known as Compare School and College
# Performance)." ... "The Department will, however, still produce the
# normal suite of key stage 2 accountability measures at school and
# multi-academy trust level and share these securely with primary
# schools, academy trusts and local authorities to inform school
# improvement discussions."
# (i.e. school-level 202122 KS2 results exist internally at DfE but were
# withheld from every public channel: performance tables/CSCP, EES, and by
# extension the legacy DfE archives the current legacy_ks2_urls entries in
# meltano.yml were sourced from.)
#
# What was tried:
# 1. Direct download URL pattern from the task brief:
# https://www.compare-school-performance.service.gov.uk/download-data?download=true&regions=0&filters=KS2&fileformat=csv&year=2021-2022&meta=false
# -> HTTP 404, HTML error page (not a CSV/ZIP). Saved response inspected;
# confirmed 404 via response headers (`content-type: text/html`).
# 2. Walked the actual multi-step download wizard at
# https://www.compare-school-performance.service.gov.uk/download-data
# with a browser User-Agent and a cookie jar, replicating the GET-based
# form steps: currentstep=year (downloadYear=2021-2022) -> currentstep=
# region (regiontype=all&la=0) -> currentstep=datatypes. On the final
# "datatypes" step, the checkbox list for 2021-2022 has NO "ks2" (or
# "ks2mats") option at all -- only ks4/ks4prov/ks4underlying/ks5* /
# pupil-destination/absence/census/mats checkboxes are present.
# Control check: repeating the same wizard walk for downloadYear=
# 2018-2019, 2022-2023 and 2023-2024 shows a "ks2" (and "ks2mats")
# checkbox present in all three; downloadYear=2020-2021 (COVID-cancelled
# KS2 SATs year) also has NO ks2 checkbox, matching the pattern for a
# year where school-level KS2 genuinely isn't published. 2021-2022
# behaves identically to the cancelled 2020-2021 year, not like the
# normal 2018-2019/2022-2023/2023-2024 years.
# 3. Web search for a standalone KS2 2022 performance-tables archive
# (e.g. "england_ks2final" for 2022) on assets.publishing.service.gov.uk
# found no such file; only unrelated 2022/2023-dated documents.
#
# No ZIP was ever obtained -- /tmp/dfe-2021-2022-ks2.zip contains the 404
# HTML error page from attempt (1) above, not a real archive. It contains
# no england_ks2final.csv (there is no ZIP to look inside).
#
# Column-map check (brief's Step 1): NOT RUN -- there is no 2021/22
# england_ks2final.csv to check headers against. This is moot until/unless
# a non-public source (e.g. a manual/internal DfE extract) becomes
# available; _LEGACY_KS2_COLUMN_MAP itself is unchanged and untested here.
#
# Recommendation: mark 202122 school-level KS2 as a genuine, permanent
# source-data gap (not a backfill candidate) unless the project can obtain
# the internal DfE extract DfE says it shared "securely with primary
# schools, academy trusts and local authorities" -- that is not a route
# available to this pipeline. Task 6's meltano.yml change (Step 2) and the
# filebrowser upload should NOT proceed for 202122; there is nothing to
# upload.
# TASK 7 VALUE SAMPLE 2026-07-12: live value_counts() over the 7 report-card
# columns (plus the related safeguarding-effective flag) in the same MI CSV
# resolved by discover_csv_url() as at run time (31 May 2026 inspections
# file). Blank cells read as the literal string 'NULL' (matches
# keep_default_na=False in tap.py). Observed non-blank values, verbatim:
#
# 'Safeguarding standards': 'Met' (1319), 'Not met' (10)
# 'Inclusion': 'Expected standard' (710),
# 'Strong standard' (447), 'Needs attention' (130), 'Exceptional' (23),
# 'Urgent improvement' (19)
# 'Curriculum and teaching': 'Expected standard' (797),
# 'Needs attention' (287), 'Strong standard' (206),
# 'Urgent improvement' (28), 'Exceptional' (11)
# 'Achievement': 'Expected standard' (701),
# 'Needs attention' (364), 'Strong standard' (207),
# 'Urgent improvement' (39), 'Exceptional' (18)
# 'Attendance and behaviour': 'Expected standard' (699),
# 'Strong standard' (405), 'Needs attention' (188),
# 'Urgent improvement' (21), 'Exceptional' (16)
# 'Personal development and wellbeing': 'Expected standard' (728),
# 'Strong standard' (504), 'Needs attention' (66), 'Exceptional' (23),
# 'Urgent improvement' (8)
# 'Leadership and governance': 'Expected standard' (813),
# 'Strong standard' (292), 'Needs attention' (172),
# 'Urgent improvement' (34), 'Exceptional' (18)
# 'Latest OEIF safeguarding is effective?' (note double space, not used by
# Task 7 -- kept for completeness): 'Yes' (12970), 'No' (96)
#
# So the 6 graded report-card columns share exactly one 5-value vocabulary:
# {'Exceptional', 'Strong standard', 'Expected standard', 'Needs attention',
# 'Urgent improvement'} -- no 'Attention needed' variant was observed
# anywhere, so parse_report_card_grade.sql does NOT need that speculative
# branch from the task brief. 'Safeguarding standards' is a separate
# two-value vocabulary {'Met', 'Not met'}.
#
# Collision check: 'Achievement' matches by EXACT list-membership
# (`candidate in df_columns`, a Python list containment check against the
# full column-name list, not a substring/regex match) against only
# ['Achievement', 'Achievement - date of grade'] -- the date-paired column
# has a different exact string and is never selected. Same check for
# 'Safeguarding standards' found only itself, its own date-of-grade column,
# and the unrelated 'Latest OEIF safeguarding is effective?' column (not
# mapped to any rc_* field). No legacy OEIF column is accidentally consumed
# by an rc_ mapping.
+144
View File
@@ -0,0 +1,144 @@
"""Generate GIAS code->name dictionaries from the live bulk CSV.
Writes:
- backend/gias_codes.py (canonical Python module)
- pipeline/scripts/gias_codes.py (byte-identical copy)
- pipeline/transform/seeds/gias_code_names.csv (dbt seed for drift test)
Run from the repo root whenever the dbt drift test warns that DfE
added/renamed a value: python pipeline/scripts/generate_gias_codes.py
"""
from __future__ import annotations
import io
import sys
from datetime import date, timedelta
from pathlib import Path
import pandas as pd
import requests
GIAS_URL = (
"https://ea-edubase-api-prod.azurewebsites.net"
"/edubase/downloads/public/edubasealldata{date}.csv"
)
# (CSV code column, CSV name column, python dict name, seed field key)
FIELDS = [
("TypeOfEstablishment (code)", "TypeOfEstablishment (name)", "SCHOOL_TYPE", "school_type"),
("EstablishmentStatus (code)", "EstablishmentStatus (name)", "ESTABLISHMENT_STATUS", "establishment_status"),
("PhaseOfEducation (code)", "PhaseOfEducation (name)", "PHASE_OF_EDUCATION", "phase_of_education"),
("OfficialSixthForm (code)", "OfficialSixthForm (name)", "OFFICIAL_SIXTH_FORM", "official_sixth_form"),
("ReligiousCharacter (code)", "ReligiousCharacter (name)", "RELIGIOUS_CHARACTER", "religious_character"),
("AdmissionsPolicy (code)", "AdmissionsPolicy (name)", "ADMISSIONS_POLICY", "admissions_policy"),
]
MODULE_HEADER = '''"""GIAS code -> name dictionaries.
GENERATED by pipeline/scripts/generate_gias_codes.py from the GIAS bulk CSV
— do not edit by hand; rerun the script when the dbt drift test warns.
The canonical file is backend/gias_codes.py; pipeline/scripts/gias_codes.py
must be byte-identical (enforced by backend/tests/test_gias_codes.py).
"""
from __future__ import annotations
import logging
import math
logger = logging.getLogger(__name__)
'''
MODULE_FOOTER = '''
def translate(code, mapping: dict[int, str]) -> str | None:
"""Translate a GIAS code to its display name.
None/NaN -> None (column absent or suppressed). Unknown codes degrade to
"Unknown (<code>)" with a warning so a new DfE value never blanks the UI.
"""
if code is None or (isinstance(code, float) and math.isnan(code)):
return None
code = int(code)
if code not in mapping:
logger.warning("Unknown GIAS code %s (not in dictionary)", code)
return f"Unknown ({code})"
return mapping[code]
'''
def download_csv() -> pd.DataFrame:
for day in (date.today(), date.today() - timedelta(days=1)):
url = GIAS_URL.format(date=day.strftime("%Y%m%d"))
print(f"Downloading {url}")
resp = requests.get(url, timeout=300)
if resp.status_code == 404:
continue
resp.raise_for_status()
return pd.read_csv(
io.StringIO(resp.content.decode("latin-1")),
dtype=str, keep_default_na=False,
)
sys.exit("GIAS CSV not available for today or yesterday")
def main() -> None:
repo = Path(__file__).resolve().parents[2]
df = download_csv()
module_parts = [MODULE_HEADER]
seed_rows: list[tuple[str, int, str]] = []
for code_col, name_col, dict_name, field_key in FIELDS:
pairs = (
df[[code_col, name_col]]
.loc[lambda d: d[code_col] != ""]
.drop_duplicates()
)
by_code: dict[int, set] = {}
for c, n in pairs.itertuples(index=False):
by_code.setdefault(int(c), set()).add(n)
mapping = []
for code, names in sorted(by_code.items()):
named = sorted(n for n in names if n != "")
if len(named) > 1:
sys.exit(f"{code_col}: code {code} maps to multiple names {named} — investigate before generating")
# Codes that only ever appear with a blank (name) are GIAS
# "not recorded" sentinels (e.g. ReligiousCharacter 99,
# AdmissionsPolicy 9). Map them to "" so the API serves the same
# empty string the old name pipeline did — the "Unknown (<code>)"
# path is reserved for genuinely new codes.
mapping.append((code, named[0] if named else ""))
lines = [f"{dict_name}: dict[int, str] = {{"]
for code, name in mapping:
escaped = name.replace('"', '\\"')
lines.append(f' {code}: "{escaped}",')
lines.append("}\n")
module_parts.append("\n".join(lines))
seed_rows += [(field_key, code, name) for code, name in mapping]
module = "\n".join(module_parts) + MODULE_FOOTER
(repo / "backend" / "gias_codes.py").write_text(module)
(repo / "pipeline" / "scripts" / "gias_codes.py").write_text(module)
seed_path = repo / "pipeline" / "transform" / "seeds" / "gias_code_names.csv"
with open(seed_path, "w", newline="") as fh:
import csv
w = csv.writer(fh)
w.writerow(["field", "code", "name"])
w.writerows(seed_rows)
print(f"Wrote backend/gias_codes.py, pipeline/scripts/gias_codes.py, {seed_path.name}")
print("\nKey codes for the dbt work (Task 3):")
for field in ("establishment_status", "phase_of_education", "official_sixth_form"):
print(f" {field}:")
for f, code, name in seed_rows:
if f == field:
print(f" {code} = {name}")
if __name__ == "__main__":
main()
+155
View File
@@ -0,0 +1,155 @@
"""GIAS code -> name dictionaries.
GENERATED by pipeline/scripts/generate_gias_codes.py from the GIAS bulk CSV
— do not edit by hand; rerun the script when the dbt drift test warns.
The canonical file is backend/gias_codes.py; pipeline/scripts/gias_codes.py
must be byte-identical (enforced by backend/tests/test_gias_codes.py).
"""
from __future__ import annotations
import logging
import math
logger = logging.getLogger(__name__)
SCHOOL_TYPE: dict[int, str] = {
1: "Community school",
2: "Voluntary aided school",
3: "Voluntary controlled school",
5: "Foundation school",
6: "City technology college",
7: "Community special school",
8: "Non-maintained special school",
10: "Other independent special school",
11: "Other independent school",
12: "Foundation special school",
14: "Pupil referral unit",
15: "Local authority nursery school",
18: "Further education",
24: "Secure units",
25: "Offshore schools",
26: "Service children's education",
27: "Miscellaneous",
28: "Academy sponsor led",
29: "Higher education institutions",
30: "Welsh establishment",
31: "Sixth form centres",
32: "Special post 16 institution",
33: "Academy special sponsor led",
34: "Academy converter",
35: "Free schools",
36: "Free schools special",
37: "British schools overseas",
38: "Free schools alternative provision",
39: "Free schools 16 to 19",
40: "University technical college",
41: "Studio schools",
42: "Academy alternative provision converter",
43: "Academy alternative provision sponsor led",
44: "Academy special converter",
45: "Academy 16-19 converter",
46: "Academy 16 to 19 sponsor led",
49: "Online provider",
56: "Institution funded by other government department",
57: "Academy secure 16 to 19",
}
ESTABLISHMENT_STATUS: dict[int, str] = {
1: "Open",
2: "Closed",
3: "Open, but proposed to close",
4: "Proposed to open",
}
PHASE_OF_EDUCATION: dict[int, str] = {
0: "Not applicable",
1: "Nursery",
2: "Primary",
3: "Middle deemed primary",
4: "Secondary",
5: "Middle deemed secondary",
6: "16 plus",
7: "All-through",
}
OFFICIAL_SIXTH_FORM: dict[int, str] = {
0: "Not applicable",
1: "Has a sixth form",
2: "Does not have a sixth form",
9: "",
}
RELIGIOUS_CHARACTER: dict[int, str] = {
0: "Does not apply",
2: "Church of England",
3: "Roman Catholic",
4: "Methodist",
5: "Jewish",
6: "None",
7: "Muslim",
8: "Seventh Day Adventist",
9: "Church of England/Methodist",
10: "Methodist/Church of England",
11: "Church of England/Roman Catholic",
12: "Church of England/United Reformed Church",
13: "Roman Catholic/Church of England",
14: "Quaker",
15: "Christian",
16: "United Reformed Church",
17: "Congregational Church",
18: "Free Church",
19: "Church of England/Free Church",
20: "Church of England/Christian",
21: "Sikh",
22: "Greek Orthodox",
24: "Buddhist",
25: "Hindu",
26: "Moravian",
28: "Inter- / non- denominational",
29: "Multi-faith",
30: "Church of England/Methodist/United Reform Church/Baptist",
31: "Anglican",
32: "Anglican/Christian",
33: "Anglican/Evangelical",
34: "Anglican/Church of England",
35: "Catholic",
36: "Charadi Jewish",
37: "Christian/Evangelical",
38: "Christian Science",
39: "Christian/Methodist",
40: "Christian/non-denominational",
41: "Church of England/Evangelical",
42: "Islam",
43: "Orthodox Jewish",
44: "Plymouth Brethren Christian Church",
45: "Protestant",
46: "Protestant/Evangelical",
47: "Reformed Baptist",
48: "Roman Catholic/Anglican",
49: "Sunni Deobandi",
99: "",
}
ADMISSIONS_POLICY: dict[int, str] = {
0: "Not applicable",
2: "Selective",
4: "Non-selective",
9: "",
}
def translate(code, mapping: dict[int, str]) -> str | None:
"""Translate a GIAS code to its display name.
None/NaN -> None (column absent or suppressed). Unknown codes degrade to
"Unknown (<code>)" with a warning so a new DfE value never blanks the UI.
"""
if code is None or (isinstance(code, float) and math.isnan(code)):
return None
code = int(code)
if code not in mapping:
logger.warning("Unknown GIAS code %s (not in dictionary)", code)
return f"Unknown ({code})"
return mapping[code]
+10 -7
View File
@@ -19,6 +19,8 @@ import psycopg2
import psycopg2.extras import psycopg2.extras
import typesense import typesense
from gias_codes import PHASE_OF_EDUCATION, RELIGIOUS_CHARACTER, SCHOOL_TYPE, translate
COLLECTION_SCHEMA = { COLLECTION_SCHEMA = {
"fields": [ "fields": [
{"name": "urn", "type": "int32"}, {"name": "urn", "type": "int32"},
@@ -44,10 +46,10 @@ QUERY_BASE = """
SELECT SELECT
s.urn, s.urn,
s.school_name, s.school_name,
s.phase, s.phase_code,
s.school_type, s.school_type_code,
l.local_authority_name as local_authority, l.local_authority_name as local_authority,
s.religious_character, s.religious_character_code,
s.ofsted_grade, s.ofsted_grade,
l.postcode, l.postcode,
s.headteacher_name, s.headteacher_name,
@@ -85,14 +87,15 @@ def build_document(row: dict) -> dict:
"id": str(row["urn"]), "id": str(row["urn"]),
"urn": row["urn"], "urn": row["urn"],
"school_name": row["school_name"] or "", "school_name": row["school_name"] or "",
"phase": row["phase"] or "", "phase": translate(row["phase_code"], PHASE_OF_EDUCATION) or "",
"school_type": row["school_type"] or "", "school_type": translate(row["school_type_code"], SCHOOL_TYPE) or "",
"local_authority": row["local_authority"] or "", "local_authority": row["local_authority"] or "",
"postcode": row["postcode"] or "", "postcode": row["postcode"] or "",
} }
if row.get("religious_character"): religious_character = translate(row.get("religious_character_code"), RELIGIOUS_CHARACTER)
doc["religious_character"] = row["religious_character"] if religious_character:
doc["religious_character"] = religious_character
if row.get("ofsted_grade"): if row.get("ofsted_grade"):
doc["ofsted_rating"] = OFSTED_LABELS.get(row["ofsted_grade"], "") doc["ofsted_rating"] = OFSTED_LABELS.get(row["ofsted_grade"], "")
if row.get("headteacher_name"): if row.get("headteacher_name"):
@@ -0,0 +1,17 @@
-- Macro: Parse Ofsted Report Card grade (post-Nov 2025 framework) from text
-- into the 5-point scale. Real values confirmed via a live sample of the MI
-- CSV (see pipeline/scripts/diagnose_compare_gaps.py's
-- "TASK 7 VALUE SAMPLE 2026-07-12" note) -- unrecognised text (including the
-- 'NULL' sentinel used by the source CSV for blanks) parses to NULL, never
-- errors.
{% macro parse_report_card_grade(column_name) %}
case lower(trim(nullif({{ column_name }}, 'NULL')))
when 'exceptional' then 1
when 'strong standard' then 2
when 'expected standard' then 3
when 'needs attention' then 4
when 'urgent improvement' then 5
else null
end
{% endmacro %}
@@ -15,8 +15,11 @@ current_ks2 as (
year, total_pupils, eligible_pupils, year, total_pupils, eligible_pupils,
rwm_expected_pct, rwm_high_pct, rwm_expected_pct, rwm_high_pct,
reading_expected_pct, reading_high_pct, reading_avg_score, reading_progress, reading_expected_pct, reading_high_pct, reading_avg_score, reading_progress,
reading_progress_lower_ci, reading_progress_upper_ci,
writing_expected_pct, writing_high_pct, writing_progress, writing_expected_pct, writing_high_pct, writing_progress,
writing_progress_lower_ci, writing_progress_upper_ci, writing_working_towards_pct,
maths_expected_pct, maths_high_pct, maths_avg_score, maths_progress, maths_expected_pct, maths_high_pct, maths_avg_score, maths_progress,
maths_progress_lower_ci, maths_progress_upper_ci,
gps_expected_pct, gps_high_pct, gps_avg_score, science_expected_pct, gps_expected_pct, gps_high_pct, gps_avg_score, science_expected_pct,
reading_absence_pct, writing_absence_pct, maths_absence_pct, gps_absence_pct, science_absence_pct, reading_absence_pct, writing_absence_pct, maths_absence_pct, gps_absence_pct, science_absence_pct,
rwm_expected_boys_pct, rwm_high_boys_pct, rwm_expected_girls_pct, rwm_high_girls_pct, rwm_expected_boys_pct, rwm_high_boys_pct, rwm_expected_girls_pct, rwm_high_girls_pct,
@@ -33,8 +36,11 @@ predecessor_ks2 as (
ks2.year, ks2.total_pupils, ks2.eligible_pupils, ks2.year, ks2.total_pupils, ks2.eligible_pupils,
ks2.rwm_expected_pct, ks2.rwm_high_pct, ks2.rwm_expected_pct, ks2.rwm_high_pct,
ks2.reading_expected_pct, ks2.reading_high_pct, ks2.reading_avg_score, ks2.reading_progress, ks2.reading_expected_pct, ks2.reading_high_pct, ks2.reading_avg_score, ks2.reading_progress,
ks2.reading_progress_lower_ci, ks2.reading_progress_upper_ci,
ks2.writing_expected_pct, ks2.writing_high_pct, ks2.writing_progress, ks2.writing_expected_pct, ks2.writing_high_pct, ks2.writing_progress,
ks2.writing_progress_lower_ci, ks2.writing_progress_upper_ci, ks2.writing_working_towards_pct,
ks2.maths_expected_pct, ks2.maths_high_pct, ks2.maths_avg_score, ks2.maths_progress, ks2.maths_expected_pct, ks2.maths_high_pct, ks2.maths_avg_score, ks2.maths_progress,
ks2.maths_progress_lower_ci, ks2.maths_progress_upper_ci,
ks2.gps_expected_pct, ks2.gps_high_pct, ks2.gps_avg_score, ks2.science_expected_pct, ks2.gps_expected_pct, ks2.gps_high_pct, ks2.gps_avg_score, ks2.science_expected_pct,
ks2.reading_absence_pct, ks2.writing_absence_pct, ks2.maths_absence_pct, ks2.gps_absence_pct, ks2.science_absence_pct, ks2.reading_absence_pct, ks2.writing_absence_pct, ks2.maths_absence_pct, ks2.gps_absence_pct, ks2.science_absence_pct,
ks2.rwm_expected_boys_pct, ks2.rwm_high_boys_pct, ks2.rwm_expected_girls_pct, ks2.rwm_high_girls_pct, ks2.rwm_expected_boys_pct, ks2.rwm_high_boys_pct, ks2.rwm_expected_girls_pct, ks2.rwm_high_girls_pct,
@@ -18,7 +18,8 @@ current_ks4 as (
english_maths_strong_pass_pct, english_maths_standard_pass_pct, english_maths_strong_pass_pct, english_maths_standard_pass_pct,
ebacc_entry_pct, ebacc_strong_pass_pct, ebacc_standard_pass_pct, ebacc_avg_score, ebacc_entry_pct, ebacc_strong_pass_pct, ebacc_standard_pass_pct, ebacc_avg_score,
gcse_grade_91_pct, gcse_grade_91_pct,
sen_pct, sen_support_pct, sen_ehcp_pct sen_pct, sen_support_pct, sen_ehcp_pct,
progress_8_banding, attainment_8_disadvantage_gap, progress_8_disadvantage_gap
from all_ks4 from all_ks4
), ),
@@ -34,7 +35,8 @@ predecessor_ks4 as (
ks4.english_maths_strong_pass_pct, ks4.english_maths_standard_pass_pct, ks4.english_maths_strong_pass_pct, ks4.english_maths_standard_pass_pct,
ks4.ebacc_entry_pct, ks4.ebacc_strong_pass_pct, ks4.ebacc_standard_pass_pct, ks4.ebacc_avg_score, ks4.ebacc_entry_pct, ks4.ebacc_strong_pass_pct, ks4.ebacc_standard_pass_pct, ks4.ebacc_avg_score,
ks4.gcse_grade_91_pct, ks4.gcse_grade_91_pct,
ks4.sen_pct, ks4.sen_support_pct, ks4.sen_ehcp_pct ks4.sen_pct, ks4.sen_support_pct, ks4.sen_ehcp_pct,
ks4.progress_8_banding, ks4.attainment_8_disadvantage_gap, ks4.progress_8_disadvantage_gap
from all_ks4 ks4 from all_ks4 ks4
inner join {{ ref('int_school_lineage') }} lin inner join {{ ref('int_school_lineage') }} lin
on ks4.urn = lin.predecessor_urn on ks4.urn = lin.predecessor_urn
@@ -8,9 +8,10 @@ models:
tests: [not_null, unique] tests: [not_null, unique]
- name: school_name - name: school_name
tests: [not_null] tests: [not_null]
- name: phase - name: phase_code
description: > description: >
Primary / Secondary / All-through etc. May be null for a small number GIAS PhaseOfEducation code (2 = Primary, 4 = Secondary, 7 = All-through,
etc. — see seeds/gias_code_names.csv). May be null for a small number
of independent schools where GIAS publishes "Not Applicable", no of independent schools where GIAS publishes "Not Applicable", no
statutory age range, and the school name gives no hint. statutory age range, and the school name gives no hint.
tests: tests:
@@ -27,10 +28,26 @@ models:
- not_null - not_null
- accepted_values: - accepted_values:
values: [true, false] values: [true, false]
- name: status - name: status_code
description: GIAS EstablishmentStatus code (1 = Open, 3 = Open but proposed to close)
tests: tests:
- accepted_values: - accepted_values:
values: ["Open"] values: [1, 3]
- name: school_type_code
tests:
- accepted_values:
severity: warn
values: [1, 2, 3, 5, 6, 7, 8, 10, 11, 12, 14, 15, 18, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 49, 56, 57]
- name: religious_character_code
tests:
- accepted_values:
severity: warn
values: [0, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 24, 25, 26, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 99]
- name: admissions_policy_code
tests:
- accepted_values:
severity: warn
values: [0, 2, 4, 9]
- name: dim_location - name: dim_location
description: School location dimension with PostGIS geometry description: School location dimension with PostGIS geometry
@@ -69,6 +86,13 @@ models:
tests: [not_null] tests: [not_null]
- name: year - name: year
tests: [not_null] tests: [not_null]
- name: reading_progress_lower_ci
- name: reading_progress_upper_ci
- name: writing_progress_lower_ci
- name: writing_progress_upper_ci
- name: writing_working_towards_pct
- name: maths_progress_lower_ci
- name: maths_progress_upper_ci
tests: tests:
- unique: - unique:
column_name: "urn || '-' || year" column_name: "urn || '-' || year"
@@ -80,6 +104,15 @@ models:
tests: [not_null] tests: [not_null]
- name: year - name: year
tests: [not_null] tests: [not_null]
- name: progress_8_banding
tests:
- accepted_values:
values: ['Well above average', 'Above average', 'Average', 'Below average', 'Well below average']
config:
where: "progress_8_banding is not null"
severity: warn
- name: attainment_8_disadvantage_gap
- name: progress_8_disadvantage_gap
tests: tests:
- unique: - unique:
column_name: "urn || '-' || year" column_name: "urn || '-' || year"
@@ -107,6 +140,11 @@ models:
tests: [not_null] tests: [not_null]
- name: year - name: year
tests: [not_null] tests: [not_null]
- name: second_preference_offers
- name: third_preference_offers
- name: cross_la_applications
- name: cross_la_offers
- name: total_offers
- name: fact_finance - name: fact_finance
description: School financial data — one row per URN per year description: School financial data — one row per URN per year
@@ -31,4 +31,5 @@ select
else null else null
end as longitude end as longitude
from {{ ref('stg_gias_establishments') }} s from {{ ref('stg_gias_establishments') }} s
where s.status = 'Open' -- Must match dim_school's status filter exactly (the API inner-joins the two).
where s.status_code in (1, 3)
+23 -24
View File
@@ -19,16 +19,17 @@ select
s.urn, s.urn,
s.local_authority_code * 1000 + s.establishment_number as laestab, s.local_authority_code * 1000 + s.establishment_number as laestab,
s.school_name, s.school_name,
-- Phase in GIAS code space (see seeds/gias_code_names.csv):
-- 2 = Primary, 4 = Secondary, 7 = All-through, 0 = Not applicable.
case case
-- 1. Trust GIAS phase when it's a real value (not the catch-all "Not Applicable") -- 1. Trust GIAS phase when it's a real value (0 = the catch-all "Not Applicable")
when s.phase is not null when s.phase_code is not null and s.phase_code != 0
and lower(trim(s.phase)) not in ('not applicable', '', 'unknown') then s.phase_code
then s.phase
-- 2. Infer from statutory age range (independent schools still publish these) -- 2. Infer from statutory age range (independent schools still publish these)
when s.statutory_high_age is not null and s.statutory_high_age <= 11 then 'Primary' when s.statutory_high_age is not null and s.statutory_high_age <= 11 then 2
when s.statutory_low_age is not null and s.statutory_low_age >= 11 then 'Secondary' when s.statutory_low_age is not null and s.statutory_low_age >= 11 then 4
when s.statutory_low_age is not null and s.statutory_high_age is not null when s.statutory_low_age is not null and s.statutory_high_age is not null
and s.statutory_low_age < 11 and s.statutory_high_age > 11 then 'All-through' and s.statutory_low_age < 11 and s.statutory_high_age > 11 then 7
-- 3. Fallback: infer from school name (covers independents with missing ages) -- 3. Fallback: infer from school name (covers independents with missing ages)
when s.school_name ilike '%primary%' when s.school_name ilike '%primary%'
or s.school_name ilike '%infant%' or s.school_name ilike '%infant%'
@@ -36,31 +37,27 @@ select
or s.school_name ilike '%preparatory%' or s.school_name ilike '%preparatory%'
or s.school_name ilike '% prep school%' or s.school_name ilike '% prep school%'
or s.school_name ilike '% prep %' or s.school_name ilike '% prep %'
then 'Primary' then 2
when s.school_name ilike '%secondary%' when s.school_name ilike '%secondary%'
or s.school_name ilike '%high school%' or s.school_name ilike '%high school%'
or s.school_name ilike '%grammar%' or s.school_name ilike '%grammar%'
or s.school_name ilike '%senior school%' or s.school_name ilike '%senior school%'
or s.school_name ilike '%upper school%' or s.school_name ilike '%upper school%'
then 'Secondary' then 4
-- 4. Give up — leave phase null so the UI renders no pill -- 4. Give up — null renders no phase pill
else null else null
end as phase, end as phase_code,
s.school_type, s.school_type_code,
s.academy_trust_name, s.academy_trust_name,
s.academy_trust_uid, s.academy_trust_uid,
s.religious_character, s.religious_character_code,
s.gender, s.gender,
s.statutory_low_age || '-' || s.statutory_high_age as age_range, s.statutory_low_age || '-' || s.statutory_high_age as age_range,
-- Authoritative sixth-form flag (spec §3): GIAS OfficialSixthForm. -- GIAS OfficialSixthForm in code space: 1 = has, 2 = does not, 0 = N/A.
-- "Not applicable" (nurseries, primaries, PRUs) => false. Blank GIAS -- Null (rare, new establishments) falls back to the statutory age range.
-- value (rare, new establishments) falls back to the statutory age range.
-- lower(trim()) guards against casing/whitespace variants in raw GIAS
-- data, same as the phase derivation above — an unmatched variant would
-- otherwise silently fall through to the age-range fallback.
case case
when lower(trim(s.official_sixth_form)) = 'has a sixth form' then true when s.official_sixth_form_code = 1 then true
when lower(trim(s.official_sixth_form)) in ('does not have a sixth form', 'not applicable') then false when s.official_sixth_form_code in (0, 2) then false
else coalesce(s.statutory_high_age >= 18, false) else coalesce(s.statutory_high_age >= 18, false)
end as has_sixth_form, end as has_sixth_form,
s.capacity, s.capacity,
@@ -70,9 +67,9 @@ select
s.telephone, s.telephone,
s.open_date, s.open_date,
s.close_date, s.close_date,
s.status, s.status_code,
s.nursery_provision, s.nursery_provision,
s.admissions_policy, s.admissions_policy_code,
-- Latest Ofsted (populated after monthly Ofsted pipeline runs) -- Latest Ofsted (populated after monthly Ofsted pipeline runs)
{% if ofsted_relation is not none %} {% if ofsted_relation is not none %}
@@ -91,4 +88,6 @@ from schools s
{% if ofsted_relation is not none %} {% if ofsted_relation is not none %}
left join {{ ref('int_ofsted_latest') }} o on s.urn = o.urn left join {{ ref('int_ofsted_latest') }} o on s.urn = o.urn
{% endif %} {% endif %}
where s.status = 'Open' -- 1 = Open; 3 = Open, but proposed to close (still operating; drops out when
-- GIAS flips to Closed — marts fully rebuild each run).
where s.status_code in (1, 3)
@@ -5,9 +5,14 @@ select
year, year,
school_phase, school_phase,
places_offered, places_offered,
total_offers,
total_applications, total_applications,
first_preference_applications, first_preference_applications,
first_preference_offers, first_preference_offers,
second_preference_offers,
third_preference_offers,
cross_la_applications,
cross_la_offers,
first_preference_offer_pct, first_preference_offer_pct,
oversubscription_ratio, oversubscription_ratio,
oversubscribed, oversubscribed,
@@ -15,13 +15,20 @@ select
reading_high_pct, reading_high_pct,
reading_avg_score, reading_avg_score,
reading_progress, reading_progress,
reading_progress_lower_ci,
reading_progress_upper_ci,
writing_expected_pct, writing_expected_pct,
writing_high_pct, writing_high_pct,
writing_progress, writing_progress,
writing_progress_lower_ci,
writing_progress_upper_ci,
writing_working_towards_pct,
maths_expected_pct, maths_expected_pct,
maths_high_pct, maths_high_pct,
maths_avg_score, maths_avg_score,
maths_progress, maths_progress,
maths_progress_lower_ci,
maths_progress_upper_ci,
gps_expected_pct, gps_expected_pct,
gps_high_pct, gps_high_pct,
gps_avg_score, gps_avg_score,
@@ -16,6 +16,9 @@ select
progress_8_score, progress_8_score,
progress_8_lower_ci, progress_8_lower_ci,
progress_8_upper_ci, progress_8_upper_ci,
progress_8_banding,
attainment_8_disadvantage_gap,
progress_8_disadvantage_gap,
progress_8_english, progress_8_english,
progress_8_maths, progress_8_maths,
progress_8_ebacc, progress_8_ebacc,
@@ -32,6 +32,11 @@ renamed as (
{{ safe_numeric('times_put_as_any_preferred_school') }}::integer as total_applications, {{ safe_numeric('times_put_as_any_preferred_school') }}::integer as total_applications,
{{ safe_numeric('times_put_as_1st_preference') }}::integer as first_preference_applications, {{ safe_numeric('times_put_as_1st_preference') }}::integer as first_preference_applications,
-- Cross-borough demand: applications naming this school from families
-- living in another local authority, and offers made to them.
{{ safe_numeric('"all_applications_from_another_LA"') }}::integer as cross_la_applications,
{{ safe_numeric('"offers_to_applicants_from_another_LA"') }}::integer as cross_la_offers,
-- Proportions -- Proportions
-- first_preference_offer_pct: of families who listed this school FIRST, -- first_preference_offer_pct: of families who listed this school FIRST,
-- the percentage that received an offer. 0100 scale. -- the percentage that received an offer. 0100 scale.
@@ -39,6 +39,12 @@ pivoted as (
max(case when subject = 'Reading' max(case when subject = 'Reading'
and breakdown_topic = 'All pupils' and breakdown = 'Total' and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_score') }} end) as reading_progress, then {{ safe_numeric('progress_measure_score') }} end) as reading_progress,
max(case when subject = 'Reading'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as reading_progress_lower_ci,
max(case when subject = 'Reading'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as reading_progress_upper_ci,
max(case when subject = 'Reading' max(case when subject = 'Reading'
and breakdown_topic = 'All pupils' and breakdown = 'Total' and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('absent_or_not_able_to_access_percent') }} end) as reading_absence_pct, then {{ safe_numeric('absent_or_not_able_to_access_percent') }} end) as reading_absence_pct,
@@ -53,6 +59,15 @@ pivoted as (
max(case when subject = 'Writing' max(case when subject = 'Writing'
and breakdown_topic = 'All pupils' and breakdown = 'Total' and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_score') }} end) as writing_progress, then {{ safe_numeric('progress_measure_score') }} end) as writing_progress,
max(case when subject = 'Writing'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as writing_progress_lower_ci,
max(case when subject = 'Writing'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as writing_progress_upper_ci,
max(case when subject = 'Writing'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('working_towards_expected_standard_pupil_percent') }} end) as writing_working_towards_pct,
max(case when subject = 'Writing' max(case when subject = 'Writing'
and breakdown_topic = 'All pupils' and breakdown = 'Total' and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('absent_or_not_able_to_access_percent') }} end) as writing_absence_pct, then {{ safe_numeric('absent_or_not_able_to_access_percent') }} end) as writing_absence_pct,
@@ -70,6 +85,12 @@ pivoted as (
max(case when subject = 'Maths' max(case when subject = 'Maths'
and breakdown_topic = 'All pupils' and breakdown = 'Total' and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_score') }} end) as maths_progress, then {{ safe_numeric('progress_measure_score') }} end) as maths_progress,
max(case when subject = 'Maths'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_lower_conf_interval') }} end) as maths_progress_lower_ci,
max(case when subject = 'Maths'
and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('progress_measure_upper_conf_interval') }} end) as maths_progress_upper_ci,
max(case when subject = 'Maths' max(case when subject = 'Maths'
and breakdown_topic = 'All pupils' and breakdown = 'Total' and breakdown_topic = 'All pupils' and breakdown = 'Total'
then {{ safe_numeric('absent_or_not_able_to_access_percent') }} end) as maths_absence_pct, then {{ safe_numeric('absent_or_not_able_to_access_percent') }} end) as maths_absence_pct,
@@ -143,13 +164,20 @@ select
p.reading_high_pct, p.reading_high_pct,
p.reading_avg_score, p.reading_avg_score,
p.reading_progress, p.reading_progress,
p.reading_progress_lower_ci,
p.reading_progress_upper_ci,
p.writing_expected_pct, p.writing_expected_pct,
p.writing_high_pct, p.writing_high_pct,
p.writing_progress, p.writing_progress,
p.writing_progress_lower_ci,
p.writing_progress_upper_ci,
p.writing_working_towards_pct,
p.maths_expected_pct, p.maths_expected_pct,
p.maths_high_pct, p.maths_high_pct,
p.maths_avg_score, p.maths_avg_score,
p.maths_progress, p.maths_progress,
p.maths_progress_lower_ci,
p.maths_progress_upper_ci,
p.gps_expected_pct, p.gps_expected_pct,
p.gps_high_pct, p.gps_high_pct,
p.gps_avg_score, p.gps_avg_score,
@@ -31,4 +31,10 @@ select
from {{ source('raw', 'ees_ks2_national') }} from {{ source('raw', 'ees_ks2_national') }}
where time_period ~ '^[0-9]+$' where time_period ~ '^[0-9]+$'
and cast(trim(time_period) as integer) >= 201617 -- 2015/16 was the first year of the current expected-standard tests, so it's
-- the correct floor (not 2016/17 -- that excluded a real, comparable national
-- row). GPS/science/scaled-score columns are already mapped correctly end to
-- end (tap.py's _KS2_NATIONAL_COL_MAP + this model select them fine); the
-- prod NULLs for those fields are stale raw.ees_ks2_national data from before
-- the map covered them, not a mapping bug -- no map change accompanies this fix.
and cast(trim(time_period) as integer) >= 201516
@@ -62,7 +62,16 @@ info as (
{{ safe_numeric('ks2_scaledscore_average') }} as prior_attainment_avg, {{ safe_numeric('ks2_scaledscore_average') }} as prior_attainment_avg,
{{ safe_numeric('sen_pupil_percent') }} as sen_pct, {{ safe_numeric('sen_pupil_percent') }} as sen_pct,
{{ safe_numeric('sen_with_ehcp_pupil_percent') }} as sen_ehcp_pct, {{ safe_numeric('sen_with_ehcp_pupil_percent') }} as sen_ehcp_pct,
{{ safe_numeric('sen_no_ehcp_pupil_percent') }} as sen_support_pct {{ safe_numeric('sen_no_ehcp_pupil_percent') }} as sen_support_pct,
-- EES suppression sentinels (z/c/x/q/u) and blanks must not reach the
-- mart as banding labels
case
when lower(trim(progress8_banding)) in ('', 'z', 'c', 'x', 'q', 'u', 'null')
then null
else trim(progress8_banding)
end as progress_8_banding,
{{ safe_numeric('attainment8_diffn') }} as attainment_8_disadvantage_gap,
{{ safe_numeric('progress8_diffn') }} as progress_8_disadvantage_gap
from {{ source('raw', 'ees_ks4_info') }} from {{ source('raw', 'ees_ks4_info') }}
where school_urn is not null where school_urn is not null
) )
@@ -102,7 +111,10 @@ select
-- Context -- Context
i.sen_pct, i.sen_pct,
i.sen_ehcp_pct, i.sen_ehcp_pct,
i.sen_support_pct i.sen_support_pct,
i.progress_8_banding,
i.attainment_8_disadvantage_gap,
i.progress_8_disadvantage_gap
from all_pupils p from all_pupils p
left join info i on p.urn = i.urn and p.year = i.year left join info i on p.urn = i.urn and p.year = i.year
@@ -12,12 +12,12 @@ renamed as (
"LA (name)" as local_authority_name, "LA (name)" as local_authority_name,
cast(nullif("EstablishmentNumber", '') as integer) as establishment_number, cast(nullif("EstablishmentNumber", '') as integer) as establishment_number,
"EstablishmentName" as school_name, "EstablishmentName" as school_name,
"TypeOfEstablishment (name)" as school_type, cast(nullif(trim("TypeOfEstablishment (code)"), '') as integer) as school_type_code,
"PhaseOfEducation (name)" as phase, cast(nullif(trim("PhaseOfEducation (code)"), '') as integer) as phase_code,
nullif(trim("OfficialSixthForm (name)"), '') as official_sixth_form, cast(nullif(trim("OfficialSixthForm (code)"), '') as integer) as official_sixth_form_code,
"Gender (name)" as gender, "Gender (name)" as gender,
"ReligiousCharacter (name)" as religious_character, cast(nullif(trim("ReligiousCharacter (code)"), '') as integer) as religious_character_code,
"AdmissionsPolicy (name)" as admissions_policy, cast(nullif(trim("AdmissionsPolicy (code)"), '') as integer) as admissions_policy_code,
"SchoolCapacity" as capacity, "SchoolCapacity" as capacity,
cast(nullif("NumberOfPupils", '') as integer) as total_pupils, cast(nullif("NumberOfPupils", '') as integer) as total_pupils,
"HeadTitle (name)" as head_title, "HeadTitle (name)" as head_title,
@@ -30,7 +30,7 @@ renamed as (
"Town" as town, "Town" as town,
"County (name)" as county, "County (name)" as county,
"Postcode" as postcode, "Postcode" as postcode,
"EstablishmentStatus (name)" as status, cast(nullif(trim("EstablishmentStatus (code)"), '') as integer) as status_code,
case when "OpenDate" = '' then null else to_date("OpenDate", 'DD-MM-YYYY') end as open_date, case when "OpenDate" = '' then null else to_date("OpenDate", 'DD-MM-YYYY') end as open_date,
case when "CloseDate" = '' then null else to_date("CloseDate", 'DD-MM-YYYY') end as close_date, case when "CloseDate" = '' then null else to_date("CloseDate", 'DD-MM-YYYY') end as close_date,
"Trusts (name)" as academy_trust_name, "Trusts (name)" as academy_trust_name,
@@ -17,13 +17,23 @@ select
{{ safe_numeric('reading_high_pct') }} as reading_high_pct, {{ safe_numeric('reading_high_pct') }} as reading_high_pct,
{{ safe_numeric('reading_avg_score') }} as reading_avg_score, {{ safe_numeric('reading_avg_score') }} as reading_avg_score,
{{ safe_numeric('reading_progress') }} as reading_progress, {{ safe_numeric('reading_progress') }} as reading_progress,
-- Progress CIs / working-towards: not published in the legacy CSVs.
-- Typed placeholders keep positional alignment with stg_ees_ks2 in
-- int_ks2_with_lineage's UNION ALL.
null::numeric as reading_progress_lower_ci,
null::numeric as reading_progress_upper_ci,
{{ safe_numeric('writing_expected_pct') }} as writing_expected_pct, {{ safe_numeric('writing_expected_pct') }} as writing_expected_pct,
{{ safe_numeric('writing_high_pct') }} as writing_high_pct, {{ safe_numeric('writing_high_pct') }} as writing_high_pct,
{{ safe_numeric('writing_progress') }} as writing_progress, {{ safe_numeric('writing_progress') }} as writing_progress,
null::numeric as writing_progress_lower_ci,
null::numeric as writing_progress_upper_ci,
null::numeric as writing_working_towards_pct,
{{ safe_numeric('maths_expected_pct') }} as maths_expected_pct, {{ safe_numeric('maths_expected_pct') }} as maths_expected_pct,
{{ safe_numeric('maths_high_pct') }} as maths_high_pct, {{ safe_numeric('maths_high_pct') }} as maths_high_pct,
{{ safe_numeric('maths_avg_score') }} as maths_avg_score, {{ safe_numeric('maths_avg_score') }} as maths_avg_score,
{{ safe_numeric('maths_progress') }} as maths_progress, {{ safe_numeric('maths_progress') }} as maths_progress,
null::numeric as maths_progress_lower_ci,
null::numeric as maths_progress_upper_ci,
{{ safe_numeric('gps_expected_pct') }} as gps_expected_pct, {{ safe_numeric('gps_expected_pct') }} as gps_expected_pct,
{{ safe_numeric('gps_high_pct') }} as gps_high_pct, {{ safe_numeric('gps_high_pct') }} as gps_high_pct,
{{ safe_numeric('gps_avg_score') }} as gps_avg_score, {{ safe_numeric('gps_avg_score') }} as gps_avg_score,
@@ -41,8 +41,13 @@ select
-- SEN -- SEN
null::numeric as sen_pct, null::numeric as sen_pct,
{{ safe_numeric('sen_ehcp_pct') }} as sen_ehcp_pct,
{{ safe_numeric('sen_support_pct') }} as sen_support_pct, {{ safe_numeric('sen_support_pct') }} as sen_support_pct,
{{ safe_numeric('sen_ehcp_pct') }} as sen_ehcp_pct
-- Progress 8 banding & disadvantage gaps (not published in legacy format)
null::text as progress_8_banding,
null::numeric as attainment_8_disadvantage_gap,
null::numeric as progress_8_disadvantage_gap
from {{ source('raw', 'legacy_ks4') }} from {{ source('raw', 'legacy_ks4') }}
where urn is not null where urn is not null
@@ -33,17 +33,23 @@ renamed as (
nullif(trim(ungraded_outcome), 'NULL') as ungraded_outcome, nullif(trim(ungraded_outcome), 'NULL') as ungraded_outcome,
{{ parse_ungraded_outcome('ungraded_outcome') }}::integer as ungraded_grade, {{ parse_ungraded_outcome('ungraded_outcome') }}::integer as ungraded_grade,
-- Report Card fields (post-Nov 2025 framework) -- Report Card fields (post-Nov 2025 framework), 5-point scale:
-- TODO: add rc_* columns to tap-uk-ofsted schema once CSV column names are confirmed -- 1 Exceptional · 2 Strong standard · 3 Expected standard
null::text as rc_safeguarding_met, -- · 4 Needs attention · 5 Urgent improvement
null::text as rc_inclusion, case lower(trim(nullif(rc_safeguarding_met, 'NULL')))
null::text as rc_curriculum_teaching, when 'met' then true
null::text as rc_achievement, when 'not met' then false
null::text as rc_attendance_behaviour, end as rc_safeguarding_met,
null::text as rc_personal_development, {{ parse_report_card_grade('rc_inclusion') }}::integer as rc_inclusion,
null::text as rc_leadership_governance, {{ parse_report_card_grade('rc_curriculum_teaching') }}::integer as rc_curriculum_teaching,
null::text as rc_early_years, {{ parse_report_card_grade('rc_achievement') }}::integer as rc_achievement,
null::text as rc_sixth_form, {{ parse_report_card_grade('rc_attendance_behaviour') }}::integer as rc_attendance_behaviour,
{{ parse_report_card_grade('rc_personal_development') }}::integer as rc_personal_development,
{{ parse_report_card_grade('rc_leadership_governance') }}::integer as rc_leadership_governance,
-- No MI column exists for these yet (see tap.py); the tap never
-- emits rc_early_years/rc_sixth_form, so these stay NULL.
null::integer as rc_early_years,
null::integer as rc_sixth_form,
report_url report_url
from source from source
@@ -0,0 +1,108 @@
field,code,name
school_type,1,Community school
school_type,2,Voluntary aided school
school_type,3,Voluntary controlled school
school_type,5,Foundation school
school_type,6,City technology college
school_type,7,Community special school
school_type,8,Non-maintained special school
school_type,10,Other independent special school
school_type,11,Other independent school
school_type,12,Foundation special school
school_type,14,Pupil referral unit
school_type,15,Local authority nursery school
school_type,18,Further education
school_type,24,Secure units
school_type,25,Offshore schools
school_type,26,Service children's education
school_type,27,Miscellaneous
school_type,28,Academy sponsor led
school_type,29,Higher education institutions
school_type,30,Welsh establishment
school_type,31,Sixth form centres
school_type,32,Special post 16 institution
school_type,33,Academy special sponsor led
school_type,34,Academy converter
school_type,35,Free schools
school_type,36,Free schools special
school_type,37,British schools overseas
school_type,38,Free schools alternative provision
school_type,39,Free schools 16 to 19
school_type,40,University technical college
school_type,41,Studio schools
school_type,42,Academy alternative provision converter
school_type,43,Academy alternative provision sponsor led
school_type,44,Academy special converter
school_type,45,Academy 16-19 converter
school_type,46,Academy 16 to 19 sponsor led
school_type,49,Online provider
school_type,56,Institution funded by other government department
school_type,57,Academy secure 16 to 19
establishment_status,1,Open
establishment_status,2,Closed
establishment_status,3,"Open, but proposed to close"
establishment_status,4,Proposed to open
phase_of_education,0,Not applicable
phase_of_education,1,Nursery
phase_of_education,2,Primary
phase_of_education,3,Middle deemed primary
phase_of_education,4,Secondary
phase_of_education,5,Middle deemed secondary
phase_of_education,6,16 plus
phase_of_education,7,All-through
official_sixth_form,0,Not applicable
official_sixth_form,1,Has a sixth form
official_sixth_form,2,Does not have a sixth form
official_sixth_form,9,
religious_character,0,Does not apply
religious_character,2,Church of England
religious_character,3,Roman Catholic
religious_character,4,Methodist
religious_character,5,Jewish
religious_character,6,None
religious_character,7,Muslim
religious_character,8,Seventh Day Adventist
religious_character,9,Church of England/Methodist
religious_character,10,Methodist/Church of England
religious_character,11,Church of England/Roman Catholic
religious_character,12,Church of England/United Reformed Church
religious_character,13,Roman Catholic/Church of England
religious_character,14,Quaker
religious_character,15,Christian
religious_character,16,United Reformed Church
religious_character,17,Congregational Church
religious_character,18,Free Church
religious_character,19,Church of England/Free Church
religious_character,20,Church of England/Christian
religious_character,21,Sikh
religious_character,22,Greek Orthodox
religious_character,24,Buddhist
religious_character,25,Hindu
religious_character,26,Moravian
religious_character,28,Inter- / non- denominational
religious_character,29,Multi-faith
religious_character,30,Church of England/Methodist/United Reform Church/Baptist
religious_character,31,Anglican
religious_character,32,Anglican/Christian
religious_character,33,Anglican/Evangelical
religious_character,34,Anglican/Church of England
religious_character,35,Catholic
religious_character,36,Charadi Jewish
religious_character,37,Christian/Evangelical
religious_character,38,Christian Science
religious_character,39,Christian/Methodist
religious_character,40,Christian/non-denominational
religious_character,41,Church of England/Evangelical
religious_character,42,Islam
religious_character,43,Orthodox Jewish
religious_character,44,Plymouth Brethren Christian Church
religious_character,45,Protestant
religious_character,46,Protestant/Evangelical
religious_character,47,Reformed Baptist
religious_character,48,Roman Catholic/Anglican
religious_character,49,Sunni Deobandi
religious_character,99,
admissions_policy,0,Not applicable
admissions_policy,2,Selective
admissions_policy,4,Non-selective
admissions_policy,9,
1 field code name
2 school_type 1 Community school
3 school_type 2 Voluntary aided school
4 school_type 3 Voluntary controlled school
5 school_type 5 Foundation school
6 school_type 6 City technology college
7 school_type 7 Community special school
8 school_type 8 Non-maintained special school
9 school_type 10 Other independent special school
10 school_type 11 Other independent school
11 school_type 12 Foundation special school
12 school_type 14 Pupil referral unit
13 school_type 15 Local authority nursery school
14 school_type 18 Further education
15 school_type 24 Secure units
16 school_type 25 Offshore schools
17 school_type 26 Service children's education
18 school_type 27 Miscellaneous
19 school_type 28 Academy sponsor led
20 school_type 29 Higher education institutions
21 school_type 30 Welsh establishment
22 school_type 31 Sixth form centres
23 school_type 32 Special post 16 institution
24 school_type 33 Academy special sponsor led
25 school_type 34 Academy converter
26 school_type 35 Free schools
27 school_type 36 Free schools special
28 school_type 37 British schools overseas
29 school_type 38 Free schools alternative provision
30 school_type 39 Free schools 16 to 19
31 school_type 40 University technical college
32 school_type 41 Studio schools
33 school_type 42 Academy alternative provision converter
34 school_type 43 Academy alternative provision sponsor led
35 school_type 44 Academy special converter
36 school_type 45 Academy 16-19 converter
37 school_type 46 Academy 16 to 19 sponsor led
38 school_type 49 Online provider
39 school_type 56 Institution funded by other government department
40 school_type 57 Academy secure 16 to 19
41 establishment_status 1 Open
42 establishment_status 2 Closed
43 establishment_status 3 Open, but proposed to close
44 establishment_status 4 Proposed to open
45 phase_of_education 0 Not applicable
46 phase_of_education 1 Nursery
47 phase_of_education 2 Primary
48 phase_of_education 3 Middle deemed primary
49 phase_of_education 4 Secondary
50 phase_of_education 5 Middle deemed secondary
51 phase_of_education 6 16 plus
52 phase_of_education 7 All-through
53 official_sixth_form 0 Not applicable
54 official_sixth_form 1 Has a sixth form
55 official_sixth_form 2 Does not have a sixth form
56 official_sixth_form 9
57 religious_character 0 Does not apply
58 religious_character 2 Church of England
59 religious_character 3 Roman Catholic
60 religious_character 4 Methodist
61 religious_character 5 Jewish
62 religious_character 6 None
63 religious_character 7 Muslim
64 religious_character 8 Seventh Day Adventist
65 religious_character 9 Church of England/Methodist
66 religious_character 10 Methodist/Church of England
67 religious_character 11 Church of England/Roman Catholic
68 religious_character 12 Church of England/United Reformed Church
69 religious_character 13 Roman Catholic/Church of England
70 religious_character 14 Quaker
71 religious_character 15 Christian
72 religious_character 16 United Reformed Church
73 religious_character 17 Congregational Church
74 religious_character 18 Free Church
75 religious_character 19 Church of England/Free Church
76 religious_character 20 Church of England/Christian
77 religious_character 21 Sikh
78 religious_character 22 Greek Orthodox
79 religious_character 24 Buddhist
80 religious_character 25 Hindu
81 religious_character 26 Moravian
82 religious_character 28 Inter- / non- denominational
83 religious_character 29 Multi-faith
84 religious_character 30 Church of England/Methodist/United Reform Church/Baptist
85 religious_character 31 Anglican
86 religious_character 32 Anglican/Christian
87 religious_character 33 Anglican/Evangelical
88 religious_character 34 Anglican/Church of England
89 religious_character 35 Catholic
90 religious_character 36 Charadi Jewish
91 religious_character 37 Christian/Evangelical
92 religious_character 38 Christian Science
93 religious_character 39 Christian/Methodist
94 religious_character 40 Christian/non-denominational
95 religious_character 41 Church of England/Evangelical
96 religious_character 42 Islam
97 religious_character 43 Orthodox Jewish
98 religious_character 44 Plymouth Brethren Christian Church
99 religious_character 45 Protestant
100 religious_character 46 Protestant/Evangelical
101 religious_character 47 Reformed Baptist
102 religious_character 48 Roman Catholic/Anglican
103 religious_character 49 Sunni Deobandi
104 religious_character 99
105 admissions_policy 0 Not applicable
106 admissions_policy 2 Selective
107 admissions_policy 4 Non-selective
108 admissions_policy 9
@@ -0,0 +1,33 @@
-- Warn when the live GIAS CSV carries a (code, name) pair we don't have in
-- the dictionary seed — i.e. DfE added or renamed a value. Fix by rerunning
-- pipeline/scripts/generate_gias_codes.py and committing the regenerated
-- dictionaries + seed together.
{{ config(severity='warn') }}
with raw_pairs as (
{% for field_key, code_col, name_col in [
('school_type', 'TypeOfEstablishment (code)', 'TypeOfEstablishment (name)'),
('establishment_status', 'EstablishmentStatus (code)', 'EstablishmentStatus (name)'),
('phase_of_education', 'PhaseOfEducation (code)', 'PhaseOfEducation (name)'),
('official_sixth_form', 'OfficialSixthForm (code)', 'OfficialSixthForm (name)'),
('religious_character', 'ReligiousCharacter (code)', 'ReligiousCharacter (name)'),
('admissions_policy', 'AdmissionsPolicy (code)', 'AdmissionsPolicy (name)')
] %}
select distinct
'{{ field_key }}' as field,
cast(nullif(trim("{{ code_col }}"), '') as integer) as code,
nullif(trim("{{ name_col }}"), '') as name
from {{ source('raw', 'gias_establishments') }}
where nullif(trim("{{ code_col }}"), '') is not null
and nullif(trim("{{ name_col }}"), '') is not null
{% if not loop.last %}union all{% endif %}
{% endfor %}
)
select r.*
from raw_pairs r
left join {{ ref('gias_code_names') }} s
on s.field = r.field
and s.code = r.code
and s.name = r.name
where s.field is null