Beechwood College (URN 142458, Sully, CF64 5SE) survived the England-only
filter. GIAS types it a Special post 16 institution (32), not a Welsh
establishment (30), so filtering on establishment type alone left it behind —
the last Welsh school on the site, and the reason Vale of Glamorgan was still
in the authority list.
The earlier verification claimed the type codes mapped onto the Welsh
authorities in both directions. That was checked exhaustively for Cardiff and
by count for three others; Vale of Glamorgan was never checked, and it was the
one that did not hold.
LA code is the reliable discriminator: GIAS gives the 22 Welsh unitary
authorities the contiguous block 660-681, which English authorities never use.
Filtering on postcode would have been wrong — Redbrook, Tutshill, Wyedean and
two other Gloucestershire schools carry NP16/NP25 postcodes because Royal Mail
areas straddle the border, and they are English schools with English data. An
e2e test now pins those five so the fix cannot be simplified into a postcode
filter later.
The 660-681 range is documented GIAS structure this project cannot verify from
its own data, so the range filters and the authority NAME checks: if the range
is ever wrong, a Welsh authority reappears in assert_england_only_schools and
the pipeline fails loudly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
GIAS ships the whole UK plus overseas and offshore establishments. None of
them carry comparable DfE performance data — Wales does not publish on the
English measures at all — so every one of these pages rendered with null
results, null Ofsted and null phase. There were 2,036 of them: 1,569 Welsh,
123 offshore (Jersey, Guernsey, Isle of Man, Gibraltar), 316 British schools
overseas and 28 service children's schools. All 2,036 were being submitted to
search engines, alongside 29 local authorities that existed in the filters
purely to list them.
Filter at the mart boundary rather than the view layer. dim_school and
dim_location both exclude TypeOfEstablishment in {25, 26, 30, 37}, listed once
as vars.non_england_school_type_codes. Everything downstream reads those two
marts — search, the school page, /api/filters, rankings, Typesense and
build_sitemap() — so one filter removes them from the site and the sitemap
together, and Typesense drops them on its next rebuild since it recreates the
collection and swaps the alias rather than upserting in place.
coalesce rather than a bare NOT IN: a null type code would make the predicate
null and drop the row silently, and an unknown type is not grounds for
exclusion. No establishment has a null type today, but a future GIAS refresh
could ship one and the loss would be invisible.
assert_england_only_schools guards both directions: no excluded type survives
in dim_school, and dim_location holds no URN dim_school lacks — the API
inner-joins them, so the two filters drifting apart would silently shrink the
corpus.
Corpus goes from 27,229 schools to 25,193, and the authority list from 182 to
153. The 1,569 Welsh URLs now 404.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Adds the cut-off distance a parent actually asks about — "how close do we
need to live?" — end to end: a Singer tap, dbt staging and mart models, an
Airflow DAG, and a tile on both detail templates. 3,597 schools across 57
local authorities carry a figure; the rest are unchanged.
There is no national source for this. Each LA publishes its own cut-offs in
its own format, and the collected CSV is transcribed from PDFs, spreadsheets
and web pages — so most of the work here is deciding what is safe to show.
Data
* tap-uk-school-distance loads the CSV verbatim into raw. Keyed on
(urn, year, school_name), because school_name carries the admission
route: (urn, year) alone collides on 118 keys and a reload would have
silently dropped every band but one.
* stg_school_distance applies a 25 m – 25 km plausibility band. The source
contains 0.0-mile rows (published where a school filled on a higher
criterion), 1-metre cut-offs, and one reading 533 miles — ~4% of rows,
all of which would put a visibly wrong number on a live page.
* fact_admission_distance collapses routes to one row per school per year
using the furthest, and keeps route_count so the page can say the figure
is the widest of several bands rather than the one for a given child.
Serving
* Kept out of fact_admissions: that mart is EES-derived and near-complete
for England, this one covers 57 LAs, and the two refresh independently.
* Latest year only. Coverage is ragged — a school may have 2021 and 2026
and nothing between — so a history array would invite a trend line drawn
through gaps that are absences of publication, not of a cut-off.
* The Admissions section now renders on either source. 3% of the schools
that render have a cut-off and no EES admissions row, and gating on
admissions alone would have hidden the figure on those pages.
Interface
* The year travels with the figure everywhere it appears; a cut-off
detached from its admissions round is not a fact about anything.
* "Not a fixed catchment — it moves every year" sits under every instance,
because that is the inference a parent will otherwise draw.
* Replaces a hardcoded "Historical distance cut-off data is not available
for this school" that appeared on every secondary page, including the
ones whose council does publish it. The absence is now stated only when
it is real, and names the authority that would hold it.
The tint costs the muted tokens their AA margin: measured on the composited
backdrop (not the computed one, which reports the untinted card), --text-muted
falls to 4.09:1 in dark theme. The tile uses --text-secondary instead — 6.50:1
dark, 6.60:1 light.
The DAG is manual, like the other annual ones: councils publish on allocation
day, each on its own timetable, so there is no date worth scheduling against.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
The DAG's dbt --select list predates the official-KS4-nationals stream,
so the extract loaded raw.ees_ks4_national but the staging model and
fact_ks4_national_averages were never rebuilt — staging kept serving the
old computed means after the DAG run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
Swept in accidentally by a broad 'git add pipeline'. They embed local
absolute paths and a personal usage-tracking UUID, and a stale committed
manifest causes partial-parse/version-mismatch noise for others.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
New ees_ks4_national stream ingests the EES 'National characteristics
summary data' series (England, state-funded, all pupils). The old mart's
unweighted school means were 7-15 points off every headline measure and
produced an impossible national Progress 8 (-0.27). The API's computed
fallback is gone too: the footnote calls these figures official, so an
unbuilt mart now yields an empty series, never a stand-in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
The FSM chip anchored against disadvantaged_pct (a different measure,
FSM6+CLA) whenever fsm_pct was null — which it always was, since the
performance df has no fsm_pct. New fact_census_benchmarks mart supplies
pupil-weighted FSM/EAL means per phase; the KS2-column medians that
produced a bogus 50% 'secondary disadvantaged' anchor are gone.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
The MI file's report-card grade columns belong to the latest FULL
inspection (col 'Inspection start date'), but inspection_date maps to the
legacy OEIF graded/ungraded dates — so report cards were being dated with
pre-Nov-2025 inspections. Also discover_csv_url() returned matches[0],
the oldest (2017) link on the GOV.UK page.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
fact_ks4_national_averages is computed once at dbt build time (covered by
the EES DAG's stg_ees_ks4+ selector). _national_averages_payload now reads
both national-averages marts instead of scanning the performance dataframe
per year on every /api/compare request (~250ms saved per call). Fallback
for the deploy-before-DAG window computes the latest year only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
The AI review gate caught that stg_ees_ks2's 7 new columns broke the
positional UNION ALL with stg_legacy_ks2 in int_ks2_with_lineage, and
that the lineage CTEs never emitted them (same class of bug fixed for
KS4 in 34a5de2). Legacy gets typed null placeholders at matching
positions; both lineage CTEs pass the columns through. 45/45 columns
verified name-identical in order across both union branches.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
Evidence trail for the rc_* mapping in the prior commit: real value_counts()
over the 7 MI report-card columns, confirming the 5-value grade vocabulary
and that 'Achievement'/'Safeguarding standards' match by exact string only
(no legacy OEIF column accidentally consumed).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
Wires the tap TODO in stg_ofsted_inspections.sql: maps the 7 confirmed
report-card MI columns (Safeguarding standards, Inclusion, Curriculum
and teaching, Achievement, Attendance and behaviour, Personal
development and wellbeing, Leadership and governance) into rc_*
fields, parsed via the new parse_report_card_grade macro against
real sampled grade values (Exceptional/Strong standard/Expected
standard/Needs attention/Urgent improvement). rc_safeguarding_met
becomes boolean from Met/Not met. rc_early_years/rc_sixth_form have
no MI column yet and are intentionally omitted from COLUMN_PRIORITY,
staying NULL.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
DfE never published school-level KS2 2021/22 data publicly (confirmed via
EES release notes and by walking the Compare School Performance download
wizard, which has no ks2 checkbox for 2021-2022, same as the COVID-cancelled
2020-2021 year). No archive exists to verify column headers against or
upload to the filebrowser; Task 6 is blocked at the source-data level.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
Widen the year filter in stg_ees_ks2_national.sql from >= 201617 to
>= 201516 so the England national-averages line no longer starts a
year late; the catalogue CSV has a real, comparable 201516 row (2015/16
was the first year of the current expected-standard tests, so it's the
correct floor).
GPS/science/scaled-score national columns confirmed present at source
with correct mapping; prod NULLs are stale raw data, backfilled by the
next extract run. No _KS2_NATIONAL_COL_MAP change accompanies this fix.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
int_ks4_with_lineage.sql unions stg_ees_ks4 and stg_legacy_ks4 via
`select *`, which PostgreSQL aligns positionally. stg_legacy_ks4 listed
sen_support_pct before sen_ehcp_pct while stg_ees_ks4 lists sen_ehcp_pct
before sen_support_pct, swapping the two values for legacy-sourced rows
in marts.fact_ks4_performance. Reordered stg_legacy_ks4's final select
to match.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
ReligiousCharacter 99 (~4k schools) and AdmissionsPolicy 9 (~5.6k) carry a
code with a blank name in the GIAS CSV; the generator skipped them so they
hit the Unknown(<code>) path — wrongly triggering the Faith-priority tag
and polluting filters. Blank-only codes now map to "" (byte-identical to
the old name pipeline); accepted_values lists extended to match the seed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Addresses AI-review findings: the annual IDACI DAG also rebuilds a mart
(fact_deprivation) and needs the reload; curl gets connect/max timeouts
so an unreachable backend fails fast instead of hanging the task.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The daily/monthly/annual DAG docstring promised an Invalidate Cache step
that never existed — after a marts rebuild the backend kept serving its
startup-cached (possibly empty) DataFrame until a container restart.
Add a POST /api/admin/reload task at the end of each pipeline DAG,
mirroring the sitemap DAG's admin-call pattern.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
These schools are still operating and publish results; they drop out
automatically when GIAS flips them to Closed since marts fully rebuild
each run.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Matches the phase derivation's guard against casing/whitespace variants in
raw GIAS data; an unmatched variant previously fell through silently to the
statutory-age fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Removes the 'What Parents Say' section and all supporting elements:
Frontend:
- Drop the OfstedParentView type, the parent_view field, the survey
section and the 'X% would recommend' callouts in the primary and
secondary detail views, the Parents nav item, and the parent-view CSS.
Backend:
- Remove the FactParentView model, its loading in data_loader, and
parent_view from the school-details API response.
- Bump SCHEMA_VERSION to 6 and add an idempotent drop step
(DROP TABLE IF EXISTS marts.fact_parent_view) to the CLI migration;
add scripts/sql/drop_fact_parent_view.sql to apply directly to the
dbt-owned marts DBs on staging and prod.
Pipeline:
- Delete the stg_parent_view + fact_parent_view dbt models and their
source/schema entries, the tap-uk-parent-view Meltano extractor, and
the monthly Parent View DAG; drop it from the Dockerfile and the
staging bootstrap docs.
The rest of dbt (which builds every mart the app reads) is untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
int_ofsted_latest is only ref()'d inside a conditional block, so dbt
couldn't infer the edge and failed to compile dim_school. Add the
-- depends_on hint dbt recommends. No runtime behaviour change: the
adapter.get_relation guard still handles the pre-Ofsted-pipeline case.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Ungraded (Section 8) inspections don't assign a fresh grade — the export
only gives free text like "School remains Good". Parse that text into a
grade (remains Outstanding -> 1, remains Good -> 2, else null) and use it
as a last-resort fallback when no graded overall effectiveness exists.
Also retain schools that have only an ungraded inspection (no graded date)
by coalescing the inspection date, so ~8.5k previously-dropped schools now
carry a grade.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The standalone dbt Fusion binary (dbt-core 2.x) on PATH shadows the
pip-installed classic dbt-postgres ~=1.10 and rejects the Postgres
adapter (dbt1005), breaking every DAG's dbt_build task. Invoke dbt via
`python -m dbt.cli.main` in the DAGs and the Dockerfile dbt deps step so
the classic Postgres-capable engine is always used regardless of PATH.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
int_ks4_with_lineage references stg_legacy_ks4 but the model was never
selected for build, causing a missing relation error.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
LegacyKS2Stream now auto-detects ZIP vs bare CSV — if the download is a ZIP
it extracts england_ks2final.csv; if it's a plain CSV file it reads directly.
This keeps backwards compatibility while allowing both streams to share the
same DfE annual archive URLs.
legacy_ks2_urls updated to point at the same 4 ZIPs as legacy_ks4_urls so
only one set of archives needs to be maintained going forward.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Mirrors the existing legacy KS2 pattern to fill the gap before EES hosted
KS4 data. Four files changed:
- tap-uk-ees: LegacyKS4Stream downloads each year's DfE Compare School
Performance ZIP, extracts england_ks4final.csv, maps 416 legacy columns
to Singer fields, strips % suffixes. Registered in discover_streams().
TapUKEES.config_jsonschema gains legacy_ks4_urls setting.
- stg_legacy_ks4.sql: safe_numeric casts + NULL placeholders for columns
not present in legacy format (ebacc_avg_score, gcse_grade_91_pct,
prior_attainment_avg, sen_pct).
- int_ks4_with_lineage.sql: adds all_ks4 CTE unioning stg_ees_ks4 and
stg_legacy_ks4, matching the int_ks2_with_lineage pattern.
- _stg_sources.yml + meltano.yml: source declaration and setting definition
for legacy_ks4. URLs configured per-year once provided.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Two bugs prevented historical secondary school data from loading:
1. stg_ees_ks4.sql filtered breakdown_topic = 'Total' only, but EES
releases prior to 2023/24 use breakdown_topic = 'All pupils' (matching
the KS2 convention). All older years were silently dropped to zero rows.
Fix: accept both values with an IN clause.
2. get_all_releases() in tap-uk-ees fetched only the first page of the
EES releases API. Now follows all pages via the paging.totalPages field
so no historical release is missed when more than 20 exist.
After re-running the annual EES pipeline, secondary school comparison
charts should show data across all available years (2018/19 onwards).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The staging model aliased EES's total_number_places_offered column as
published_admission_number, but PAN is the school's published capacity
(not exposed by EES at school level) — what we actually have is the
count of places offered in a given admissions round. The misnomer
propagated to the mart, SQLAlchemy model, API response, TS types, and
UI copy ("places per year", "(PAN)").
Rename end-to-end and fix the UI labels:
- "29 places for 42 first-choice applications"
→ "29 places offered for 42 first-choice applications"
- "Reception/Year 7 places per year"
→ "Reception/Year 7 places offered"
- drop the misleading "(PAN)" suffix in the secondary view
Also add a comment in stg_ees_admissions clarifying this is the number
of places offered, not PAN. Requires dbt to rebuild fact_admissions
(marts are materialized as tables) before the backend can start.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
dim_school.sql was checking for int_ofsted_latest in target.schema (wrong schema)
due to the custom generate_schema_name macro using literal schema names. The
model lives in 'intermediate', so ofsted_grade/date/framework were always NULL
in dim_school, causing all list cards to show 'Not yet inspected'.
Fix 1: data_loader.py joins marts.fact_ofsted_inspection with DISTINCT ON to
get latest inspection per school — no pipeline re-run needed.
Fix 2: dim_school.sql uses schema='intermediate' so future dbt runs correctly
denormalise the Ofsted summary into dim_school.
meltano run does not support --select; the full tap-uk-ees run already
includes EESKs2NationalStream so no separate task is needed.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>