Sitemap regeneration failed on staging. 'richmond' in the curated locality
list collides with the GIAS town Richmond in North Yorkshire (37 schools), the
registry raised, and the admin endpoint 500d — taking down sitemap generation
for all 25,000 school pages over one bad row of curated data.
The guard now skips the colliding locality and logs an error. Skipping still
achieves what the guard was for — a locality never silently shadows a town —
without letting curated data break the site. That matters beyond this bug:
GIAS town names change with no code change here, so a raise could fire
spontaneously in production later.
Also removes four localities that were London boroughs rather than districts.
Hackney, Islington, Greenwich and Ealing are local authorities with 104, 72,
108 and 115 schools and already have authority pages; a locality defined by
two or three outcodes would have been a partial near-duplicate of one — the
thin-content failure the two-namespace design exists to avoid. A test now
guards the whole borough list.
Validated against the live corpus: 15 localities, no town collisions, no
authority duplicates, all 15 clear the threshold.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Separate children per family so Search Console reports the location layer's
indexation apart from the school pages' — which is the point of the index
built in W1, and the number the stop condition watches.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The registry is cached for the process and reset by the same admin endpoint
that rebuilds the sitemaps, so places and sitemap always describe the same
corpus rather than drifting apart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The GIAS town field puts 1,819 London schools under the single value
'London', so it cannot answer 'schools in Battersea' — a query that appears in
the baseline. No single field can: parliamentary constituency gives Battersea
but not Canary Wharf, admin_ward gives Canary Wharf but not Battersea, and
neither gives Clapham or Shoreditch. So a locality is curated, defined by the
postcode districts it covers, which needs no new ingestion.
A locality may not shadow a published town: the registry raises rather than
silently costing a page that carries real demand. One below the threshold is
logged rather than raising, because a locality can legitimately be too small.
The pipeline seed mirrors the module, with a test guarding the drift — the
same arrangement gias_codes has, and for the same reason: the backend image
does not contain pipeline/.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Two namespaces because 67 town names collide with an authority name and
neither set contains the other — postal towns cross authority boundaries, so
Bedford the town holds 104 schools against the authority's 86.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
_school_sitemap_rows tested only the latest year's row, which quietly dropped
every school with results in its history but a null row for the most recent
year — a school that stopped reporting, or whose figures were suppressed for
small-cohort disclosure.
The Mallard Academy (150367) is the case that caught it: real KS2 results for
2015-16 through 2018-19, then null rows from 2022-23 on. Its detail page shows
all four years; the sitemap omitted it. Sampling 40 of the 2,206 excluded
schools found 4 like this, so roughly 220 real pages were being withheld.
Publishable is now a property of the school, computed across every row, while
lastmod still comes from the latest row so the most recent Ofsted date wins.
The field list is a module constant shared with _has_publishable_data so the
two checks cannot drift.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Search Console reports coverage per submitted sitemap, so one file per page
family is what will make W2's location pages measurable when they land. The
index's lastmod is generation time, which is the correct semantic there —
unlike on a <url>, where it would be a claim we cannot support.
Children sit under /sitemaps/ because Next only treats a whole bracketed path
segment as dynamic; a route folder named sitemap-[...parts] would be read as a
literal static segment and never match. Confirmed by the build output, which
lists /sitemaps/[...parts] as a dynamic route.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Drops the schools with neither results nor an Ofsted grade, adds /admissions
which was never listed, replaces the invented priority and changefreq with a
lastmod taken from each school's Ofsted date.
lastmod is omitted where no date is known rather than defaulted to now. An
always-now lastmod is a claim Google learns to distrust; absent honestly
means unknown.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The apex 301s to www at Cloudflare, but metadataBase, the school-page
canonical, robots.txt's Sitemap: line and the sitemap's own <loc> entries all
named the apex. Every one of those pointed Google at a redirect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Earlier years are to become a paid feature, so they stop being published.
The load-bearing part is that this is a change to the API, not only to the
page. /api/schools/{urn} is public and unauthenticated: leaving
admission_distance_history in the payload while declining to render it would
have handed the whole record to anyone who opened the network tab. It is
withheld at the source, and the page follows.
Nothing changes upstream. The tap, the plausibility band and
fact_admission_distance are untouched and still load every published year, so
restoring history for entitled callers is a change to one function in
data_loader rather than a re-collection.
What the reader now gets is the latest figure on the Admissions tile, and a
Distance section that answers the question the number alone cannot: whether
their own address falls inside it. Retitled to "How far away are you?", which
is what it now does — the previous title described a record that is no longer
there.
Removed with the history: the trend chart, the year-by-year table, the
per-year verdict strip, the trend summary and the coverage note, along with
their CSS. The section goes from 743px to 417px.
One consequence worth naming. A run of years used to soften a single close
call — a home just outside one year's cut-off was usually inside another. With
one year published, the "too close to call" band is the entire safety margin
between a parent and a place they do not have, so the verdict now names its
year, and the three outcomes are tinted apart rather than distinguished by
wording alone.
The existing stylesheet test earned its keep here: the three verdict classes
were referenced before they were written, and it caught them. Unstyled, a
"beyond the cut-off" result would have been indistinguishable from an "inside"
one — the exact failure the longhand class map was written to prevent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
Completes the last-distance-offered feature against the mockup: the
year-by-year record, the same numbers drawn over real streets, and the
reader's own address measured against them.
Serving the history
The first cut deliberately served only the latest year, because a plain
series would draw a trend line straight through gaps that are absences of
publication, not of a cut-off. That reasoning is answered rather than
abandoned: cutoffYearRows classifies every year in the span, and the chart
breaks the line rather than interpolating across it.
A missing year is not one fact but three. It may be unpublished; it may be
a year the school was not oversubscribed; or there may be no record at all.
Collapsing them into "no data" throws away the reassuring case and hides
the important caveat, so each is stated in words in the table.
The claim is held to what the data supports. fact_admissions.oversubscribed
compares FIRST PREFERENCES against places, which does not establish that
every applicant was offered one — so the copy says "places available on
first preferences" and a test asserts the stronger claim never appears.
The trend summary is not a verdict
It names both endpoints and their years and lets the reader conclude. It is
withheld below four published points, and a swing under a tenth of the
earlier figure is reported as "broadly the same" rather than dressed up as
a direction.
The postcode check
This is the only place on the site that answers a question about a family
rather than a school, so most of the care went into what it refuses to say.
postcodes.io returns a centroid covering roughly fifteen addresses, which
against a 500 m cut-off is a fifth of the whole distance — so a margin
inside 100 m returns "too close to call" rather than a place a family does
not have. Unpublished years count as unknown, never as a pass. The limits
are stated before the check is used, not revealed with the answer.
The postcode is geocoded in the browser and never stored.
Both templates
Banded and selective secondaries are exactly where this matters most, so
the detail is shared. The primary page gives it a third tab; the secondary
page is one flat panel by design and renders it inline.
Absence is explained rather than reported. A selective school's missing
figure is explained by how it admits; a consistently undersubscribed school
reads as good news.
Also makes the batch loader's test double honour ORDER BY. It was a no-op,
so "latest row per URN" was really "first row in the fixture" and the test
would have passed with the sort reversed or removed.
Verified: 214 frontend tests, 54 backend, 45/53 e2e green against staging
(the 8 cut-off journeys skip until the DAG runs). Rendered offline against
the real compiled CSS in both themes and at 390px; every new surface clears
WCAG AA, measured on composited pixels.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
Adds the cut-off distance a parent actually asks about — "how close do we
need to live?" — end to end: a Singer tap, dbt staging and mart models, an
Airflow DAG, and a tile on both detail templates. 3,597 schools across 57
local authorities carry a figure; the rest are unchanged.
There is no national source for this. Each LA publishes its own cut-offs in
its own format, and the collected CSV is transcribed from PDFs, spreadsheets
and web pages — so most of the work here is deciding what is safe to show.
Data
* tap-uk-school-distance loads the CSV verbatim into raw. Keyed on
(urn, year, school_name), because school_name carries the admission
route: (urn, year) alone collides on 118 keys and a reload would have
silently dropped every band but one.
* stg_school_distance applies a 25 m – 25 km plausibility band. The source
contains 0.0-mile rows (published where a school filled on a higher
criterion), 1-metre cut-offs, and one reading 533 miles — ~4% of rows,
all of which would put a visibly wrong number on a live page.
* fact_admission_distance collapses routes to one row per school per year
using the furthest, and keeps route_count so the page can say the figure
is the widest of several bands rather than the one for a given child.
Serving
* Kept out of fact_admissions: that mart is EES-derived and near-complete
for England, this one covers 57 LAs, and the two refresh independently.
* Latest year only. Coverage is ragged — a school may have 2021 and 2026
and nothing between — so a history array would invite a trend line drawn
through gaps that are absences of publication, not of a cut-off.
* The Admissions section now renders on either source. 3% of the schools
that render have a cut-off and no EES admissions row, and gating on
admissions alone would have hidden the figure on those pages.
Interface
* The year travels with the figure everywhere it appears; a cut-off
detached from its admissions round is not a fact about anything.
* "Not a fixed catchment — it moves every year" sits under every instance,
because that is the inference a parent will otherwise draw.
* Replaces a hardcoded "Historical distance cut-off data is not available
for this school" that appeared on every secondary page, including the
ones whose council does publish it. The absence is now stated only when
it is real, and names the authority that would hold it.
The tint costs the muted tokens their AA margin: measured on the composited
backdrop (not the computed one, which reports the untinted card), --text-muted
falls to 4.09:1 in dark theme. The tile uses --text-secondary instead — 6.50:1
dark, 6.60:1 light.
The DAG is manual, like the other annual ones: councils publish on allocation
day, each on its own timetable, so there is no date worth scheduling against.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
The graceful-degradation fallback keys off the column named in a Postgres
UndefinedColumn error, but the matcher only stripped an `s.` alias. The two
new dim_location columns (county, parliamentary_constituency) are selected
via the `l.` alias and Postgres reports them unquoted as
"column l.county does not exist" — which the old regex failed to match at
all, returning None.
If dim_school is rebuilt (telephone/nursery present) but dim_location is not
yet (county/parliamentary_constituency missing) — plausible since they are
independently-rebuilt dbt models — the fallback branch never matched and
load_school_data_as_dataframe() returned an empty DataFrame, showing zero
schools sitewide instead of degrading those columns to NULL.
Generalise the alias prefix to `\w+\.` and cover the l.-qualified case in
tests.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The KS2 SATs chart drew a single national-average line spanning the
full height of each subject's chart area, positioned at the national
*expected* value. But the area stacks two bars — Expected and Exceeding
— and the higher-standard/greater-depth national is a very different,
much lower figure (e.g. reading higher standard ~29% vs expected ~75%).
So the line crossed the Exceeding bar at the wrong place, making every
school's exceeding result look far below national when it wasn't.
The per-subject higher-standard nationals were already computed in the
fact_ks2_national_averages mart; they just weren't serialized. Fix:
- backend: add reading_high_pct, writing_gd_pct (writing = greater
depth) and maths_high_pct to the national-averages payload.
- SchoolDetailView: pass a nationalExceedingPct per subject, mapping
writing to the greater-depth figure.
- SatsChart: replace the single full-height line with a national marker
on each bar's own track (coral tick + "nat X%" in the bar header), so
Expected and Exceeding each sit against the correct benchmark.
KS2 only; the secondary Attainment 8 chart already uses one line for
one measure and is untouched.
Verified: tsc --noEmit, next build, and backend pytest (national
averages marts, incl. a new test guarding the per-subject nationals).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
New ees_ks4_national stream ingests the EES 'National characteristics
summary data' series (England, state-funded, all pupils). The old mart's
unweighted school means were 7-15 points off every headline measure and
produced an impossible national Progress 8 (-0.27). The API's computed
fallback is gone too: the footnote calls these figures official, so an
unbuilt mart now yields an empty series, never a stand-in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
The FSM chip anchored against disadvantaged_pct (a different measure,
FSM6+CLA) whenever fsm_pct was null — which it always was, since the
performance df has no fsm_pct. New fact_census_benchmarks mart supplies
pupil-weighted FSM/EAL means per phase; the KS2-column medians that
produced a bogus 50% 'secondary disadvantaged' anchor are gone.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
get_supplementary_data ran ~5 sequential DB round-trips per URN, so
/api/compare scaled at ~37ms/school (measured on staging: 1 school 155ms,
3 schools 220ms, 6 schools 340ms). get_supplementary_data_batch fetches
each table once with WHERE urn IN (...) and groups in Python, collapsing
5*N round-trips to a constant 5. get_supplementary_data is now a thin
wrapper so the detail endpoint is unchanged; the compare endpoint makes
one batched call. Each table degrades independently on failure.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
fact_ks4_national_averages is computed once at dbt build time (covered by
the EES DAG's stg_ees_ks4+ selector). _national_averages_payload now reads
both national-averages marts instead of scanning the performance dataframe
per year on every /api/compare request (~250ms saved per call). Fallback
for the deploy-before-DAG window computes the latest year only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
ReligiousCharacter 99 (~4k schools) and AdmissionsPolicy 9 (~5.6k) carry a
code with a blank name in the GIAS CSV; the generator skipped them so they
hit the Unknown(<code>) path — wrongly triggering the Faith-priority tag
and polluting filters. Blank-only codes now map to "" (byte-identical to
the old name pipeline); accepted_values lists extended to match the seed.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
str(ProgrammingError) embeds the full SQL, which contains every column
name — the substring check matched any error and could take the wrong
retry branch. Parse the missing column from exc.orig instead.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Closes the deploy window flagged by CI review — the backend now works
against both the old (name) and new (code) mart schemas.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- data_loader.load_school_data_as_dataframe now catches a ProgrammingError
whose message mentions has_sixth_form (psycopg2 UndefinedColumn) and
retries with a NULL-AS-has_sixth_form query variant, so the API keeps
serving data (and the app.py column-fallback branch stays reachable)
even before the nightly pipeline has rebuilt marts.dim_school.
- utils.convert_to_native now handles numpy.bool_ so GET /api/schools/{urn}
doesn't 500 once has_sixth_form is a populated bool-dtype column.
- Update the now-stale comment on the app.py age-range fallback branch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Schools without KS2/KS4 results (special post-16 institutions, sixth-form
centres, PRUs, new schools) come back from the marts LEFT JOIN with NaN in
every numeric column. school_info passed those raw pandas values straight
into JSONResponse, which renders with allow_nan=False, so the detail
endpoint 500d and the frontend turned that into a 404 on every such SEO
landing page.
Run school_info values through convert_to_native (the same treatment
yearly_data already gets), add backend unit tests plus a pytest step in PR
checks, and an e2e journey that finds a results-less school via the search
API and asserts its page renders.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>