The serialiser carries status through and computes no totals of its own.
The only aggregates in the payload are ones DfE published itself; whether
showing one is safe depends on how many of its components are suppressed,
which the frontend decides.
The batch guard grows from six tables to eight.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
The location tables carried one column: a percentage. A parent
shortlisting from a town page is asking a different question first —
does it take my child's age, is it a faith school, does it have a
nursery — and the page could not answer any of it.
Primary tables gain Ages, Religious character, Nursery and
Constituency; secondary tables the same minus Nursery, which is a
question about a different intake. An all-through school renders in
both groups, so its nursery shows under primary alone.
The measure moves to the second column rather than the last. Six
columns overflow a phone and .tableWrap turns that into a horizontal
swipe; with the measure last, the one number the page exists for is
the one scrolled off the screen.
Cell rules are the ones the school page already uses, so the two
surfaces cannot disagree about the same school: "Does not apply",
"None" and "Not applicable" all read as no religious character, and
the en-dash age normalisation moves into formatAgeSpan, which
formatAgeRange now delegates to.
Backend: nursery_provision and parliamentary_constituency were not in
the place response. Both are optional GIAS mart columns that
data_loader degrades to NULL, and the `in rows.columns` guard keeps a
mart the pipeline has not rebuilt working.
Also fixes a live bug on the same line: SCHOOL_COLUMNS already ends
with latitude and longitude, and the endpoint concatenated them again,
so pandas dropped one of every duplicated pair and warned "columns are
not unique" on each request. Ordered de-duplication removes the
warning and the silent drop.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
Code review, both findings valid.
The design doc claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the X-Forwarded-For fallback was
forgeable. That is true only for traffic that actually passed through
Cloudflare, and nothing in this process can verify that it did. Reaching
the origin directly, both headers are equally attacker-controlled — and
rotating CF-Connecting-IP mints a fresh rate-limit bucket per request,
defeating per-client limits on every endpoint including the
DataFrame-heavy /api/schools. Against abuse that is worse than the
shared bucket it replaced, which at least capped everyone together.
So the ceiling comes back. I dropped it earlier arguing it belonged at
Cloudflare; that argument assumed the keying was sound, and it is not.
GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window
against a total, independent of client identity, outermost so it refuses
before any work happens. Written by hand because slowapi cannot express
a global cap: default_limits and application_limits are both keyed by
key_func, and the latter needs middleware this app does not install.
It does not make the header trustworthy — it makes trusting it
survivable. The real fix is Authenticated Origin Pulls or an origin
firewall, now documented in DEPLOY.md as the open gap it is.
127.0.0.1 is exempt: the healthcheck curls localhost from inside the
container, and starving it would restart the container and turn a load
spike into an outage loop. Keyed on the peer address, never the Host
header, which the caller sets.
Second finding: suggest_schools_typesense promised "never raises" while
the parsing loop sat outside the try, so int(None) on a malformed
document would have made a keystroke a 500. The loop now skips bad rows
rather than dropping the whole list — and a hit with no document no
longer becomes a suggestion pointing at /school/0.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
A dedicated endpoint rather than a mode of /api/schools, because that
path filters and sorts 25,000 pandas rows per query while holding the
GIL — affordable once per search, not once per keystroke. A test asserts
the distinction directly by making load_school_data raise and requiring
the endpoint to answer anyway.
Nothing errors on ordinary input: a short query, no matches, or
Typesense being down are all 200 with an empty list.
Cached deliberately. Prefix queries repeat enormously across users and
school names change once a year, so s-maxage plus the existing ETag
middleware turns most keystrokes into 304s.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The limiter keyed on request.client.host, which in staging and prod is
the Next container — the backend has no published ports and nothing else
can reach it. So every browser user on the site shared one 60/minute
bucket per route. Measured against staging: 70 concurrent requests to
/api/schools returned exactly 60 OK and 10 refused, from one machine.
CF-Connecting-IP first. Cloudflare fronts both environments and
overwrites any client-supplied value, which a parsed X-Forwarded-For
chain does not guarantee. The XFF fallback is forgeable only from inside
the Docker network.
Named rather than hidden: the shared bucket was an accidental global
throttle on a single-process backend, and correct per-user keying
removes it. A real global ceiling belongs at Cloudflare, which is
already in the path; slowapi cannot express one without a second Limiter
and middleware this app does not install.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
One gate, at the source. The frontend needs no change: DistanceSection
already returns null when distance_m is missing, and the admissions
block already conditions on (admissions || admissionDistance). Only 57
local authorities publish cut-offs, so the off-path is the commonest
path on the site and is well covered already.
Absent, not null. /api/schools/ is public and unauthenticated, so a
field left in the payload is a published field — the reasoning already
recorded in c9a1892 when history was withheld. The two are also
different claims: null says this school has no cut-off, absent says
cut-offs are not being published at all. The frontend type now says so.
The feature is on main and live on staging and has never reached
production, which is what makes it the right first consumer: the flag
lets the code promote without the feature appearing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The endpoint and its exposure control ship together on purpose. The
moment /api/flags exists, app/api/[...path] forwards it — and the
response names every unreleased feature the codebase knows about, along
with whether it is on. Publishing that is the opposite of shipping dark.
Denied on an exact first-segment match, not a prefix, so /api/flagship
does not go down with /api/flags. Next reads the endpoint server-side
over the Docker network, which never transits the public proxy.
jest.setup.js now guards its browser globals. It runs for every suite,
including the one that declares @jest-environment node to exercise the
route handler — NextRequest needs Fetch API globals jsdom lacks, and
there is no window there to define matchMedia on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Every place page built its phase links as /schools/[slug]/[phase], the
shape that belongs to towns alone.
On an authority page that pointed into the town namespace. For 87 of the
151 authorities the target does not exist and the link 404s; for the
other 64 it resolves to the town of the same name — a different set of
schools, which is precisely the near-duplicate the two namespaces were
introduced to prevent. On an outcode page it 404s outright.
Two causes behind it, both a rule written twice and inherited by only
one of the places that needed it.
The authority phase route was in the spec and dropped by the plan, which
built the three bare routes and no fourth. The sitemap is generated from
the place registry, which was right about them all along, so 302
authority phase URLs have been submitted to Google and every one 404s.
Adding the route makes the sitemap true and serves a real query —
admissions are authority-run, so "primary schools in Kent" is how a
parent searches before they have settled on a town.
The outcode variants were the opposite: the registry computed phases for
outcodes although the spec gives them no route, and the sitemap knew to
skip them while the API did not. The registry now decides alone, and the
sitemap's duplicate of that rule is gone.
Also: an authority under the five-school threshold has no page, so the
API sends a null slug for it and the page names it without linking.
Two English authorities are in that position. It was unreachable in
today's data — verified across the EC and TR outcodes — but the thin
place redirect would have sent a reader to a 404 the year it isn't.
The e2e journey now walks every /schools link a page of each family
emits and requires a 200, which is the check that was missing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Someone on a place page is usually looking for a school they can name, so the
order should serve scanning for it rather than ranking. /api/rankings keeps
its league-table ordering; this is a place-page decision, not a site-wide one.
Sorted case-insensitively, or a capitalised name would sort ahead of every
lowercase one.
The change made five pieces of copy untrue, so they go with it. The phase
variant titled itself "— Ranked", and all four route families described
themselves as "ranked by SATs and GCSE results". A page that opens by claiming
an order it does not keep is worse than one that claims nothing.
The ItemList markup carried `position` with no declared order, which reads as
a ranking. It now declares ItemListOrderAscending, so the structured data says
what the table does.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
SW19 is mostly Merton but partly Wandsworth, and the page said only Merton.
The cause was one field doing two jobs: _parent_authority takes the modal
authority, which is right for a 301 target and wrong as a statement about
where a place is.
This is not a corner case. A quarter of viable outcodes (425 of 1,760) and a
third of viable towns (263 of 783) cross an authority boundary — Bedford the
town spans Bedford and Central Bedfordshire.
Place now carries `authorities`, every authority holding at least a tenth of
the schools and at least two of them, largest first. parent_authority stays
single and unchanged, because a redirect still needs one target.
The share threshold exists because GIAS carries postcode errors: EN6 lists two
Shropshire schools among fourteen in Hertfordshire, and a bare "any authority
present" rule would print those as though they were real. A place too small or
too fragmented to clear the threshold still names its largest, so the page
never goes silent about where it is.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
/schools/[place]/[phase] shipped as routes but reached nothing. The sitemap
emitted one URL per registry entry and the registry had no phase dimension, so
~950 pages were absent from every sitemap — and PlaceView did not link them
either, leaving them reachable by nothing at all.
That is the query shape the baseline actually showed: 'primary schools in
beccles', 'secondary schools in brentwood'. Publishing the routes without a
path in meant building for the demand and then hiding from it.
Place now carries phase_urns so the per-phase threshold can be applied without
re-querying, the sitemap emits a variant wherever a phase clears the threshold
on its own, and the API exposes the qualifying phases so the place page links
only variants that exist. Outcodes are excluded: nobody searches 'primary
schools in SW11' and those routes do not exist.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Separate children per family so Search Console reports the location layer's
indexation apart from the school pages' — which is the point of the index
built in W1, and the number the stop condition watches.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The registry is cached for the process and reset by the same admin endpoint
that rebuilds the sitemaps, so places and sitemap always describe the same
corpus rather than drifting apart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
_school_sitemap_rows tested only the latest year's row, which quietly dropped
every school with results in its history but a null row for the most recent
year — a school that stopped reporting, or whose figures were suppressed for
small-cohort disclosure.
The Mallard Academy (150367) is the case that caught it: real KS2 results for
2015-16 through 2018-19, then null rows from 2022-23 on. Its detail page shows
all four years; the sitemap omitted it. Sampling 40 of the 2,206 excluded
schools found 4 like this, so roughly 220 real pages were being withheld.
Publishable is now a property of the school, computed across every row, while
lastmod still comes from the latest row so the most recent Ofsted date wins.
The field list is a module constant shared with _has_publishable_data so the
two checks cannot drift.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Search Console reports coverage per submitted sitemap, so one file per page
family is what will make W2's location pages measurable when they land. The
index's lastmod is generation time, which is the correct semantic there —
unlike on a <url>, where it would be a claim we cannot support.
Children sit under /sitemaps/ because Next only treats a whole bracketed path
segment as dynamic; a route folder named sitemap-[...parts] would be read as a
literal static segment and never match. Confirmed by the build output, which
lists /sitemaps/[...parts] as a dynamic route.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Drops the schools with neither results nor an Ofsted grade, adds /admissions
which was never listed, replaces the invented priority and changefreq with a
lastmod taken from each school's Ofsted date.
lastmod is omitted where no date is known rather than defaulted to now. An
always-now lastmod is a claim Google learns to distrust; absent honestly
means unknown.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The apex 301s to www at Cloudflare, but metadataBase, the school-page
canonical, robots.txt's Sitemap: line and the sitemap's own <loc> entries all
named the apex. Every one of those pointed Google at a redirect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Earlier years are to become a paid feature, so they stop being published.
The load-bearing part is that this is a change to the API, not only to the
page. /api/schools/{urn} is public and unauthenticated: leaving
admission_distance_history in the payload while declining to render it would
have handed the whole record to anyone who opened the network tab. It is
withheld at the source, and the page follows.
Nothing changes upstream. The tap, the plausibility band and
fact_admission_distance are untouched and still load every published year, so
restoring history for entitled callers is a change to one function in
data_loader rather than a re-collection.
What the reader now gets is the latest figure on the Admissions tile, and a
Distance section that answers the question the number alone cannot: whether
their own address falls inside it. Retitled to "How far away are you?", which
is what it now does — the previous title described a record that is no longer
there.
Removed with the history: the trend chart, the year-by-year table, the
per-year verdict strip, the trend summary and the coverage note, along with
their CSS. The section goes from 743px to 417px.
One consequence worth naming. A run of years used to soften a single close
call — a home just outside one year's cut-off was usually inside another. With
one year published, the "too close to call" band is the entire safety margin
between a parent and a place they do not have, so the verdict now names its
year, and the three outcomes are tinted apart rather than distinguished by
wording alone.
The existing stylesheet test earned its keep here: the three verdict classes
were referenced before they were written, and it caught them. Unstyled, a
"beyond the cut-off" result would have been indistinguishable from an "inside"
one — the exact failure the longhand class map was written to prevent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
Completes the last-distance-offered feature against the mockup: the
year-by-year record, the same numbers drawn over real streets, and the
reader's own address measured against them.
Serving the history
The first cut deliberately served only the latest year, because a plain
series would draw a trend line straight through gaps that are absences of
publication, not of a cut-off. That reasoning is answered rather than
abandoned: cutoffYearRows classifies every year in the span, and the chart
breaks the line rather than interpolating across it.
A missing year is not one fact but three. It may be unpublished; it may be
a year the school was not oversubscribed; or there may be no record at all.
Collapsing them into "no data" throws away the reassuring case and hides
the important caveat, so each is stated in words in the table.
The claim is held to what the data supports. fact_admissions.oversubscribed
compares FIRST PREFERENCES against places, which does not establish that
every applicant was offered one — so the copy says "places available on
first preferences" and a test asserts the stronger claim never appears.
The trend summary is not a verdict
It names both endpoints and their years and lets the reader conclude. It is
withheld below four published points, and a swing under a tenth of the
earlier figure is reported as "broadly the same" rather than dressed up as
a direction.
The postcode check
This is the only place on the site that answers a question about a family
rather than a school, so most of the care went into what it refuses to say.
postcodes.io returns a centroid covering roughly fifteen addresses, which
against a 500 m cut-off is a fifth of the whole distance — so a margin
inside 100 m returns "too close to call" rather than a place a family does
not have. Unpublished years count as unknown, never as a pass. The limits
are stated before the check is used, not revealed with the answer.
The postcode is geocoded in the browser and never stored.
Both templates
Banded and selective secondaries are exactly where this matters most, so
the detail is shared. The primary page gives it a third tab; the secondary
page is one flat panel by design and renders it inline.
Absence is explained rather than reported. A selective school's missing
figure is explained by how it admits; a consistently undersubscribed school
reads as good news.
Also makes the batch loader's test double honour ORDER BY. It was a no-op,
so "latest row per URN" was really "first row in the fixture" and the test
would have passed with the sort reversed or removed.
Verified: 214 frontend tests, 54 backend, 45/53 e2e green against staging
(the 8 cut-off journeys skip until the DAG runs). Rendered offline against
the real compiled CSS in both themes and at 390px; every new surface clears
WCAG AA, measured on composited pixels.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
Adds the cut-off distance a parent actually asks about — "how close do we
need to live?" — end to end: a Singer tap, dbt staging and mart models, an
Airflow DAG, and a tile on both detail templates. 3,597 schools across 57
local authorities carry a figure; the rest are unchanged.
There is no national source for this. Each LA publishes its own cut-offs in
its own format, and the collected CSV is transcribed from PDFs, spreadsheets
and web pages — so most of the work here is deciding what is safe to show.
Data
* tap-uk-school-distance loads the CSV verbatim into raw. Keyed on
(urn, year, school_name), because school_name carries the admission
route: (urn, year) alone collides on 118 keys and a reload would have
silently dropped every band but one.
* stg_school_distance applies a 25 m – 25 km plausibility band. The source
contains 0.0-mile rows (published where a school filled on a higher
criterion), 1-metre cut-offs, and one reading 533 miles — ~4% of rows,
all of which would put a visibly wrong number on a live page.
* fact_admission_distance collapses routes to one row per school per year
using the furthest, and keeps route_count so the page can say the figure
is the widest of several bands rather than the one for a given child.
Serving
* Kept out of fact_admissions: that mart is EES-derived and near-complete
for England, this one covers 57 LAs, and the two refresh independently.
* Latest year only. Coverage is ragged — a school may have 2021 and 2026
and nothing between — so a history array would invite a trend line drawn
through gaps that are absences of publication, not of a cut-off.
* The Admissions section now renders on either source. 3% of the schools
that render have a cut-off and no EES admissions row, and gating on
admissions alone would have hidden the figure on those pages.
Interface
* The year travels with the figure everywhere it appears; a cut-off
detached from its admissions round is not a fact about anything.
* "Not a fixed catchment — it moves every year" sits under every instance,
because that is the inference a parent will otherwise draw.
* Replaces a hardcoded "Historical distance cut-off data is not available
for this school" that appeared on every secondary page, including the
ones whose council does publish it. The absence is now stated only when
it is real, and names the authority that would hold it.
The tint costs the muted tokens their AA margin: measured on the composited
backdrop (not the computed one, which reports the untinted card), --text-muted
falls to 4.09:1 in dark theme. The tile uses --text-secondary instead — 6.50:1
dark, 6.60:1 light.
The DAG is manual, like the other annual ones: councils publish on allocation
day, each on its own timetable, so there is no date worth scheduling against.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
Add seven school-identity fields to the detail header (both primary and
secondary views): age range, religious character, nursery and sixth-form
indicators as chips; telephone (tel: link), county and parliamentary
constituency as header details. religious_denomination, age_range and
has_sixth_form were already served; telephone, nursery_provision, county
and parliamentary_constituency are newly wired through the marts query
(with a NULL fallback for un-rebuilt marts, mirroring has_sixth_form) and
the school_info API response.
Remove three UI sections the backend never populated (always null): Year 1
Phonics, the SEN "types of additional needs" breakdown, and the average
class-size card — along with their now-dead props, route plumbing, and the
SenDetail/Phonics types + class_size_avg field.
Extend the e2e detail journey to assert the Phonics section is gone and the
new header fields render when the record carries them.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The KS2 SATs chart drew a single national-average line spanning the
full height of each subject's chart area, positioned at the national
*expected* value. But the area stacks two bars — Expected and Exceeding
— and the higher-standard/greater-depth national is a very different,
much lower figure (e.g. reading higher standard ~29% vs expected ~75%).
So the line crossed the Exceeding bar at the wrong place, making every
school's exceeding result look far below national when it wasn't.
The per-subject higher-standard nationals were already computed in the
fact_ks2_national_averages mart; they just weren't serialized. Fix:
- backend: add reading_high_pct, writing_gd_pct (writing = greater
depth) and maths_high_pct to the national-averages payload.
- SchoolDetailView: pass a nationalExceedingPct per subject, mapping
writing to the greater-depth figure.
- SatsChart: replace the single full-height line with a national marker
on each bar's own track (coral tick + "nat X%" in the bar header), so
Expected and Exceeding each sit against the correct benchmark.
KS2 only; the secondary Attainment 8 chart already uses one line for
one measure and is untouched.
Verified: tsc --noEmit, next build, and backend pytest (national
averages marts, incl. a new test guarding the per-subject nationals).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
M1: admissions rounds are now selected by the active phase tab
(admissionsForPhase) — an all-through school's Year 7 round no longer
masquerades as Reception odds beside pure primaries; honest per-cell and
section fallbacks name the round (Reception / Year 7).
M2: Ofsted sentinel codes (9 = not applicable) never render as judgement
chips, and the sixth-form judgement — previously dropped — now renders
for schools that have one.
M3: 'What this means' is phase- and type-aware: selective schools get
entrance-test framing, secondary faith schools a faith-criteria note, and
the primaries' distance template never appears on the secondary tab
(admissions_policy now exposed in compare school_info).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
New ees_ks4_national stream ingests the EES 'National characteristics
summary data' series (England, state-funded, all pupils). The old mart's
unweighted school means were 7-15 points off every headline measure and
produced an impossible national Progress 8 (-0.27). The API's computed
fallback is gone too: the footnote calls these figures official, so an
unbuilt mart now yields an empty series, never a stand-in.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
The FSM chip anchored against disadvantaged_pct (a different measure,
FSM6+CLA) whenever fsm_pct was null — which it always was, since the
performance df has no fsm_pct. New fact_census_benchmarks mart supplies
pupil-weighted FSM/EAL means per phase; the KS2-column medians that
produced a bogus 50% 'secondary disadvantaged' anchor are gone.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
get_supplementary_data ran ~5 sequential DB round-trips per URN, so
/api/compare scaled at ~37ms/school (measured on staging: 1 school 155ms,
3 schools 220ms, 6 schools 340ms). get_supplementary_data_batch fetches
each table once with WHERE urn IN (...) and groups in Python, collapsing
5*N round-trips to a constant 5. get_supplementary_data is now a thin
wrapper so the detail endpoint is unchanged; the compare endpoint makes
one batched call. Each table degrades independently on failure.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
fact_ks4_national_averages is computed once at dbt build time (covered by
the EES DAG's stg_ees_ks4+ selector). _national_averages_payload now reads
both national-averages marts instead of scanning the performance dataframe
per year on every /api/compare request (~250ms saved per call). Fallback
for the deploy-before-DAG window computes the latest year only.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
- data_loader.load_school_data_as_dataframe now catches a ProgrammingError
whose message mentions has_sixth_form (psycopg2 UndefinedColumn) and
retries with a NULL-AS-has_sixth_form query variant, so the API keeps
serving data (and the app.py column-fallback branch stays reachable)
even before the nightly pipeline has rebuilt marts.dim_school.
- utils.convert_to_native now handles numpy.bool_ so GET /api/schools/{urn}
doesn't 500 once has_sixth_form is a populated bool-dtype column.
- Update the now-stale comment on the app.py age-range fallback branch.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Schools without KS2/KS4 results (special post-16 institutions, sixth-form
centres, PRUs, new schools) come back from the marts LEFT JOIN with NaN in
every numeric column. school_info passed those raw pandas values straight
into JSONResponse, which renders with allow_nan=False, so the detail
endpoint 500d and the frontend turned that into a 404 on every such SEO
landing page.
Run school_info values through convert_to_native (the same treatment
yearly_data already gets), add backend unit tests plus a pytest step in PR
checks, and an e2e journey that finds a results-less school via the search
API and asserts its page renders.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Removes the 'What Parents Say' section and all supporting elements:
Frontend:
- Drop the OfstedParentView type, the parent_view field, the survey
section and the 'X% would recommend' callouts in the primary and
secondary detail views, the Parents nav item, and the parent-view CSS.
Backend:
- Remove the FactParentView model, its loading in data_loader, and
parent_view from the school-details API response.
- Bump SCHEMA_VERSION to 6 and add an idempotent drop step
(DROP TABLE IF EXISTS marts.fact_parent_view) to the CLI migration;
add scripts/sql/drop_fact_parent_view.sql to apply directly to the
dbt-owned marts DBs on staging and prod.
Pipeline:
- Delete the stg_parent_view + fact_parent_view dbt models and their
source/schema entries, the tap-uk-parent-view Meltano extractor, and
the monthly Parent View DAG; drop it from the Dockerfile and the
staging bootstrap docs.
The rest of dbt (which builds every mart the app reads) is untouched.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The rankings endpoint validated year with le=2100, but the database
stores academic-year codes like 201819, so any explicit year selection
returned a 422 and the rankings page rendered its empty state. Widen
the bound to cover the codes and extend the e2e journey to pick a
specific year and assert the table stays populated.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The school detail page only showed the latest admissions year. We store
every year, which is more decision-relevant for parents (the trend and its
consistency matter more than a single noisy year).
Backend now returns the full admissions_history (oldest first) alongside the
existing latest-year object. The primary SchoolDetailView gains a header
toggle ("This year | N-year trend") that swaps the Q&A for an SVG sparkline
of the first-choice offer rate. The toggle only appears when >=2 years carry
an offer rate; otherwise it falls back to the single-year card. Both views
share one CSS-grid cell so switching causes no layout shift.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Frontend
- Dynamic-import Chart.js components on detail/compare views so Chart.js
no longer ships in initial JS.
- Drop force-dynamic on home, compare, rankings so internal data fetches
reuse Next.js's per-call revalidate cache.
- Switch /school/[slug] to ISR with a 7-day revalidate window (school
data updates annually).
- Preconnect to analytics + postcodes.io; remove redundant defer on the
Umami Script tag (afterInteractive already covers it).
- Bump images.minimumCacheTTL to 1 year.
- Extract HowItWorks and Editorial sections as server components passed
to HomeView via slot props so their JSX stays out of the client bundle.
Backend
- Add GZipMiddleware (min 512 bytes).
- Add CacheAndETagMiddleware: per-path Cache-Control with long s-maxage
+ stale-while-revalidate, ETag generation, and 304 on If-None-Match.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The /api/rankings endpoint returned each row keyed by the metric's
column name (e.g. rwm_high_pct) but never under a generic `value`
field. The frontend RankingItem type and RankingsView both read
ranking.value, so every row rendered "—" for every metric — the
default rwm_expected_pct included.
Add `df["value"] = df[metric]` before JSON serialisation so the
frontend gets the value it has always expected. The raw metric
column is still in the row for any caller that wants it explicitly.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
P1 (backend/data_loader.py): Add load_latest_school_data() which pre-computes
the one-row-per-school latest-year snapshot (groupby, prev-year trend merge)
once at startup instead of on every /api/schools request. get_schools route
now starts from the cached snapshot rather than rebuilding it.
S3 (backend/app.py): Wrap synchronous geocode_single_postcode() call in
asyncio.to_thread() so postcode lookups no longer block the uvicorn event
loop. Admin reload endpoint also uses to_thread for both cache primes.
P2 (nextjs-app/components/HomeView.tsx): Add mapParamsRef guard so switching
back to map view does not re-fetch 500 schools when search params haven't
changed. Reset ref on new searches so fresh data is always fetched.
P3 (nextjs-app/lib/chartSetup.ts): Extract Chart.js registration into a
shared side-effect module. ComparisonChart and PerformanceChart now import
it instead of each calling ChartJS.register() independently.
P4 (backend/database.py): Remove unnecessary db.commit() from the read-only
get_db_session() context manager — saves a DB round-trip on every request.
P5 (backend/database.py): Add pool_recycle=1800 to SQLAlchemy engine to
prevent stale TCP connections from accumulating in long-running processes.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Redesign the School Details page for better parent comprehension:
- New SatsChart component: horizontal cascade bars with ruler scale and
national average marker (teal/coral palette matching site theme)
- Admissions section: visual progress bar showing 1st-preference demand
vs available places, colour-coded by oversubscription status
- Historical data: collapse raw year-by-year table behind a disclosure
element while keeping the performance line chart always visible
- EAL metric: add national average comparison via DeltaChip (backend now
includes eal_pct in national averages endpoint)
- New formatWithSuppression utility for null/suppressed data handling
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The school_info object was missing total_pupils entirely, so the frontend
always fell back to the KS4 exam cohort from yearly_data. Now selects
s.total_pupils (GIAS NumberOfPupils — full school roll) as gias_total_pupils
in the main query and exposes it as total_pupils on school_info.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Replaces computed means from our school dataset with the published DfE
national headline figures for the KS2 chart reference line.
- tap-uk-ees: new EESKs2NationalStream fetches the stable EES data-catalogue
CSV (one row per year, England national total, AllSchools filter)
- dbt staging: stg_ees_ks2_national normalises columns, casts to float,
filters to years >= 201617
- dbt mart: fact_ks2_national_averages — one row per year, official figures
- backend/models: Ks2NationalAverage SQLAlchemy model
- backend/app: /api/national-averages queries the mart for KS2 by_year;
secondary by_year stays computed (no DfE KS4 national dataset yet)
- DAG: extract_ks2_national task added to school_data_annual_ees,
runs in parallel with the main EES extract
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Previously the dashed reference line was a flat horizontal at the latest
year's national average across all historical data, implying the national
figure was constant. Now the backend returns per-year averages in `by_year`
and the chart maps each data year to its own national average, so the
reference line correctly reflects how the national picture changed over time
(including COVID recovery dip/recovery).
- backend: /api/national-averages now includes `by_year` list alongside
existing `year`/`primary`/`secondary` latest-year snapshot
- types: NationalAverages extended with `by_year: NationalAveragesYear[]`
- PerformanceChart: accepts `nationalByYear` prop; builds per-year series
aligned to school data years, falling back to scalar prop if absent
- SchoolDetailView + SecondarySchoolDetailView: pass `nationalAvg.by_year`
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Root cause: the UNION ALL query in data_loader.py produced two rows per
all-through school per year (one KS2, one KS4), with drop_duplicates()
silently discarding the KS4 row. Fixes:
- New dbt mart `fact_performance`: FULL OUTER JOIN of fact_ks2_performance
and fact_ks4_performance on (urn, year). One row per school per year.
All-through schools have both KS2 and KS4 columns populated.
- data_loader.py: replace 175-line UNION ALL with a simple JOIN to
fact_performance. No more duplicate rows or drop_duplicates needed.
- sync_typesense.py: single LATERAL JOIN to fact_performance instead of
two separate KS2/KS4 joins.
- app.py: remove drop_duplicates (no longer needed); add PHASE_GROUPS
constant so all-through/middle schools appear in primary and secondary
filter results (were previously invisible to both); scope result_filters
gender/admissions_policies to secondary schools only.
- HomeView.tsx: isSecondaryView is now majority-based (not "any secondary")
and isMixedView shows both sort option sets for mixed result sets.
- school/[slug]/page.tsx: all-through schools route to SchoolDetailView
(renders both SATs + GCSE sections) instead of SecondarySchoolDetailView
(KS4-only). Dedicated SEO metadata for all-through schools.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Remove 10-mile radius option; cap backend radius max at 5 miles
- Raise backend page_size max to 500 so map can fetch all schools in one call
- HomeView: when map view is active, fetch all schools within radius
(page_size=500) instead of showing only the paginated first page;
falls back to initial SSR schools while loading
- SchoolMap/LeafletMapInner: accept referencePoint prop and render a
distinctive coral circle pin at the search postcode location
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Load-more requests read URL params (postcode, radius, etc.) but page_size
is never in the URL — it's hardcoded in page.tsx. Without it the backend
received page_size=None, hit a TypeError on (page-1)*None, returned 500,
and the silent catch left the user stuck on page 1.
In a dense area (e.g. Wimbledon SW19) 50 schools fit within ~1.8 miles,
so page 1 never shows anything beyond that regardless of selected radius.
Fix:
- Backend: give page_size a safe default of 25 instead of None
- Frontend: explicitly pass initialSchools.page_size in load-more params
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Backend builds sitemap.xml from school data at startup (in-memory)
- POST /api/admin/regenerate-sitemap refreshes it after data updates
- New Airflow DAG (sitemap_generate) runs Sundays 05:00 and calls the endpoint
- Next.js proxies /sitemap.xml to the backend; removes the slow dynamic sitemap.ts
- docker-compose passes BACKEND_URL + ADMIN_API_KEY to Airflow env
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
1. Simpler home page: only search box on landing, no filter dropdowns
2. Advanced filters: hidden behind toggle on results page, auto-open if active
3. Per-school phase rendering: each row renders based on its own data
4. Taller 4-line rows with context line (type, age range, denomination, gender)
5. Result-scoped filters: dropdown values reflect current search results
6. Fix blank filter values: exclude empty strings and "Not applicable"
7. Rankings: Primary/Secondary phase tabs with phase-specific metrics
8. Compare: Primary/Secondary tabs with school counts and phase metrics
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>