5c0ccc693deeca46b9ca043d4056bafc8b00b3cc
146
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
5c0ccc693d |
fix: order nearby schools by distance, not by how alike they are
Reported from staging: a Catholic primary showed six Catholic primaries, none of them close enough to be a real option, and omitted the community school down the road. Three causes, compounding. Ranking put tier before distance, so a faith match at 2.9 miles outranked a community school at 0.3. The ENOUGH=3 stopping rule — added so a cap of six would not drag in weak distant matches — filled the row from the best tier before it ever widened, which is what made every card Catholic. And a 3-mile tier-1 radius is sane for a secondary and most of a city for a primary, whose catchments are routinely under a mile. The premise was backwards. For a parent, distance is a constraint and intake is a preference; a school beyond a primary catchment is not a weaker option, it is not an option. So distance now decides the order and nothing else does. The hard filters are untouched — they were always where the defensibility lived. Similarity survives as chips on the card: reported, so a reader applies their own weighting, rather than ranked, so we apply ours for them. Reach is capped per phase (primary 2, secondary 6, post-16 10) as a sanity bound, not a target: ordering already handles density, so the cap only decides what happens where an area is sparse. A primary with nothing inside two miles now renders no section, which is the honest answer. Deleted: the tier system, the stopping rule, the tier-dependent lede, the `tier` field, the tier-3 fallback chip and its style. select_similar also stops taking is_secondary — it reads the phase from the subject's own row, so no caller can hand it one that disagrees with the data. The heading is now "Other schools nearby". The hard filters still guarantee a comparable set, but nothing ranks on likeness, so the heading no longer says it does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
2175dccb7c |
fix(api): treat "16 plus" as secondary, the way PHASE_GROUPS already does
PR Checks / Frontend Typecheck + Tests (pull_request) Canceled after 9s
PR Checks / Backend Smoke (pull_request) Canceled after 0s
PR Checks / Build Backend (no push) (pull_request) Canceled after 0s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 0s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
GIAS phase 6 is "16 plus", and PHASE_GROUPS deliberately files it in the secondary group. The payload helper decided the same question with `"secondary" in phase_text`, which that value does not satisfy — so a sixth-form college was handed the primary bucket and offered infant schools as its peers, with the KS2 metric key to label them. No crash; just a page confidently showing the wrong schools. The decision now lives in similar_schools.is_secondary_phase, beside the PHASE_GROUPS bucket it selects from, so the two cannot drift again. A test pins them together. The same binary assumption had a second output. computeSchoolFlags tests for the substring too, so a 16-plus school renders the primary template, and the composer was labelling the section from the template: "Other primary schools near <sixth form college>" above a row of secondaries. The section now derives its noun from the school's own phase, which also removes the duplicated wording from both composers. A 16-plus school's candidates span the whole secondary group, so no single noun fits and it gets the honest general one. Reported in review on #150. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
52b00ac752 |
feat(api): serve similar schools on the detail endpoint
Rides in the existing payload rather than a new endpoint: the page already makes one server fetch for its data, and /school/[slug] regenerates weekly, so the per-request cost is paid once per school per week. Wrapped so a failure in selection never 500s a page that is otherwise complete — the posture get_supplementary_data already takes. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
b571d9c549 |
feat(api): decide which nearby schools a page may offer
Hard filters encode claims the section may not make — a selective school is not an alternative to a non-selective one, a special school is not comparable to a mainstream one, a Girls school is not an option for a Boys school's reader — so they never relax. Soft preferences describe closeness of fit, so they relax across three tiers, and only far enough to reach three; the remaining slots up to six fill from the tiers already opened. PHASE_GROUPS moves to schemas.py so this module can share it without importing app, which would be a cycle. _mask() exists because Series.apply on an empty Series returns a DataFrame, and using that as a mask drops every column — so the next lookup raises KeyError instead of yielding no rows. A special school with no special school near it empties the frame at the provision filter, which is the ordinary case for most special schools, so this was a crash on a common path rather than an edge case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
5ad1cbfb53 |
fix(api): bound the search candidate set instead of draining Typesense
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 39s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 6m44s
Fetching every match kept scoped searches correct but left the number of round trips in the caller's hands: a one-letter query, or a deliberately broad one, walked the whole collection a page at a time. Cap the candidate set at 1,000 URNs — four pages — and return the relevance-ordered prefix when the ceiling is hit. That is still far more than one page, so the API's own authority, phase and postcode filters keep the matches they need, while latency and upstream load stay bounded. A capped query is logged so a genuinely truncated search is visible. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
9b75f54206 |
fix(api): publish a validated dataset, and stop truncating search
Reload cleared the caches first and rebuilt afterwards, so any failure left the API serving nothing, and requests arriving mid-reload saw a half-swapped state. It now builds and validates the replacement frames, place registry, reverse index and sitemaps off the request loop, then publishes them in one synchronous step under a lock. A failed reload returns 503 and keeps the previous data. Sitemap regeneration takes the same path rather than clearing the live registry up front. Typesense search returned at most one page of hits and used an empty list for both "no matches" and "search is down", so a genuine empty result silently fell back to substring matching. It now pages through every candidate and returns None only on failure; the fallback matches literally, since a query containing regex metacharacters used to throw. Empty datasets answer 503 rather than 200-with-nothing or a misleading 404, so callers can tell an outage from an absent school. Adds /api/release, which reports the build identity baked into the image. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
1d8858fbda |
chore: remove the code the legacy CSV importer left behind
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m2s
`backend/migration.py` and `scripts/migrate_csv_to_db.py` import `School`, `SchoolResult`, `init_db` and `set_db_schema_version` — names that no longer exist. `scripts/geocode_schools.py` imports the same removed ORM model. None of the three can be imported against the current backend, so they were not dormant utilities anyone could fall back on; they were files that would fail on the first line. `backend/version.py` existed only to hand `SCHEMA_VERSION` to that importer, and the FastAPI lifespan performs no version-triggered import. Three symbols go with them, each confirmed to have no caller: the unvectorised `haversine_distance`, superseded by the inline NumPy calculation in search; `fetcher`, an SWR helper for a dependency this project does not install; and `kmToMiles`. `calculateDistance` stays — CutoffMapPanel uses it. Two comments pointed at `migrate_csv_to_db.py --drop` to explain why Payload owns its own schema. The reason survives the script: blog content must stay clear of the school marts and Airflow's metadata. Reworded rather than deleted, so the constraint keeps its justification. docs/LEGACY_CODE.md records what was removed and where to find it in history. It also records what was deliberately *not* removed, which is the more useful half: unused UI components awaiting a design decision, manual data utilities whose operators a repository search cannot see, and fallbacks that look obsolete but are load-bearing — `data_loader.py`'s older-mart branches, the generated GIAS dictionary copies, and the `legacy`-named dbt models that annual DAG selectors explicitly include. A zero-import count is evidence, not a verdict. The scripts that fetch DfE CSVs are marked historical and kept, pending confirmation that nobody runs them by hand. Checked: 190 backend tests, 429 frontend tests, `tsc --noEmit` clean. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan |
||
|
|
b0d5334e06 |
perf(places): index the reverse lookup, and isolate the registry in tests
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m47s
Two review findings, both confirmed before fixing.
The client fixture in test_school_details.py patched load_school_data
but not _place_registry, which is a module-level cache. A probe settled
it rather than an argument: poisoning the global with a registry built
from data the fixture never saw, then issuing the fixture's own
request, returned that other dataset's places. So the new
`places == []` assertion was satisfied by a stale registry exactly as
well as by the fixture's own data, and proved nothing. Every other test
that touches place data already reset it; the fixture predates places
existing and was never updated. It resets it now.
places_for_urn walked every place in the registry and did a tuple
membership test against each, on /api/schools/{urn}, the site's
highest-traffic endpoint. It now reads a dict built once per registry.
Measured against a synthetic corpus of 27k schools in 1,650 places:
0.118ms per request becomes 0.0001ms, with the index built once in
21ms. Production carries ~5,000 places, so the scan there is larger
again. The absolute saving per request is small; the point is that it
is repeated on every school page view and costs nothing to remove.
The index is cached against the registry by identity rather than
behind a second flag. Anything that drops _place_registry — every test
that touches place data does — gets a fresh registry object, which no
longer matches what the index was built from, so the index rebuilds
with it. A separate _place_index = None would be one more thing to
forget, and a stale reverse index is precisely the first finding's bug
wearing a different hat.
That invalidation has its own test, and the test was checked by
breaking the identity check: five tests fail without it, so three
existing ones were already relying on it too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
|
||
|
|
d65eb58883 |
fix(seo): build the phase links the docstring already promised
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 3m48s
Review caught _places_payload documenting a `phase_url` the function never returned. The docstring was not stray prose: the approved design included the phase variant — "Primary schools in Beccles" was one of its four example links — and it was dropped during implementation without being mentioned. Deleting the sentence would have closed the report while losing the feature, so the links are built instead. These are the pages that most needed them. ~950 phase variants were once reachable by nothing at all: absent from every sitemap and unlinked from the place page. "Primary schools in brentwood" is the query they exist to answer. Membership is read from the registry's own `phase_urns` rather than re-derived from the school's phase string. The registry already decides which phases a place publishes and which schools are listed on each, so asking it is both shorter and the only way the link cannot disagree with the page it points at. It also means outcodes need no special case: they carry empty `phase_urns` by design, because nobody searches "primary schools in SW11", so they report no phase links on their own. `phases` is a list rather than a single url. An all-through school is listed on both the primary and secondary pages, so there is no tie to break and no reason to invent one. Each entry renders directly after its own place, so "22 primary schools in Brentwood" reads as part of Brentwood rather than as an unrelated link further along the row. The e2e journey now follows a phase link where the town publishes one and asserts it resolves. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq |
||
|
|
7f5f0fb676 |
feat(seo): link school pages into the location layer
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
W2 shipped ~5,000 place pages and nothing linked into them. The
location layer pointed down at school pages; school pages pointed
nowhere on the site. Their only anchor was the school's own website, so
the ~27k pages carrying most of the site's inbound authority passed it
straight off-site, and the new corpus was reachable mainly through the
sitemap.
Three things close the loop.
A reverse index over the place registry, places_for_urn, answers which
published places contain a school. Derived from the registry rather
than stored beside it, so the two cannot disagree about which places
exist: a place below the publish threshold is absent from the registry
and therefore never offered as a link. A test asserts that invariant
across every place in a built registry.
GET /api/schools/{urn} gains a `places` array carrying the name, count
and canonical path for each. It rides on the request the page already
makes, so the school page costs no extra round trip. The frontend types
it optional and defaults it to empty, because the two images deploy
separately and a frontend ahead of the API must render without it.
The page gains a "More schools near here" module and a BreadcrumbList.
The module orders narrowest first, because a reader on a school page
wants its town before its county, while the API orders widest first for
the trail. Anchors state their destination's size — "37 schools in
Brentwood" — which is worth more to a reader and a crawler than "see
more". With no published places it renders nothing rather than an empty
heading.
The trail is rooted at the homepage, not /schools. There is no /schools
index page; the location layer lives only at /schools/[place],
/schools/authority/[la] and /schools/near/[outcode]. Rooting it at the
bare path would have opened every breadcrumb with a link to a 404.
Outcodes are omitted from the trail: "schools near CM15" is a real
query and a useful link, but nobody navigates Essex to CM15 to a
school, and a breadcrumb claiming that describes a hierarchy the site
does not have.
School pages also now declare the School type rather than
EducationalOrganization, the parent type that covers universities and
nurseries alike.
The e2e journey asserts the round trip in both directions, following a
place page's own first school so the pair is genuinely related rather
than hardcoded. A one-way link is what already existed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
|
||
|
|
6fc7fce948 |
feat(flags): put /about and /blog behind flags, dark by default
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 20s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m21s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 12s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m39s
Both features ship dark. Neither is reachable in an environment where its flag is off, and every flag in this system starts off, so a deploy of this commit makes both disappear until someone turns them on deliberately. Two independent flags rather than one, which makes blog-on-about-off a reachable state. That state is the whole reason the change is larger than four notFound() calls: the blog leans on the About page for its author identity. The Person entity is anchored at /about#tudor, and that URL 404s while about_page is dark, so a post published in that state would claim an author resolving to nothing. Worse than having no named author. Both bylines fall back to unlinked text and the BlogPosting attributes to the publisher instead, so every combination of the two flags renders something correct. Gated: /about, /blog, /blog/[slug], the RSS feed, both footer links, and the matching content-sitemap entries. A sitemap must never advertise a URL that 404s. With both dark it emits a valid empty urlset rather than a 404, because robots.txt names it unconditionally. Not gated: /admin. Posts have to be writable before the blog is worth switching on, so flagging the panel would make the flag unflippable. getFlags takes a revalidate rather than always using the 300s constant. Reading a flag pins the calling route to the lowest revalidate among its fetches, and the footer links live in the root layout, so a naive gate there would have dropped every school and place page from a weekly cache to a 5-minute one. The layout passes 604800, the floor those routes already declare, and the build confirms all four SSG route families still prerender. The cost is one-way latency: pages follow a flip in minutes, footer links within a week. The e2e journeys follow the existing paired shape from the admission_distance flag: a lit journey and a dark one for each flag, reading state from whether /about and /blog respond rather than from /api/flags, which another journey asserts is not publicly reachable. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq |
||
|
|
2e9b5c83c5 |
fix(destinations): the masking pass can no longer exit unsafely
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 7m23s
Review found _mask_for_disclosure could return with its invariant broken and say nothing. add_companion only ever withheld a *published* cell, so a group with one suppressed category and every other one not_applicable — routine in special schools and AP, where few categories apply — left the loop with the lone suppressed cell still solvable. Reproduced on a nine-pupil cohort: one hidden cell, cohort served, residual intact. A disclosure-control pass that fails silently is worse than none, because everything downstream trusts it. The loop now runs until the invariant holds and escalates when no companion exists: the pupil group is dropped from the payload, and an empty block serialises as None so the section is absent rather than an empty shell. disclosure_invariant_holds() is exported so tests assert it directly instead of re-deriving it, and an exhaustive test sweeps all 81 suppression patterns of a four-category group. Also fixes a test that set up six measures and checked one: the loop was `for measure in ["school_sixth_form"]`. It now checks every measure, and against the real invariant — none hidden, or at least two, rather than "at least two", which the five published measures would have failed. No regression on real data: 262 mainstream secondaries, all-pupils bar still drawable on 94%, zero invariant violations, one disadvantaged group dropped by the new escalation. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob |
||
|
|
102397fe69 |
fix(destinations): withhold at the API, not just in the chart
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m14s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m49s
Code review found the disclosure the whole design was meant to prevent.
R1 was written as a rendering rule and implemented as one: canRenderBar
stopped the bar being drawn, but GET /api/schools/{urn} still carried the
cohort and every published category. cohort - sum(published) returned
Whitley Bay's withheld further-education figure exactly — 18 pupils — to
any caller, and the RSC payload put it in the browser too.
app.py already stated the principle for admission_distance: this endpoint
is public and unauthenticated, so a field left in the payload is a
published field. The same reasoning applies here and did not get applied.
_mask_for_disclosure now closes both identities before serialisation —
categories sum to the cohort, and disadvantaged + other = all — by adding
secondary suppression until every row and column hides none or at least
two. My first attempt picked the smallest published cell as the companion
and a new test caught it choosing a zero, which protects nothing: the
residual still resolved to 18. The companion must carry pupils.
DfE's own aggregates are no longer served. Nothing rendered them, and one
spanning a single suppressed component names it.
Cost, measured over 262 mainstream secondaries: the all-pupils bar
survives on 94% rather than 100%. Zero lone-suppressed groups remain.
The e2e helper now tells a missing feature apart from missing data: it
fails if the API serves no destinations key at all, and skips if the key
is served but the annual DAG has not populated the marts. Failing on the
second would redden the staging gate for unrelated commits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
|
||
|
|
c5719ef362 |
feat(destinations): serve destinations without closing the gaps
The serialiser carries status through and computes no totals of its own. The only aggregates in the payload are ones DfE published itself; whether showing one is safe depends on how many of its components are suppressed, which the frontend decides. The batch guard grows from six tables to eight. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob |
||
|
|
9a1f56c431 |
feat(places): say what each school is, not only how it scored
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m44s
The location tables carried one column: a percentage. A parent shortlisting from a town page is asking a different question first — does it take my child's age, is it a faith school, does it have a nursery — and the page could not answer any of it. Primary tables gain Ages, Religious character, Nursery and Constituency; secondary tables the same minus Nursery, which is a question about a different intake. An all-through school renders in both groups, so its nursery shows under primary alone. The measure moves to the second column rather than the last. Six columns overflow a phone and .tableWrap turns that into a horizontal swipe; with the measure last, the one number the page exists for is the one scrolled off the screen. Cell rules are the ones the school page already uses, so the two surfaces cannot disagree about the same school: "Does not apply", "None" and "Not applicable" all read as no religious character, and the en-dash age normalisation moves into formatAgeSpan, which formatAgeRange now delegates to. Backend: nursery_provision and parliamentary_constituency were not in the place response. Both are optional GIAS mart columns that data_loader degrades to NULL, and the `in rows.columns` guard keeps a mart the pipeline has not rebuilt working. Also fixes a live bug on the same line: SCHOOL_COLUMNS already ends with latitude and longitude, and the endpoint concatenated them again, so pandas dropped one of every duplicated pair and warned "columns are not unique" on each request. Ordered de-duplication removes the warning and the silent drop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM |
||
|
|
0fa1a292c7 |
fix(api): bound what a forged CF-Connecting-IP can buy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Code review, both findings valid. The design doc claimed Cloudflare "replaces the header, so a browser cannot forge it", and that only the X-Forwarded-For fallback was forgeable. That is true only for traffic that actually passed through Cloudflare, and nothing in this process can verify that it did. Reaching the origin directly, both headers are equally attacker-controlled — and rotating CF-Connecting-IP mints a fresh rate-limit bucket per request, defeating per-client limits on every endpoint including the DataFrame-heavy /api/schools. Against abuse that is worse than the shared bucket it replaced, which at least capped everyone together. So the ceiling comes back. I dropped it earlier arguing it belonged at Cloudflare; that argument assumed the keying was sound, and it is not. GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window against a total, independent of client identity, outermost so it refuses before any work happens. Written by hand because slowapi cannot express a global cap: default_limits and application_limits are both keyed by key_func, and the latter needs middleware this app does not install. It does not make the header trustworthy — it makes trusting it survivable. The real fix is Authenticated Origin Pulls or an origin firewall, now documented in DEPLOY.md as the open gap it is. 127.0.0.1 is exempt: the healthcheck curls localhost from inside the container, and starving it would restart the container and turn a load spike into an outage loop. Keyed on the peer address, never the Host header, which the caller sets. Second finding: suggest_schools_typesense promised "never raises" while the parsing loop sat outside the try, so int(None) on a malformed document would have made a keystroke a 500. The loop now skips bad rows rather than dropping the whole list — and a hit with no document no longer becomes a suggestion pointing at /school/0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
28cf0a342c |
feat(suggest): wire autosuggest into the search box behind a flag
Off means off — no combobox role, no listener, no fetch. A test asserts the absence of the request, not just the absence of the dropdown, because a hidden-but-fetching control would still be spending the rate limit on a feature nobody can see. Enter with no active option falls through to the form's submit handler and searches the typed text exactly as before. The existing behaviour is preserved, not replaced, and that has its own test. Suppressed once the value parses as a postcode: the box takes a name OR a postcode, and suggesting schools during postcode entry fights the user. .omniBoxContainer gains position: relative — the dropdown is absolutely positioned and without it would have anchored to the page instead. Four render sites, all wired: page.tsx renders HomeView in the success path AND the catch fallback, and HomeView renders FilterBar as hero AND sticky. Missing any one would make the flag silently do nothing somewhere. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
1a6d349dad |
feat(suggest): GET /api/suggest, cacheable and DataFrame-free
A dedicated endpoint rather than a mode of /api/schools, because that path filters and sorts 25,000 pandas rows per query while holding the GIL — affordable once per search, not once per keystroke. A test asserts the distinction directly by making load_school_data raise and requiring the endpoint to answer anyway. Nothing errors on ordinary input: a short query, no matches, or Typesense being down are all 200 with an empty list. Cached deliberately. Prefix queries repeat enormously across users and school names change once a year, so s-maxage plus the existing ETag middleware turns most keystrokes into 304s. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
75d3534d82 |
feat(suggest): Typesense rows for autosuggest, no DataFrame
search_schools_typesense returns URNs, which forces the caller to hydrate from the 25,000-row in-memory frame. Every field a suggestion needs is already in the Typesense document, so this returns documents and the caller needs no pandas at all — the difference between a query that can run per keystroke and one that cannot. Never raises. Typesense unreachable or erroring gives an empty list, because a dropdown that quietly stops appearing is the right failure for a keystroke path. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
ff041544f2 |
fix(api): rate-limit per caller, not per proxy
The limiter keyed on request.client.host, which in staging and prod is the Next container — the backend has no published ports and nothing else can reach it. So every browser user on the site shared one 60/minute bucket per route. Measured against staging: 70 concurrent requests to /api/schools returned exactly 60 OK and 10 refused, from one machine. CF-Connecting-IP first. Cloudflare fronts both environments and overwrites any client-supplied value, which a parsed X-Forwarded-For chain does not guarantee. The XFF fallback is forgeable only from inside the Docker network. Named rather than hidden: the shared bucket was an accidental global throttle on a single-process backend, and correct per-user keying removes it. A real global ceiling belongs at Cloudflare, which is already in the path; slowapi cannot express one without a second Limiter and middleware this app does not install. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
c3ba7aae0d |
feat(flags): ship last-distance-offered dark behind a flag
One gate, at the source. The frontend needs no change: DistanceSection
already returns null when distance_m is missing, and the admissions
block already conditions on (admissions || admissionDistance). Only 57
local authorities publish cut-offs, so the off-path is the commonest
path on the site and is well covered already.
Absent, not null. /api/schools/ is public and unauthenticated, so a
field left in the payload is a published field — the reasoning already
recorded in
|
||
|
|
c30ad1db07 |
feat(flags): serve /api/flags, and keep the public proxy off it
The endpoint and its exposure control ship together on purpose. The moment /api/flags exists, app/api/[...path] forwards it — and the response names every unreleased feature the codebase knows about, along with whether it is on. Publishing that is the opposite of shipping dark. Denied on an exact first-segment match, not a prefix, so /api/flagship does not go down with /api/flags. Next reads the endpoint server-side over the Docker network, which never transits the public proxy. jest.setup.js now guards its browser globals. It runs for every suite, including the one that declares @jest-environment node to exercise the route handler — NextRequest needs Fetch API globals jsdom lacks, and there is no window there to define matchMedia on. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
7424cef7c6 |
feat(flags): the registry and a fail-closed Unleash client
Unleash holds flag state; it does not hold the list of flags. REGISTRY is that list, because the SDK evaluates an unknown flag to False and without a registry that is an undeclared False — indistinguishable from a typo in a flag name. Fail-closed throughout, and never raises: an unset UNLEASH_URL, an unreachable server, a client that throws, an undeclared name — all False. A flag layer that can 500 a request path or stop the API booting is worse than one that is switched off. Every flag defaults to False, with no per-flag override, because a flag that defaults on is a kill switch and this is deliberately not one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
d1358cc00f |
fix(places): phase links must stay in their own namespace
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m37s
Every place page built its phase links as /schools/[slug]/[phase], the shape that belongs to towns alone. On an authority page that pointed into the town namespace. For 87 of the 151 authorities the target does not exist and the link 404s; for the other 64 it resolves to the town of the same name — a different set of schools, which is precisely the near-duplicate the two namespaces were introduced to prevent. On an outcode page it 404s outright. Two causes behind it, both a rule written twice and inherited by only one of the places that needed it. The authority phase route was in the spec and dropped by the plan, which built the three bare routes and no fourth. The sitemap is generated from the place registry, which was right about them all along, so 302 authority phase URLs have been submitted to Google and every one 404s. Adding the route makes the sitemap true and serves a real query — admissions are authority-run, so "primary schools in Kent" is how a parent searches before they have settled on a town. The outcode variants were the opposite: the registry computed phases for outcodes although the spec gives them no route, and the sitemap knew to skip them while the API did not. The registry now decides alone, and the sitemap's duplicate of that rule is gone. Also: an authority under the five-school threshold has no page, so the API sends a null slug for it and the page names it without linking. Two English authorities are in that position. It was unreachable in today's data — verified across the EC and TR outcodes — but the thin place redirect would have sent a reader to a 404 the year it isn't. The e2e journey now walks every /schools link a page of each family emits and requires a 200, which is the check that was missing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
9cc87c41bb |
fix(places): a phase page needs results, not merely publishable schools
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m4s
Asked where schools with no results should sit in an alphabetical list, and found that some pages were almost entirely made of them. The per-phase threshold counted schools that were publishable — a result OR an Ofsted grade — while a phase page exists for its results column. /schools/kent/primary published with none of its five rows carrying a result; Minehead had one of seven, Buntingford one of five. Forty-four phase pages were majority-blank. It is the same rule as "no page without a local average", which was written into the spec as a thin-page control and never extended per phase. The threshold now counts schools with a result for that phase. It gates whether the page exists; it does not filter rows — a page that publishes still lists every school of the phase, because someone looking up a school by name has to find it whether or not it published results. 126 of 1,012 variant pages stop publishing: 62 primary, 64 secondary. Every one of them was a table with too little in it to be worth a page. The ordering itself is unchanged: pure A-Z, blanks interleaved. A school sits where its name says it does, and at roughly a tenth of rows that reads fine. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
8967966eef |
feat(places): list schools alphabetically on place pages
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m21s
Someone on a place page is usually looking for a school they can name, so the order should serve scanning for it rather than ranking. /api/rankings keeps its league-table ordering; this is a place-page decision, not a site-wide one. Sorted case-insensitively, or a capitalised name would sort ahead of every lowercase one. The change made five pieces of copy untrue, so they go with it. The phase variant titled itself "— Ranked", and all four route families described themselves as "ranked by SATs and GCSE results". A page that opens by claiming an order it does not keep is worse than one that claims nothing. The ItemList markup carried `position` with no declared order, which reads as a ranking. It now declares ItemListOrderAscending, so the structured data says what the table does. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
bb2f7a5841 |
fix(places): address review, and merge places GIAS spells more than one way
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m34s
Two findings from review on #121, plus a third the review prompted. The cap at three authorities silently dropped the fourth in exactly the case where the information matters most — a genuinely fragmented place — and contradicted the stated goal of naming every authority a place sits in. It is gone. The share rule was always the real limit and already bounds the list at ten. Measured against the live corpus, one town would have been truncated today: LONDON, split evenly between Hackney, Lambeth, Westminster and Lewisham. parent_authority used mode() while authorities used value_counts(), and on an exact tie pandas does not guarantee the two pick the same name, so the 301 could have pointed somewhere other than the authority named first on the page. The parent is now derived from authorities[0]: one computation, one answer. It also inherits the sentinel filter, so a place can no longer redirect to /schools/authority/does-not-apply. Chasing the truncation case surfaced a worse bug. Places were grouped by raw town value, but the registry is keyed by slug, and GIAS spells the same place several ways. Five town slugs come from more than one spelling: "London" (1,819 schools) and "LONDON" (12) both slugify to `london`, so the later group simply overwrote the earlier one — /schools/london could have shown twelve schools, silently, depending on row order. Weston-super-Mare was split 14/19 across two spellings and Newcastle-under-Lyme across three. Grouping is now by slug, and the display name is the most common spelling. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
1cb5314c53 |
feat(places): name every authority a place sits in
SW19 is mostly Merton but partly Wandsworth, and the page said only Merton. The cause was one field doing two jobs: _parent_authority takes the modal authority, which is right for a 301 target and wrong as a statement about where a place is. This is not a corner case. A quarter of viable outcodes (425 of 1,760) and a third of viable towns (263 of 783) cross an authority boundary — Bedford the town spans Bedford and Central Bedfordshire. Place now carries `authorities`, every authority holding at least a tenth of the schools and at least two of them, largest first. parent_authority stays single and unchanged, because a redirect still needs one target. The share threshold exists because GIAS carries postcode errors: EN6 lists two Shropshire schools among fourteen in Hertfordshire, and a bare "any authority present" rule would print those as though they were real. A place too small or too fragmented to clear the threshold still names its largest, so the page never goes silent about where it is. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
6f749ed21f |
fix(places): submit and link the phase variants
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m39s
/schools/[place]/[phase] shipped as routes but reached nothing. The sitemap emitted one URL per registry entry and the registry had no phase dimension, so ~950 pages were absent from every sitemap — and PlaceView did not link them either, leaving them reachable by nothing at all. That is the query shape the baseline actually showed: 'primary schools in beccles', 'secondary schools in brentwood'. Publishing the routes without a path in meant building for the demand and then hiding from it. Place now carries phase_urns so the per-phase threshold can be applied without re-querying, the sitemap emits a variant wherever a phase clears the threshold on its own, and the API exposes the qualifying phases so the place page links only variants that exist. Outcodes are excluded: nobody searches 'primary schools in SW11' and those routes do not exist. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
d3c63ccc6d |
fix(places): a locality collision must not break the sitemap
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 34s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 34s
Sitemap regeneration failed on staging. 'richmond' in the curated locality list collides with the GIAS town Richmond in North Yorkshire (37 schools), the registry raised, and the admin endpoint 500d — taking down sitemap generation for all 25,000 school pages over one bad row of curated data. The guard now skips the colliding locality and logs an error. Skipping still achieves what the guard was for — a locality never silently shadows a town — without letting curated data break the site. That matters beyond this bug: GIAS town names change with no code change here, so a raise could fire spontaneously in production later. Also removes four localities that were London boroughs rather than districts. Hackney, Islington, Greenwich and Ealing are local authorities with 104, 72, 108 and 115 schools and already have authority pages; a locality defined by two or three outcodes would have been a partial near-duplicate of one — the thin-content failure the two-namespace design exists to avoid. A test now guards the whole borough list. Validated against the live corpus: 15 localities, no town collisions, no authority duplicates, all 15 clear the threshold. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
42138fc402 |
feat(places): submit place and outcode sitemaps
Separate children per family so Search Console reports the location layer's indexation apart from the school pages' — which is the point of the index built in W1, and the number the stop condition watches. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
c5af476213 |
feat(places): /api/places registry and place detail endpoints
The registry is cached for the process and reset by the same admin endpoint that rebuilds the sitemaps, so places and sitemap always describe the same corpus rather than drifting apart. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
de853b90b3 |
feat(places): London localities and postcode districts
The GIAS town field puts 1,819 London schools under the single value 'London', so it cannot answer 'schools in Battersea' — a query that appears in the baseline. No single field can: parliamentary constituency gives Battersea but not Canary Wharf, admin_ward gives Canary Wharf but not Battersea, and neither gives Clapham or Shoreditch. So a locality is curated, defined by the postcode districts it covers, which needs no new ingestion. A locality may not shadow a published town: the registry raises rather than silently costing a page that carries real demand. One below the threshold is logged rather than raising, because a locality can legitimately be too small. The pipeline seed mirrors the module, with a test guarding the drift — the same arrangement gias_codes has, and for the same reason: the backend image does not contain pipeline/. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
759d9f5cea |
feat(places): registry of towns and authorities
Two namespaces because 67 town names collide with an authority name and neither set contains the other — postal towns cross authority boundaries, so Bedford the town holds 104 schools against the authority's 86. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
07c97a46c5 |
fix(seo): a school is publishable on any year's results, not the latest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 53s
_school_sitemap_rows tested only the latest year's row, which quietly dropped every school with results in its history but a null row for the most recent year — a school that stopped reporting, or whose figures were suppressed for small-cohort disclosure. The Mallard Academy (150367) is the case that caught it: real KS2 results for 2015-16 through 2018-19, then null rows from 2022-23 on. Its detail page shows all four years; the sitemap omitted it. Sampling 40 of the 2,206 excluded schools found 4 like this, so roughly 220 real pages were being withheld. Publishable is now a property of the school, computed across every row, while lastmod still comes from the latest row so the most recent Ofsted date wins. The field list is a module constant shared with _has_publishable_data so the two checks cannot drift. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
2208ad93c1 |
feat(seo): split the sitemap into a per-family index
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m2s
Search Console reports coverage per submitted sitemap, so one file per page family is what will make W2's location pages measurable when they land. The index's lastmod is generation time, which is the correct semantic there — unlike on a <url>, where it would be a claim we cannot support. Children sit under /sitemaps/ because Next only treats a whole bracketed path segment as dynamic; a route folder named sitemap-[...parts] would be read as a literal static segment and never match. Confirmed by the build output, which lists /sitemaps/[...parts] as a dynamic route. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
24f3cb4c65 |
fix(seo): submit only school pages that have something to show
Drops the schools with neither results nor an Ofsted grade, adds /admissions which was never listed, replaces the invented priority and changefreq with a lastmod taken from each school's Ofsted date. lastmod is omitted where no date is known rather than defaulted to now. An always-now lastmod is a claim Google learns to distrust; absent honestly means unknown. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
5ddc7314fd |
fix(seo): canonicalise on the www host, which is the one that serves 200
The apex 301s to www at Cloudflare, but metadataBase, the school-page canonical, robots.txt's Sitemap: line and the sitemap's own <loc> entries all named the apex. Every one of those pointed Google at a redirect. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj |
||
|
|
c9a1892bfb |
feat(admissions): publish the latest cut-off only, holding history back
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m17s
Earlier years are to become a paid feature, so they stop being published.
The load-bearing part is that this is a change to the API, not only to the
page. /api/schools/{urn} is public and unauthenticated: leaving
admission_distance_history in the payload while declining to render it would
have handed the whole record to anyone who opened the network tab. It is
withheld at the source, and the page follows.
Nothing changes upstream. The tap, the plausibility band and
fact_admission_distance are untouched and still load every published year, so
restoring history for entitled callers is a change to one function in
data_loader rather than a re-collection.
What the reader now gets is the latest figure on the Admissions tile, and a
Distance section that answers the question the number alone cannot: whether
their own address falls inside it. Retitled to "How far away are you?", which
is what it now does — the previous title described a record that is no longer
there.
Removed with the history: the trend chart, the year-by-year table, the
per-year verdict strip, the trend summary and the coverage note, along with
their CSS. The section goes from 743px to 417px.
One consequence worth naming. A run of years used to soften a single close
call — a home just outside one year's cut-off was usually inside another. With
one year published, the "too close to call" band is the entire safety margin
between a parent and a place they do not have, so the verdict now names its
year, and the three outcomes are tinted apart rather than distinguished by
wording alone.
The existing stylesheet test earned its keep here: the three verdict classes
were referenced before they were written, and it caught them. Unstyled, a
"beyond the cut-off" result would have been indistinguishable from an "inside"
one — the exact failure the longhand class map was written to prevent.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
|
||
|
|
a72323874f |
feat(admissions): add the cut-off history, map and postcode check
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m39s
Completes the last-distance-offered feature against the mockup: the year-by-year record, the same numbers drawn over real streets, and the reader's own address measured against them. Serving the history The first cut deliberately served only the latest year, because a plain series would draw a trend line straight through gaps that are absences of publication, not of a cut-off. That reasoning is answered rather than abandoned: cutoffYearRows classifies every year in the span, and the chart breaks the line rather than interpolating across it. A missing year is not one fact but three. It may be unpublished; it may be a year the school was not oversubscribed; or there may be no record at all. Collapsing them into "no data" throws away the reassuring case and hides the important caveat, so each is stated in words in the table. The claim is held to what the data supports. fact_admissions.oversubscribed compares FIRST PREFERENCES against places, which does not establish that every applicant was offered one — so the copy says "places available on first preferences" and a test asserts the stronger claim never appears. The trend summary is not a verdict It names both endpoints and their years and lets the reader conclude. It is withheld below four published points, and a swing under a tenth of the earlier figure is reported as "broadly the same" rather than dressed up as a direction. The postcode check This is the only place on the site that answers a question about a family rather than a school, so most of the care went into what it refuses to say. postcodes.io returns a centroid covering roughly fifteen addresses, which against a 500 m cut-off is a fifth of the whole distance — so a margin inside 100 m returns "too close to call" rather than a place a family does not have. Unpublished years count as unknown, never as a pass. The limits are stated before the check is used, not revealed with the answer. The postcode is geocoded in the browser and never stored. Both templates Banded and selective secondaries are exactly where this matters most, so the detail is shared. The primary page gives it a third tab; the secondary page is one flat panel by design and renders it inline. Absence is explained rather than reported. A selective school's missing figure is explained by how it admits; a consistently undersubscribed school reads as good news. Also makes the batch loader's test double honour ORDER BY. It was a no-op, so "latest row per URN" was really "first row in the fixture" and the test would have passed with the sort reversed or removed. Verified: 214 frontend tests, 54 backend, 45/53 e2e green against staging (the 8 cut-off journeys skip until the DAG runs). Rendered offline against the real compiled CSS in both themes and at 390px; every new surface clears WCAG AA, measured on composited pixels. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE |
||
|
|
88c653215d |
feat(admissions): show the last distance offered where councils publish it
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m12s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m7s
Adds the cut-off distance a parent actually asks about — "how close do we
need to live?" — end to end: a Singer tap, dbt staging and mart models, an
Airflow DAG, and a tile on both detail templates. 3,597 schools across 57
local authorities carry a figure; the rest are unchanged.
There is no national source for this. Each LA publishes its own cut-offs in
its own format, and the collected CSV is transcribed from PDFs, spreadsheets
and web pages — so most of the work here is deciding what is safe to show.
Data
* tap-uk-school-distance loads the CSV verbatim into raw. Keyed on
(urn, year, school_name), because school_name carries the admission
route: (urn, year) alone collides on 118 keys and a reload would have
silently dropped every band but one.
* stg_school_distance applies a 25 m – 25 km plausibility band. The source
contains 0.0-mile rows (published where a school filled on a higher
criterion), 1-metre cut-offs, and one reading 533 miles — ~4% of rows,
all of which would put a visibly wrong number on a live page.
* fact_admission_distance collapses routes to one row per school per year
using the furthest, and keeps route_count so the page can say the figure
is the widest of several bands rather than the one for a given child.
Serving
* Kept out of fact_admissions: that mart is EES-derived and near-complete
for England, this one covers 57 LAs, and the two refresh independently.
* Latest year only. Coverage is ragged — a school may have 2021 and 2026
and nothing between — so a history array would invite a trend line drawn
through gaps that are absences of publication, not of a cut-off.
* The Admissions section now renders on either source. 3% of the schools
that render have a cut-off and no EES admissions row, and gating on
admissions alone would have hidden the figure on those pages.
Interface
* The year travels with the figure everywhere it appears; a cut-off
detached from its admissions round is not a fact about anything.
* "Not a fixed catchment — it moves every year" sits under every instance,
because that is the inference a parent will otherwise draw.
* Replaces a hardcoded "Historical distance cut-off data is not available
for this school" that appeared on every secondary page, including the
ones whose council does publish it. The absence is now stated only when
it is real, and names the authority that would hold it.
The tint costs the muted tokens their AA margin: measured on the composited
backdrop (not the computed one, which reports the untinted card), --text-muted
falls to 4.09:1 in dark theme. The tile uses --text-secondary instead — 6.50:1
dark, 6.60:1 light.
The DAG is manual, like the other annual ones: councils publish on allocation
day, each on its own timetable, so there is no date worth scheduling against.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
|
||
|
|
a102508ef1 |
fix(data): strip any table alias in missing-column matcher, not just s.
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m1s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 42s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 9s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m49s
The graceful-degradation fallback keys off the column named in a Postgres UndefinedColumn error, but the matcher only stripped an `s.` alias. The two new dim_location columns (county, parliamentary_constituency) are selected via the `l.` alias and Postgres reports them unquoted as "column l.county does not exist" — which the old regex failed to match at all, returning None. If dim_school is rebuilt (telephone/nursery present) but dim_location is not yet (county/parliamentary_constituency missing) — plausible since they are independently-rebuilt dbt models — the fallback branch never matched and load_school_data_as_dataframe() returned an empty DataFrame, showing zero schools sitewide instead of degrading those columns to NULL. Generalise the alias prefix to `\w+\.` and cover the l.-qualified case in tests. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
0186227ced |
feat(detail): surface GIAS identity/contact details, drop unwired sections
PR Checks / Frontend Typecheck + Tests (pull_request) Canceled after 58s
PR Checks / Backend Smoke (pull_request) Canceled after 0s
PR Checks / Build Backend (no push) (pull_request) Canceled after 0s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 0s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
Add seven school-identity fields to the detail header (both primary and secondary views): age range, religious character, nursery and sixth-form indicators as chips; telephone (tel: link), county and parliamentary constituency as header details. religious_denomination, age_range and has_sixth_form were already served; telephone, nursery_provision, county and parliamentary_constituency are newly wired through the marts query (with a NULL fallback for un-rebuilt marts, mirroring has_sixth_form) and the school_info API response. Remove three UI sections the backend never populated (always null): Year 1 Phonics, the SEN "types of additional needs" breakdown, and the average class-size card — along with their now-dead props, route plumbing, and the SenDetail/Phonics types + class_size_avg field. Extend the e2e detail journey to assert the Phonics section is gone and the new header fields render when the record carries them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> |
||
|
|
8a9ba30cc2 |
fix(detail): compare each SATs bar to its own national benchmark
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
The KS2 SATs chart drew a single national-average line spanning the full height of each subject's chart area, positioned at the national *expected* value. But the area stacks two bars — Expected and Exceeding — and the higher-standard/greater-depth national is a very different, much lower figure (e.g. reading higher standard ~29% vs expected ~75%). So the line crossed the Exceeding bar at the wrong place, making every school's exceeding result look far below national when it wasn't. The per-subject higher-standard nationals were already computed in the fact_ks2_national_averages mart; they just weren't serialized. Fix: - backend: add reading_high_pct, writing_gd_pct (writing = greater depth) and maths_high_pct to the national-averages payload. - SchoolDetailView: pass a nationalExceedingPct per subject, mapping writing to the greater-depth figure. - SatsChart: replace the single full-height line with a national marker on each bar's own track (coral tick + "nat X%" in the bar header), so Expected and Exceeding each sit against the correct benchmark. KS2 only; the secondary Attainment 8 chart already uses one line for one measure and is untouched. Verified: tsc --noEmit, next build, and backend pytest (national averages marts, incl. a new test guarding the per-subject nationals). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB |
||
|
|
5ec4f3f7cd |
fix(list/map): badge report-card schools as Report Card, not their carried-forward grade
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m48s
The search-result cards and map pins keyed report-card detection off
ofsted_framework === 'ReportCard', but the API sets ofsted_framework to the
raw event grouping ('Schools - S5'); worse, ofsted_grade (the carried-forward
legacy grade) was checked first and won. So report-card schools were badged
by their old grade — Barclay's pin/card read 'Outstanding · 2021' instead of
'Report Card · 2026'. Same root cause as the detail-page fix, different
surface.
Expose ofsted_rc_date on the list serialization (the report-card inspection
date, non-null only for report cards) and make both badge builders treat a
present rc-date as winning over any grade, using its year. Removes the dead
framework === 'ReportCard' branches.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB
|
||
|
|
5944d88f0b |
fix(compare): expert sign-off must-fixes — phase-matched admissions, Ofsted sentinel codes, selective-school copy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m53s
M1: admissions rounds are now selected by the active phase tab (admissionsForPhase) — an all-through school's Year 7 round no longer masquerades as Reception odds beside pure primaries; honest per-cell and section fallbacks name the round (Reception / Year 7). M2: Ofsted sentinel codes (9 = not applicable) never render as judgement chips, and the sixth-form judgement — previously dropped — now renders for schools that have one. M3: 'What this means' is phase- and type-aware: selective schools get entrance-test framing, secondary faith schools a faith-criteria note, and the primaries' distance template never appears on the secondary tab (admissions_policy now exposed in compare school_info). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB |
||
|
|
c9e324635b |
fix(data): official DfE KS4 national headline averages; drop mislabelled computed means
New ees_ks4_national stream ingests the EES 'National characteristics summary data' series (England, state-funded, all pupils). The old mart's unweighted school means were 7-15 points off every headline measure and produced an impossible national Progress 8 (-0.27). The API's computed fallback is gone too: the footnote calls these figures official, so an unbuilt mart now yields an empty series, never a stand-in. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB |
||
|
|
1d855f3c17 |
fix(compare): census-sourced FSM/EAL benchmarks; never fall back across measure definitions
The FSM chip anchored against disadvantaged_pct (a different measure, FSM6+CLA) whenever fsm_pct was null — which it always was, since the performance df has no fsm_pct. New fact_census_benchmarks mart supplies pupil-weighted FSM/EAL means per phase; the KS2-column medians that produced a bogus 50% 'secondary disadvantaged' anchor are gone. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB |
||
|
|
9773483221 |
fix(compare): date report cards with their own inspection date, never the legacy one
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0146VHeLAWjDVE2B5uU67jCB |
||
|
|
b4b0249a06 |
Fix Ofsted transitional inspections, phase tab exclusions, and FSM benchmark comparison
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m9s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 10s
|