Compare commits

...
Author SHA1 Message Date
TudorandClaude Opus 5.5 d1688ac150 copy(web): replace em dashes in public copy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m20s
Rewrites every visitor-facing string that used an em dash, choosing the
punctuation for what the dash was doing: a colon before a list or
explanation, a comma for an aside, a full stop between two thoughts,
parentheses for an aside mid-sentence. Covers page titles and meta
descriptions, the home and admissions guide copy, school page headings
and notes, the compare page, metric labels and tooltips.

Two rewrites also fix the sentence around them: the closure banner no
longer repeats "proposed for closure", and the cut-off caveat's list of
priorities now parses.

A lone dash marking a missing value in a table cell stays: it is a data
convention, not prose. A Jest guard walks the source with the TypeScript
parser and fails on any other em dash in a string or JSX text node.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-24 22:04:27 +01:00
tudor 271ffe92d4 Merge pull request 'fix(web): lift the mobile jump sheet above the bottom tab bar' (#154) from fix/jump-sheet-under-bottom-bar into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m23s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 14s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 28s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 3m0s
Reviewed-on: #154
2026-09-22 19:55:00 +00:00
TudorandClaude Opus 5 0571d1c0ff fix(web): stop the sheet-open rule stealing .sectionNav's layout
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
Review catch, and a bad one: the previous commit anchored its insertion on
`padding: 0.5rem 0.75rem;` and closed .sectionNav there. Everything that
followed in the rule — margin-bottom, box-shadow, display: flex, align-items,
gap — was orphaned into .sectionNavSheetOpen, which is only applied while the
mobile jump sheet is open.

So the sticky nav lost its flex layout, spacing and shadow in the closed
state, which is virtually every page view on every school detail page. A
site-wide regression introduced by a fix for one mobile menu.

Redone by anchoring on the complete rule, closing brace included, so nothing
can be orphaned. .sectionNav is now byte-identical to main and the diff is
purely additive; .sectionNavSheetOpen carries the z-index and nothing else.

The staging experiment that validated this fix set nav.style.zIndex = '1100'
with every other declaration intact, so it was always testing this version
rather than the broken one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 20:43:58 +01:00
TudorandClaude Opus 5 180d6e9b3e fix(web): lift the jump sheet above the bottom tab bar
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 37s
On mobile the last item in "Jump to section" was painted over by the fixed
bottom tab bar and could not be tapped. Reported as nearby schools missing
from the menu; it was there, underneath the bar.

The sticky nav sets `position: sticky` with `z-index: 10`, which makes it a
stacking context. The sheet's own `z-index: 1600` therefore orders it only
inside that context — against the tab bar (z-index 1000) the nav's 10 is what
counts, so the bar wins. Verified on staging: every menu item returns itself
from elementFromPoint except the last, which returns the tab bar.

Latent rather than new. With five sections the list stopped just above the bar;
"Nearby schools" made six, and the sixth is the first to reach it. Any section
added later would have done the same.

Lifted only while the sheet is open, and only to 1100 — above the bar, below
the comparison toast (2000), the fullscreen map (5000) and the info popover
(9999). The backdrop rises with it, so tapping over the bar now dismisses the
sheet instead of navigating away.

The journey asks what a thumb asks: for each item, whether it is the topmost
element at its own centre. A bounding-box check cannot see this — the item is
in the viewport and the right size, just underneath something. Confirmed to
fail against current staging, naming "Nearby schools", before the fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 20:35:25 +01:00
tudor 80f405123e Merge pull request 'test(e2e): wait for the carousel's smooth scroll to settle before measuring it' (#153) from fix/nearby-journey-waits-for-smooth-scroll into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m25s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 28s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m53s
Reviewed-on: #153
2026-09-22 14:55:16 +00:00
TudorandClaude Opus 5 077aca6008 test(e2e): wait for the carousel's smooth scroll to settle before measuring it
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 32s
Second failure of the same assertion, and the first fix addressed a real but
different problem. This is the one that was actually producing the number.

The arrows scroll with `behavior: 'smooth'`. The test polled for "has it moved
at all" — satisfied 50ms in, at 13px of a 1300px journey — then recorded the
offset, clicked, and recorded again. The row was still travelling throughout,
so the delta it measured was the tail of the arrow's animation, not the effect
of the selection. Hence a deterministic 659, roughly half of the 1317 this
school's row scrolls.

Measured against staging rather than reasoned about: the animation runs about
700ms, and the samples are in the helper's comment.

settledScrollLeft waits for two identical readings before trusting one. The app
was never at fault — driving staging by hand, the offset holds at 1317 across
the selection, exactly as intended.

Verified against staging both ways: main's version of this test fails there,
this version passes, along with all three mobile widths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 15:13:32 +01:00
tudor 9dba5ff1ff Merge pull request 'test(e2e): fix the nearby-schools journey clicking a card it scrolled past' (#152) from fix/nearby-scroll-journey-clicks-offscreen-card into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m23s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 25s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m49s
Reviewed-on: #152
2026-09-22 13:50:25 +00:00
TudorandClaude Opus 5 f530a912bc test(e2e): stop the nearby-schools journey clicking a card it scrolled past
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 21s
The assertion that adding to compare does not reset the carousel failed on
staging: 537 → 2. The app was not at fault. Playwright scrolls a target into
view before clicking, and the test clicked the FIRST card's button after
paging the row to the end — so Playwright scrolled the container back to the
start, and the assertion measured that.

Reproduced on a static page with no React on it: a snap scroller at 615,
Playwright clicks the off-screen first card, scrollLeft becomes 2. Scroll-snap
was ruled out first — mandatory, proximity and no-snap all behave identically
when the button mutates in place.

Now clicks the last card's button, which is visible at the end of the travel,
and allows a few pixels for snap and sub-pixel adjustment while still failing
on a reset to the start. The property was never actually under test before;
it is now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 14:19:00 +01:00
tudor 029fe8d8a6 Merge pull request 'fix: order nearby schools by distance, not by how alike they are' (#151) from fix/nearby-schools-order-by-distance into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m24s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 30s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m51s
Reviewed-on: #151
2026-09-22 13:09:38 +00:00
TudorandClaude Opus 5 cd1c5d1e1a docs: revise the spec to the design that survived staging
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 57s
The spec described the tier system as the design of record. It is gone, so
the document was describing something the code deliberately does not do.

The revision note and the "why not, having built it the other way first"
passage are kept rather than overwritten. The mistake is the instructive part:
treating a preference as a constraint inverted the ranking, and the stopping
rule added to prevent weak distant matches is what guaranteed six Catholic
schools and no community school down the road. A spec that quietly presents the
second design as the plan teaches nobody why the first one failed.

The mockup link is annotated as one revision behind rather than silently left
to look current.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
TudorandClaude Opus 5 cd6a45bf7d refactor: rename similar → nearby, so the code says what the section does
The section ranks on distance and is headed "Other schools nearby", but every
identifier still called it "similar" — the exact drift that leaves a later
reader trusting a name over the behaviour.

Mechanical: files, the module, the payload key, the type, the components, the
prop. No behaviour change; the suites are unchanged in count and still green.
Free to do now because #150 has not merged, so the payload key rename needs no
lockstep deploy. Uses of "similar" that are ordinary English — progress
measures compared to similar pupils, and unrelated comments — are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
TudorandClaude Opus 5 5c0ccc693d fix: order nearby schools by distance, not by how alike they are
Reported from staging: a Catholic primary showed six Catholic primaries, none
of them close enough to be a real option, and omitted the community school
down the road.

Three causes, compounding. Ranking put tier before distance, so a faith match
at 2.9 miles outranked a community school at 0.3. The ENOUGH=3 stopping rule —
added so a cap of six would not drag in weak distant matches — filled the row
from the best tier before it ever widened, which is what made every card
Catholic. And a 3-mile tier-1 radius is sane for a secondary and most of a city
for a primary, whose catchments are routinely under a mile.

The premise was backwards. For a parent, distance is a constraint and intake is
a preference; a school beyond a primary catchment is not a weaker option, it is
not an option. So distance now decides the order and nothing else does. The
hard filters are untouched — they were always where the defensibility lived.
Similarity survives as chips on the card: reported, so a reader applies their
own weighting, rather than ranked, so we apply ours for them.

Reach is capped per phase (primary 2, secondary 6, post-16 10) as a sanity
bound, not a target: ordering already handles density, so the cap only decides
what happens where an area is sparse. A primary with nothing inside two miles
now renders no section, which is the honest answer.

Deleted: the tier system, the stopping rule, the tier-dependent lede, the
`tier` field, the tier-3 fallback chip and its style. select_similar also stops
taking is_secondary — it reads the phase from the subject's own row, so no
caller can hand it one that disagrees with the data.

The heading is now "Other schools nearby". The hard filters still guarantee a
comparable set, but nothing ranks on likeness, so the heading no longer says it
does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
tudor 151cf4bc80 Merge pull request 'feat: similar schools nearby on the detail page' (#150) from feat/similar-schools-nearby into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 44s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m25s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m14s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 6s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 42s
Reviewed-on: #150
2026-09-22 05:53:06 +00:00
TudorandClaude Opus 5 83dc5ae5dc docs: correct the metric rule for "16 plus"
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m11s
The spec said the metric follows the template the page renders. That is no
longer true for GIAS phase 6: a sixth-form college renders the primary
template but is matched, correctly, against secondaries. The rule is phase
group membership, decided once in is_secondary_phase — and the section's lede
noun comes from the school's phase rather than its template for the same
reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 06:41:59 +01:00
TudorandClaude Opus 5 2175dccb7c fix(api): treat "16 plus" as secondary, the way PHASE_GROUPS already does
PR Checks / Frontend Typecheck + Tests (pull_request) Canceled after 9s
PR Checks / Backend Smoke (pull_request) Canceled after 0s
PR Checks / Build Backend (no push) (pull_request) Canceled after 0s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 0s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
GIAS phase 6 is "16 plus", and PHASE_GROUPS deliberately files it in the
secondary group. The payload helper decided the same question with
`"secondary" in phase_text`, which that value does not satisfy — so a
sixth-form college was handed the primary bucket and offered infant schools
as its peers, with the KS2 metric key to label them. No crash; just a page
confidently showing the wrong schools.

The decision now lives in similar_schools.is_secondary_phase, beside the
PHASE_GROUPS bucket it selects from, so the two cannot drift again. A test
pins them together.

The same binary assumption had a second output. computeSchoolFlags tests for
the substring too, so a 16-plus school renders the primary template, and the
composer was labelling the section from the template: "Other primary schools
near <sixth form college>" above a row of secondaries. The section now
derives its noun from the school's own phase, which also removes the
duplicated wording from both composers. A 16-plus school's candidates span
the whole secondary group, so no single noun fits and it gets the honest
general one.

Reported in review on #150.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 06:41:49 +01:00
TudorandClaude Opus 5 bd2a6c385b test(e2e): cover the similar-schools section, compare hand-off and mobile widths
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m13s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 37s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 3m1s
Two things here cannot be covered anywhere else. jsdom has no layout, so
scrollWidth and clientWidth are both 0 and the arrows' disabled state can only
be measured by a real engine. And the scroll position surviving a selection is
DOM state rather than React state, so only a real browser can prove the row
does not jump back when the footer re-renders.

MOBILE.md asks for a Playwright width check and records that it was not written
because Playwright was not in the project. It is — this suite — so the check
exists now, scoped to the page this feature touches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:43:38 +01:00
TudorandClaude Opus 5 4e0d8bcf87 feat(web): render similar schools on both detail templates
Inside SchoolDetailShell rather than after it, because the sticky nav's
scroll-spy finds sections with getElementById and can only reach one that
lives in the shell. Last in the order, and last in the nav, because the two
must agree or the nav links to an anchor that was never rendered.

hasSimilarSchools is optional on NavItemsInput, matching hasLocation beside
it: absent has to mean "no section", and making it required would have
churned ten unrelated call sites for no added safety.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:42:59 +01:00
TudorandClaude Opus 5 91314a80b5 feat(web): similar schools section, honest about what it matched
Server-rendered cards inside a client carousel that scrolls rather than
paginates, so all six links stay in the initial HTML and the row still works
with JavaScript off. Three client islands, split by what each needs: a school,
the whole selection, a DOM ref.

The lede claims a similar intake only when no card came from tier 3, chips
list what a school actually shares, a missing figure reads "Not published",
and the neighbour's number carries no valence colour — green and terracotta
mean "against England" everywhere else, and colouring it here would read as
ranking the neighbours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:41:40 +01:00
TudorandClaude Opus 5 52b00ac752 feat(api): serve similar schools on the detail endpoint
Rides in the existing payload rather than a new endpoint: the page already
makes one server fetch for its data, and /school/[slug] regenerates weekly,
so the per-request cost is paid once per school per week.

Wrapped so a failure in selection never 500s a page that is otherwise
complete — the posture get_supplementary_data already takes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:40:35 +01:00
TudorandClaude Opus 5 b571d9c549 feat(api): decide which nearby schools a page may offer
Hard filters encode claims the section may not make — a selective school is
not an alternative to a non-selective one, a special school is not comparable
to a mainstream one, a Girls school is not an option for a Boys school's
reader — so they never relax. Soft preferences describe closeness of fit, so
they relax across three tiers, and only far enough to reach three; the
remaining slots up to six fill from the tiers already opened.

PHASE_GROUPS moves to schemas.py so this module can share it without
importing app, which would be a cycle.

_mask() exists because Series.apply on an empty Series returns a DataFrame,
and using that as a mask drops every column — so the next lookup raises
KeyError instead of yielding no rows. A special school with no special school
near it empties the frame at the provision filter, which is the ordinary case
for most special schools, so this was a crash on a common path rather than an
edge case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:40:10 +01:00
TudorandClaude Opus 5 8a23e3657d docs: meet the mobile baseline, which the carousel was not doing
Checked the section against MOBILE.md at its three reference widths instead
of assuming the breakpoints were enough. Two real failures at 360px.

The arrows sat in the heading's flex row, taking 96px from a 328px card and
crushing the lede into a four-line column — for a control that swiping
already provides. Below 640px they are now gone, the header is a single
column, one card shows at 86% so the next one peeks, and the affordance is
the right-edge scroll-fade MOBILE.md already documents for horizontal
scrollers. The fade lifts at the end of the travel, so the at-end state is
computed whether or not an arrow exists to consume it.

The arrow and add-to-compare buttons were 40px against a 44px floor. Both
are 44 now. A card title's own box is shorter, but its hit area is the whole
card through the ::after overlay, so it passes on the target that actually
receives the tap.

360, 390 and 430 now all report zero overflow, no failing tap targets and no
text under 11px. MOBILE.md wanted a Playwright width check and recorded that
Playwright was not in the project; it is, so the journey now carries one for
this page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:37:34 +01:00
TudorandClaude Opus 5 e4e8f02599 docs: six schools behind a carousel, and a rule for when to stop widening
Three schools is not a neighbourhood in inner London, so the cap is six with
three visible and arrows for the rest.

Raising the cap exposes something the old cap hid. Tiers exist to reach a
usable set, and with six slots a naive loop would keep widening to fill them
— dragging in tier-3 schools ten miles away to sit beside three good matches
that had already earned the row. So tiers now stop relaxing once three are
found, and the remaining slots are filled only from the tiers already used.
Four tier-1 matches never open tier 2.

The carousel scrolls a list rather than swapping a view: all six cards are in
the initial HTML, so every link stays crawlable and the row still scrolls with
JavaScript off. The arrows' edge test carries an 8px tolerance because the
scroller's focus-ring padding is the first snap position — a row at rest
reports scrollLeft 2, and an exact test for 0 left the back arrow live and
pointing nowhere. Caught in the mockup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:35:11 +01:00
TudorandClaude Opus 5 b62dc17532 docs: drop the method disclosure, keep the one caveat that earns its place
The "how these schools are chosen" panel restated what the section already
shows — the phase in the lede, the shared characteristics on each card, the
distance above each name — so it cost space to say nothing new.

One line survives, and it is not a method note. A reader who sees "0.6 miles
away" and takes it for the walk has been misled by us, and no other element
on the card corrects that. The rest were claims the selection rules keep
true without narrating them.

Also records what happens past the third school: surplus matches are dropped
silently, because NearbyPlaces below already leads to the full lists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:29:07 +01:00
TudorandClaude Opus 5 4d7762d796 docs: implementation plan for similar schools nearby
Five tasks, each ending in a green test run and a commit: the pure
selection module, the endpoint key, the section and its two client
islands, the wiring into both templates, and the journey.

The selection logic gets its own module rather than another 200 lines in
app.py, which means the tier rules are testable against a synthetic frame
with no TestClient, no database and no monkeypatch. Moving PHASE_GROUPS
to schemas.py is what keeps that import acyclic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:21:14 +01:00
TudorandClaude Opus 5 3650f7d8b7 docs: design for similar schools nearby on the detail page
A school page links outward to its places and never to another school.
This section adds that edge: three nearby schools of the same phase and a
comparable intake, each a crawlable link and each addable to the basket.

The design separates hard filters from soft preferences and never confuses
them. Selectivity, provision and opposite-sex intake are claims the section
cannot make, so they never relax, even where that means no section renders.
Gender and religious character describe closeness of fit, so they relax in
tiers — and the card states what actually survived rather than padding with
a match it did not earn.

Includes the mockup the design is drawn against, with all three tier states
live in both themes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:18:38 +01:00
tudor dfce308f1f Merge pull request 'fix(ci): make a failed release check say what it actually saw' (#149) from fix/release-check-diagnostics into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m24s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 28s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m9s
Reviewed-on: #149
2026-09-15 15:09:46 +00:00
TudorandClaude Opus 5 64ae71d7ab fix(ci): make a failed release check say what it actually saw
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m4s
The staging poller swallowed every failure identically, so a run that
timed out told us only that the expected release never appeared — not
whether the proxy refused us, the endpoint was down, or the containers
were still serving an older build. The public staging proxy also answers
403 to urllib's default user agent while the release endpoint is healthy,
which looked exactly like a deployment that never arrived.

Identify the poller, and report each distinct observation once: HTTP
status, connection failure type, invalid JSON, or the release identities
actually reported. The timeout error carries the last observation and the
identity it wanted. Responses and the base URL stay out of the logs —
only validated sha/build_id fields are echoed back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 16:02:14 +01:00
tudor 7ab084dd3a Merge pull request 'feat: verify what we actually deployed, and publish data all-or-nothing' (#148) from feat/release-identity-and-reliability-gate into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 0s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m31s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m27s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Failing after 5m3s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Skipped
Reviewed-on: #148
2026-09-15 11:13:08 +00:00
TudorandClaude Opus 5 5ad1cbfb53 fix(api): bound the search candidate set instead of draining Typesense
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 39s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 6m44s
Fetching every match kept scoped searches correct but left the number of
round trips in the caller's hands: a one-letter query, or a deliberately
broad one, walked the whole collection a page at a time.

Cap the candidate set at 1,000 URNs — four pages — and return the
relevance-ordered prefix when the ceiling is hit. That is still far more
than one page, so the API's own authority, phase and postcode filters
keep the matches they need, while latency and upstream load stay bounded.
A capped query is logged so a genuinely truncated search is visible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 11:21:45 +01:00
TudorandClaude Opus 5 0c901cd0d1 feat(ci): gate promotion on the image set that actually passed E2E
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 6m3s
Staging health polling asked only whether something answered HTTP 200 at
the base URL. It could not tell the new deployment from the old one, so
journeys could pass against the previous release, and concurrent merges
could move the staging tags underneath a run in flight.

Each staging run now mints a build ID and stamps all three images with
the commit and that ID, as labels and — for frontend and backend — as a
build-time JSON file that environment overrides cannot rewrite.
/release.json reports both identities uncached, and scripts/ci/release.py
polls for the expected pair before and after the journeys. Only then are
the captured build digests tagged verified-<sha>.

Promotion resolves those verified tags to immutable digests, revalidates
their labels, and refuses a mixed or incomplete set before any :prod tag
moves. The whole staging workflow shares one concurrency group with
cancellation disabled, so releases serialise.

The scripts are stdlib-only and unit-tested against mocked registry and
HTTP behaviour; PR checks now run the pipeline and CI suites too. The
runbook records what this cannot prove locally, and that the first
rollout needs a commit built by this workflow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:17:50 +01:00
TudorandClaude Opus 5 7b41218e6e fix(web): show an outage as an outage, and drop superseded fetches
The home page caught every fetch failure and rendered its empty state,
so a backend outage looked like a site with no schools in it. School
pages turned any error into notFound(), which told visitors — and
crawlers — that a real school had ceased to exist. Place fetches did the
same by returning [] and null. Failures now reach a retryable error
boundary; only a genuine 404 still calls notFound().

"Load more" and the map fetch resolved against whatever state existed
when they returned, so results from an abandoned search appended
themselves to the new ones. Each fetch now carries an AbortController
and checks that its search scope is still current before touching state.
The map only records its cache key on success, so a failed load retries
instead of pinning the stale marker set.

Jest ignored .next/, whose build output otherwise shadowed real suites.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:17:18 +01:00
TudorandClaude Opus 5 9b75f54206 fix(api): publish a validated dataset, and stop truncating search
Reload cleared the caches first and rebuilt afterwards, so any failure
left the API serving nothing, and requests arriving mid-reload saw a
half-swapped state. It now builds and validates the replacement frames,
place registry, reverse index and sitemaps off the request loop, then
publishes them in one synchronous step under a lock. A failed reload
returns 503 and keeps the previous data. Sitemap regeneration takes the
same path rather than clearing the live registry up front.

Typesense search returned at most one page of hits and used an empty
list for both "no matches" and "search is down", so a genuine empty
result silently fell back to substring matching. It now pages through
every candidate and returns None only on failure; the fallback matches
literally, since a query containing regex metacharacters used to throw.

Empty datasets answer 503 rather than 200-with-nothing or a misleading
404, so callers can tell an outage from an absent school. Adds
/api/release, which reports the build identity baked into the image.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:16:58 +01:00
TudorandClaude Opus 5 38bc17cab3 feat(search): validate the index before the alias points at it
The old sync created a collection, imported batches without reading a
single import response, and swapped the alias regardless. A partial
import published a half-empty index, and two overlapping DAG runs could
prune each other's collections.

Publication now checks every import response and the final document
count before upserting the alias, and holds a session-scoped advisory
lock across the read and the publish so concurrent runs serialise.
Cleanup keeps the previous collection as a rollback pointer and is
best-effort: an uncertain alias response must never delete what might
still be live.

Also parses the Typesense URL properly instead of splitting on colons,
which mangled any host carrying a scheme and a default port.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:16:36 +01:00
tudor dc156058fe Merge pull request 'docs: describe the system that exists, and remove the importer's remains' (#147) from docs/current-truth-and-legacy-inventory into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 22s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m22s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m57s
Reviewed-on: #147
2026-09-15 08:58:54 +00:00
TudorandClaude Opus 5 1d8858fbda chore: remove the code the legacy CSV importer left behind
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m2s
`backend/migration.py` and `scripts/migrate_csv_to_db.py` import `School`,
`SchoolResult`, `init_db` and `set_db_schema_version` — names that no longer
exist. `scripts/geocode_schools.py` imports the same removed ORM model. None of
the three can be imported against the current backend, so they were not dormant
utilities anyone could fall back on; they were files that would fail on the
first line. `backend/version.py` existed only to hand `SCHEMA_VERSION` to that
importer, and the FastAPI lifespan performs no version-triggered import.

Three symbols go with them, each confirmed to have no caller: the unvectorised
`haversine_distance`, superseded by the inline NumPy calculation in search;
`fetcher`, an SWR helper for a dependency this project does not install; and
`kmToMiles`. `calculateDistance` stays — CutoffMapPanel uses it.

Two comments pointed at `migrate_csv_to_db.py --drop` to explain why Payload
owns its own schema. The reason survives the script: blog content must stay
clear of the school marts and Airflow's metadata. Reworded rather than deleted,
so the constraint keeps its justification.

docs/LEGACY_CODE.md records what was removed and where to find it in history. It
also records what was deliberately *not* removed, which is the more useful half:
unused UI components awaiting a design decision, manual data utilities whose
operators a repository search cannot see, and fallbacks that look obsolete but
are load-bearing — `data_loader.py`'s older-mart branches, the generated GIAS
dictionary copies, and the `legacy`-named dbt models that annual DAG selectors
explicitly include. A zero-import count is evidence, not a verdict.

The scripts that fetch DfE CSVs are marked historical and kept, pending
confirmation that nobody runs them by hand.

Checked: 190 backend tests, 429 frontend tests, `tsc --noEmit` clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan
2026-09-14 23:01:22 +01:00
TudorandClaude Opus 5 eaf5e5d180 docs: describe the system that exists, not the one we started with
The README still opened on "Primary School Compass", a KS2 tool for Wandsworth
and Merton served by FastAPI and vanilla JavaScript with Chart.js. Every layer
of that sentence is now wrong: coverage is England-wide across KS2, KS4,
all-through and post-16, Next.js owns the public UI, and school data comes from
dbt-built `marts.*` rather than CSVs loaded at startup. The setup instructions
walked a reader into a virtualenv and a CSV import that cannot build the current
schema, so following the docs produced an empty database and a wrong mental
model at the same time.

Replace the narrative docs with two reference documents that were checked
against the code: docs/ARCHITECTURE.md for request flow, data ownership, the
backend/frontend module boundaries and the real publication sequence, and
docs/DEVELOPMENT.md for the checks that actually run, including the container
and CI version skew that makes "just run pytest" misleading.

The env examples drifted the same way. ALLOWED_ORIGINS is a JSON array, not a
comma-separated list; the frontend needs FASTAPI_URL, DATABASE_URL and
PAYLOAD_SECRET, none of which were documented; and RATE_LIMIT_BURST,
DEFAULT_PAGE_SIZE and MAX_PAGE_SIZE were presented as tuning controls the routes
do not consult. Each is now stated as it behaves.

MIGRATION_SUMMARY.md keeps its content but gains a banner, because it reads like
setup instructions and is not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan
2026-09-14 23:01:15 +01:00
tudor be780ebe13 Merge pull request 'fix(seo): declare the share card, which the route group stopped inheriting' (#146) from fix/og-image-route-group into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m21s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m58s
Reviewed-on: #146
2026-09-14 21:03:14 +00:00
TudorandClaude Opus 5 b6c2cd5116 fix(seo): declare the share card, which the route group stopped inheriting
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 53s
The staging E2E gate's one failure. Every link to the site pasted into
a chat has been rendering bare.

Staging serves og:title, og:description, og:url, og:site_name, og:type
and twitter:card, and no og:image at all. So metadata from the layout
reaches the page; only the file convention does not.

app/opengraph-image.tsx does work — _not-found, which lives in the app
root segment, carries an og:image from it in the build output. It does
not reach the site's pages, which live in the (frontend) route group
whose own layout.tsx is the root layout. The icon conventions are not
affected: /icon.png and /apple-icon.png are both linked correctly on
the same page, verified against staging. The asymmetry is the whole
bug, and it arrived with the route-group split that Payload required.

The file stays at the app root. Moving metadata files into a route
group is what drops /robots.txt and hashes /icon.png, which CLAUDE.md
records and which this must not undo — the build still emits all four
of /robots.txt, /icon.png, /apple-icon.png and /opengraph-image. The
root layout points at the route instead, and metadataBase makes it
absolute, which the journey needs since it calls new URL() on the value.

twitter.images is set for the same reason: the card is declared
summary_large_image, and claiming a large-image card while supplying no
image is worse than claiming a summary card.

Checked before fixing that og:image was the only broken assertion in
that journey: the test aborts at line 911, so its apple-touch-icon and
maskable-icon assertions had never run. All four of those assets return
200 image/png from staging, so this does not simply move the failure
further down the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 21:54:32 +01:00
tudor dc79d653e5 Merge pull request 'feat(seo): link school pages into the location layer' (#145) from feat/school-page-place-links into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m26s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m29s
Reviewed-on: #145
2026-09-14 20:29:48 +00:00
TudorandClaude Opus 5 b0d5334e06 perf(places): index the reverse lookup, and isolate the registry in tests
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m47s
Two review findings, both confirmed before fixing.

The client fixture in test_school_details.py patched load_school_data
but not _place_registry, which is a module-level cache. A probe settled
it rather than an argument: poisoning the global with a registry built
from data the fixture never saw, then issuing the fixture's own
request, returned that other dataset's places. So the new
`places == []` assertion was satisfied by a stale registry exactly as
well as by the fixture's own data, and proved nothing. Every other test
that touches place data already reset it; the fixture predates places
existing and was never updated. It resets it now.

places_for_urn walked every place in the registry and did a tuple
membership test against each, on /api/schools/{urn}, the site's
highest-traffic endpoint. It now reads a dict built once per registry.
Measured against a synthetic corpus of 27k schools in 1,650 places:
0.118ms per request becomes 0.0001ms, with the index built once in
21ms. Production carries ~5,000 places, so the scan there is larger
again. The absolute saving per request is small; the point is that it
is repeated on every school page view and costs nothing to remove.

The index is cached against the registry by identity rather than
behind a second flag. Anything that drops _place_registry — every test
that touches place data does — gets a fresh registry object, which no
longer matches what the index was built from, so the index rebuilds
with it. A separate _place_index = None would be one more thing to
forget, and a stale reverse index is precisely the first finding's bug
wearing a different hat.

That invalidation has its own test, and the test was checked by
breaking the identity check: five tests fail without it, so three
existing ones were already relying on it too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 21:22:39 +01:00
TudorandClaude Opus 5 d65eb58883 fix(seo): build the phase links the docstring already promised
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 3m48s
Review caught _places_payload documenting a `phase_url` the function
never returned. The docstring was not stray prose: the approved design
included the phase variant — "Primary schools in Beccles" was one of
its four example links — and it was dropped during implementation
without being mentioned. Deleting the sentence would have closed the
report while losing the feature, so the links are built instead.

These are the pages that most needed them. ~950 phase variants were
once reachable by nothing at all: absent from every sitemap and
unlinked from the place page. "Primary schools in brentwood" is the
query they exist to answer.

Membership is read from the registry's own `phase_urns` rather than
re-derived from the school's phase string. The registry already decides
which phases a place publishes and which schools are listed on each, so
asking it is both shorter and the only way the link cannot disagree
with the page it points at. It also means outcodes need no special
case: they carry empty `phase_urns` by design, because nobody searches
"primary schools in SW11", so they report no phase links on their own.

`phases` is a list rather than a single url. An all-through school is
listed on both the primary and secondary pages, so there is no tie to
break and no reason to invent one. Each entry renders directly after
its own place, so "22 primary schools in Brentwood" reads as part of
Brentwood rather than as an unrelated link further along the row.

The e2e journey now follows a phase link where the town publishes one
and asserts it resolves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 20:54:59 +01:00
TudorandClaude Opus 5 7f5f0fb676 feat(seo): link school pages into the location layer
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
W2 shipped ~5,000 place pages and nothing linked into them. The
location layer pointed down at school pages; school pages pointed
nowhere on the site. Their only anchor was the school's own website, so
the ~27k pages carrying most of the site's inbound authority passed it
straight off-site, and the new corpus was reachable mainly through the
sitemap.

Three things close the loop.

A reverse index over the place registry, places_for_urn, answers which
published places contain a school. Derived from the registry rather
than stored beside it, so the two cannot disagree about which places
exist: a place below the publish threshold is absent from the registry
and therefore never offered as a link. A test asserts that invariant
across every place in a built registry.

GET /api/schools/{urn} gains a `places` array carrying the name, count
and canonical path for each. It rides on the request the page already
makes, so the school page costs no extra round trip. The frontend types
it optional and defaults it to empty, because the two images deploy
separately and a frontend ahead of the API must render without it.

The page gains a "More schools near here" module and a BreadcrumbList.
The module orders narrowest first, because a reader on a school page
wants its town before its county, while the API orders widest first for
the trail. Anchors state their destination's size — "37 schools in
Brentwood" — which is worth more to a reader and a crawler than "see
more". With no published places it renders nothing rather than an empty
heading.

The trail is rooted at the homepage, not /schools. There is no /schools
index page; the location layer lives only at /schools/[place],
/schools/authority/[la] and /schools/near/[outcode]. Rooting it at the
bare path would have opened every breadcrumb with a link to a 404.
Outcodes are omitted from the trail: "schools near CM15" is a real
query and a useful link, but nobody navigates Essex to CM15 to a
school, and a breadcrumb claiming that describes a hierarchy the site
does not have.

School pages also now declare the School type rather than
EducationalOrganization, the parent type that covers universities and
nurseries alike.

The e2e journey asserts the round trip in both directions, following a
place page's own first school so the pair is genuinely related rather
than hardcoded. A one-way link is what already existed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 20:45:50 +01:00
tudor 47f3591ed8 Merge pull request 'feat(flags): put /about and /blog behind flags, dark by default' (#144) from feat/about-blog-flags into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m22s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m10s
Reviewed-on: #144
2026-09-08 16:47:18 +00:00
TudorandClaude Opus 5 6fc7fce948 feat(flags): put /about and /blog behind flags, dark by default
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 20s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m21s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 12s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m39s
Both features ship dark. Neither is reachable in an environment where
its flag is off, and every flag in this system starts off, so a deploy
of this commit makes both disappear until someone turns them on
deliberately.

Two independent flags rather than one, which makes blog-on-about-off a
reachable state. That state is the whole reason the change is larger
than four notFound() calls: the blog leans on the About page for its
author identity. The Person entity is anchored at /about#tudor, and
that URL 404s while about_page is dark, so a post published in that
state would claim an author resolving to nothing. Worse than having no
named author. Both bylines fall back to unlinked text and the
BlogPosting attributes to the publisher instead, so every combination
of the two flags renders something correct.

Gated: /about, /blog, /blog/[slug], the RSS feed, both footer links,
and the matching content-sitemap entries. A sitemap must never
advertise a URL that 404s. With both dark it emits a valid empty
urlset rather than a 404, because robots.txt names it unconditionally.

Not gated: /admin. Posts have to be writable before the blog is worth
switching on, so flagging the panel would make the flag unflippable.

getFlags takes a revalidate rather than always using the 300s
constant. Reading a flag pins the calling route to the lowest
revalidate among its fetches, and the footer links live in the root
layout, so a naive gate there would have dropped every school and
place page from a weekly cache to a 5-minute one. The layout passes
604800, the floor those routes already declare, and the build confirms
all four SSG route families still prerender. The cost is one-way
latency: pages follow a flip in minutes, footer links within a week.

The e2e journeys follow the existing paired shape from the
admission_distance flag: a lit journey and a dark one for each flag,
reading state from whether /about and /blog respond rather than from
/api/flags, which another journey asserts is not publicly reachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-08 17:12:25 +01:00
tudor eb6d918650 Merge pull request 'fix(cms): regenerate the import map so the Content field renders' (#143) from fix/payload-import-map into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 15s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m9s
Reviewed-on: #143
2026-09-02 20:12:17 +00:00
TudorandClaude Opus 5 d47ac71c47 fix(cms): regenerate the import map so the Content field renders
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m10s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 52s
Creating a post in the admin panel showed no Content editor, and saving
failed validation on the field the writer was never shown.

The admin panel does not import field components. The server hands the
client a path per field, and resolves it through the generated map at
app/(payload)/admin/importMap.js. A richText field's path is
@payloadcms/richtext-lexical/rsc#RscEntryLexicalField. The committed map
held one entry, @payloadcms/next/rsc#CollectionCards, generated before
the blog collections existed and never re-run. A path missing from the
map is not an error the panel reports: the field simply does not render,
while required is still enforced server-side on save.

next build does not regenerate the map, so the stale copy shipped in the
image and the editor was equally broken on staging and production.

Regenerated with payload generate:importmap, which adds the lexical RSC
field, cell and diff components, BlocksFeatureClient for the Callout
block, and the default toolbar features.

Two things stop it drifting again. There was no script to run, so
package.json gets generate:importmap. And a test asserts the map carries
an entry for each thing the config asks for, in the source-reading style
of the other payload suites; against the old map all five fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-02 20:44:05 +01:00
tudor 17e5371e9c Merge pull request 'fix(cms): ship the initial Payload migration (recovers two commits stranded after #140 merged)' (#141) from fix/payload-initial-migration into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m13s
Reviewed-on: #141
2026-09-02 17:52:17 +00:00
TudorandClaude Opus 5 e2c63a9905 fix(cms): ship the initial migration so a container finds its tables
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m14s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m33s
Staging failed on boot with 42P01, relation "payload.users" does not
exist. The schema was empty because no migration existed, and the
adapter cannot create tables itself: db-postgres/connect.js gates push
on NODE_ENV !== 'production', so it is inert in a deployed container
regardless of config.

The generated migration is schema-qualified to "payload" throughout but
does not create that schema — schemaName says where tables go, it does
not create anything. It only worked against the throwaway database used
to generate it because the schema was created there by hand, so every
real environment would have failed on the first statement. CREATE SCHEMA
IF NOT EXISTS is hand-added at the top of up(), which makes it exactly
the kind of edit a regeneration discards silently; a test asserts it is
present and ordered before the first CREATE TABLE.

payload-types.ts is now committed rather than ignored. Ignoring it meant
CI typechecked against looser types than a developer with a generated
copy, which is how a Record<string, unknown> cast passed CI and then
failed locally the moment the file appeared. The post page uses the
generated Post and Media types instead, and narrows heroImage rather
than asserting it, since the field is an id at shallow depth.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00
TudorandClaude Opus 5 3f3c5953f6 style(copy): remove em dashes from the site's prose
The em dash is one of the clearest tells of machine-written text, which
is the exact impression this work exists to remove. Rewritten rather
than substituted: where a dash was carrying a real aside the sentence is
split or recast, not patched with a comma.

Covers the About page, the two Callout labels an editor sees in the
admin panel, and PUBLISHING.md, which defines the house style and should
follow it. The rule is now recorded in that house style and in the
spec's voice rules, so it survives this branch.

Code comments are left alone: they are not copy, and the surrounding
codebase uses the same punctuation throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00
tudor 124c6702a9 Merge pull request 'feat: a named author, an About page and a Payload blog' (#140) from feat/about-and-blog into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 45s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m32s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m8s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m36s
Reviewed-on: #140
2026-09-02 16:04:58 +00:00
TudorandClaude Opus 5 e25722d9ab fix(blog): hide drafts at the access layer, and back the --drop claim
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 32s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m9s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m26s
Review findings on #140.

Drafts were reachable. Posts granted unconditional public read and the
_status filter lived only in the pages that query the collection — which
is a convenience, not a control. Payload's documentation is explicit:
"The `draft` argument alone does not restrict documents with _status:
'draft' from being returned by the API." A direct GET /cms-api/posts
would have handed every unpublished draft to any visitor. Read access
now returns a query constraint for anonymous callers, which is the
documented mechanism.

The --drop claim was asserted across four files while the spec still
listed it as an open question. Now verified rather than assumed:
run_full_migration drops exactly ["school_results", "schools"] by name,
there is no drop_all() or DROP SCHEMA anywhere in backend/, the only
other drop is schema-qualified to marts, and nothing sets search_path.
The guarantee is stronger than schema isolation alone — those two table
names do not exist in Payload — so the claim stands, but it now rests on
cited code. The spec records the evidence and closes the open item.

findPost is wrapped in React's cache(): Next calls generateMetadata and
the page separately for one request, so every post view ran the same
query against Postgres twice.

The bare .lede rule was dead — .prose p scores (0,1,1) and outranks it —
so only .prose .lede ever applied. Removed, with the specificity noted
so the surviving selector is not "simplified" back into a silent
regression.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:48:59 +01:00
TudorandClaude Opus 5 07d586d0ad docs(blog): how to publish, and why the app has two route groups
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 36s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m11s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m18s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m43s
PUBLISHING.md carries the house style with the posts, so the standard
survives without the design doc to hand — including the rule that a post
states what a metric does not show, which is the strongest signal a
human wrote it.

CLAUDE.md gains the two constraints that are invisible from the code and
expensive to rediscover: metadata file conventions break if moved into a
route group, and the build must keep succeeding with DATABASE_URL unset.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:30:09 +01:00
TudorandClaude Opus 5 b793640507 feat(blog): add the blog index, post pages, RSS and content sitemap
The rendering split is dictated by CI building with no database.

/blog, /blog/rss.xml and /content-sitemap.xml have no dynamic params, so
Next prerenders them at build time and the build fails on a missing
Payload secret — caught here, not on staging. They are force-dynamic
instead: one indexed query against Postgres on the same Docker network,
and a newly published post appears immediately rather than waiting on a
revalidation. /blog/[slug] keeps ISR, because with no
generateStaticParams there is nothing to prerender; it is generated on
first request and cached, which is exactly what the collection's
afterChange hook exists to invalidate.

RichText takes `converters`, not `blocks`, in Payload 3.88, and the
default converters must be spread or every paragraph and heading loses
its renderer and the body comes out empty.

BlogPosting references the Person and Organization by @id rather than
repeating them, so every post and the About page resolve to one author
entity instead of declaring several people with the same name.

/sitemap.xml is proxied from FastAPI, which knows nothing about Payload,
so the Next-owned URLs get their own sitemap and robots.txt lists both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:28:58 +01:00
TudorandClaude Opus 5 21a5d18f59 feat(blog): add the posts and media collections
Drafts are on so a post can be written across sittings without saving
being publishing.

afterChange and afterDelete revalidate every path a post appears on.
Blog pages are ISR because CI builds with no database, so without these
a published post would not appear until the revalidate window expired —
up to an hour of a writer concluding that publishing is broken. Payload
runs in the same process as Next, so these are direct revalidatePath
calls with no webhook and no shared secret.

Media writes to an absolute /app/media matching the compose mount; a
mismatch would write into the container filesystem, where the next
redeploy silently discards it. Alt text is required rather than
optional.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:24:15 +01:00
TudorandClaude Opus 5 f614414070 feat(about): give the site a named author
The site had no author, no statement of why it exists and nobody
accountable for its numbers, which is most of why it reads as machine
generated.

The page states plainly that its author is not an education expert. The
credibility claim is lived experience — a parent going through primary
admissions — plus stated provenance for every figure, which is true and
cannot be undermined by someone noticing there is no teaching
qualification behind it. First name only: the Person JSON-LD carries no
familyName, worksFor or affiliation, and a test asserts it stays that
way.

The footer gains a fourth column, with a tablet breakpoint so four
columns pair up rather than crushing before the 768px collapse. The nav
is deliberately untouched — its mobile tab bar already carries four
items.

public/brand/tudor.jpg is NOT in this commit. The page references it and
will show a broken image until the photograph is supplied; a stock
portrait would defeat the entire point of the work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:22:52 +01:00
TudorandClaude Opus 5 310b63b0cb build(cms): wire Payload into the Docker image and both stacks
Uploads go to a named volume at /app/media. The directory is created in
the image before the mount and covered by the existing chown, because
Docker seeds a fresh named volume from the image path — a missing or
root-owned directory there fails every upload with EACCES at runtime,
long after the build passed.

PAYLOAD_SECRET uses the same :? form as AIRFLOW_ADMIN_PASSWORD: refuse
to start rather than boot with an empty secret and accept forged
sessions. Staging's must differ from production's, which the header
comment now says explicitly. Portainer prefixes volume names per stack,
so payload_media isolates itself.

prodMigrations is not wired yet — generating the initial migration needs
a reachable Postgres. Follows in its own commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:19:27 +01:00
TudorandClaude Opus 5 c2c76c5817 feat(cms): keep the admin panel out of the index
X-Robots-Tag rather than the robots.txt Disallow alone, for the same
reason the staging rule uses one: a Disallow blocks crawling, not
indexing, so a URL found from an external link can be indexed without
ever being fetched — and blocking the crawl means the noindex is never
seen. Both mechanisms are applied to /admin and /cms-api.

The existing CSP is frame-ancestors only, which restricts who may embed
the site rather than what a page may load, so it cannot break the panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:17:55 +01:00
TudorandClaude Opus 5 c5a4d106da feat(cms): install Payload and serve the admin panel
Payload 3.88 runs inside the Next app against the existing Postgres, in
its own 'payload' schema so no pipeline operation on public — the app
tables, Airflow's metadata, migrate_csv_to_db.py --drop — can reach blog
content.

Its REST API is mounted at /cms-api. /api is the FastAPI proxy's
catch-all, which would swallow every admin call and forward it to the
backend with no error. The mount points live in lib/payloadRoutes.ts so
there is one definition and a test can assert it without importing
Payload: it is ESM-only, next/jest will not transform it, and appending
transformIgnorePatterns cannot un-ignore a package. Forcing it through
transpilePackages would change how the production build bundles Payload
to serve a test, so the live proof that /api still reaches FastAPI stays
where it belongs — the e2e journeys, which call /api/schools.

The package becomes ESM ("type": "module"), which Payload's CLI requires:
richtext-lexical has top-level await and the config cannot be require()d.
Only two files needed renaming, jest.config.cjs and a build script.

The build is verified to succeed with DATABASE_URL and PAYLOAD_SECRET
both unset, which is how CI builds it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:17:01 +01:00
TudorandClaude Opus 5 2437ffce42 refactor(app): move site routes into a (frontend) route group
Payload's admin panel ships its own root layout rendering html/body.
Next allows multiple root layouts only when no app/layout.tsx exists, so
the site's routes move into their own group. Route groups are invisible
to routing: every public URL is unchanged, verified against the build's
route table.

The metadata file conventions deliberately stay at the app/ root. Moving
them into the group renamed /icon.png to /icon-4usi79.png (likewise
apple-icon and opengraph-image) and dropped /robots.txt altogether,
which would have broken the /icon.png cache-control rule, the
outputFileTracingIncludes entry for the share card, and robots.txt.

darkThemeSafety reads app/globals.css off disk rather than importing it,
so it needed its own path fix — a grep for import specifiers misses it,
and it fails as an unrunnable suite rather than a failed assertion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:11:48 +01:00
TudorandClaude Opus 5 eb648f3f76 build(next): convert the config to ESM so Payload can wrap it
withPayload() is ESM-only, so next.config.js has to become .mjs. That
file also carries the rule that keeps staging out of Google's index, so
the conversion goes in on its own, behind a test that asserts the rule
survived — along with the standalone output, the opengraph-image font
tracing and the analytics frame-ancestors CSP.

Jest resolves the .mjs config without extra configuration, so
jest.config.js is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:08:58 +01:00
TudorandClaude Opus 5 74e5fffc10 docs(about-blog): implementation plan for the About page and Payload blog
Nine tasks, each ending in an independently testable deliverable.

Two structural findings that the spec did not anticipate, both recorded
in the plan. Payload's admin panel ships its own root layout rendering
html/body, and Next allows multiple root layouts only when no
app/layout.tsx exists — so every existing route moves into an
app/(frontend) route group first, on its own, with the full suite as the
gate. Route groups are invisible to routing, so no public URL changes.

The second finding corrects the spec: adding /cms-api to the FastAPI
proxy's exclusion list would be dead code, because that catch-all only
ever matches /api/*. The route remap alone is sufficient.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 15:59:44 +01:00
TudorandClaude Opus 5 748ef32180 docs(about-blog): design for a named author, an About page and a Payload blog
The site reads as synthetic because nobody is accountable for the
numbers, no editorial judgement is visible, and the voice is
institutional third person. This designs the fix: a named author
(first name, photo, explicitly not an education expert), a coded
/about page, and a blog backed by Payload CMS running inside the
existing Next app.

Also fills a hole in the SEO programme, which has eight workstreams
and no E-E-A-T or authorship signal on a YMYL corpus.

Records two collisions found while designing, both of which fail
badly if missed: Payload's default /api route fights the existing
FastAPI catch-all proxy, and withPayload() is ESM-only so
next.config.js — which carries staging's noindex header — has to
become next.config.mjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FT1Ls4GbgLDXoQX7NAuHGT
2026-09-02 14:43:20 +01:00
tudor b0c4ea8282 Merge pull request 'fix(destinations): the DAG died on a null in a primary key' (#139) from fix/destinations-national-grain into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 14s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m36s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m6s
Reviewed-on: #139
2026-08-31 20:43:26 +00:00
TudorandClaude Opus 5 e236669fde fix(destinations): school rows and the England reference are different grains
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 52s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m27s
The annual DAG died with a BrokenPipeError from Meltano's log writer, which
is several frames from the cause: target-postgres exited first and the tap
saw its stdout close.

The tap declared primary_keys = [urn, ...] while emitting urn=None for the
national rows, and target-postgres turns primary_keys into a NOT NULL
constraint. The first national row of the run failed the insert and took
the loader with it. Every other tap in this repo keys on non-null columns.

Carrying two grains in one stream was the actual mistake, so the fix is to
separate them rather than paper over the null: four streams now, with
ees_ks4/ks5_destinations_national carrying no urn column at all — a school
identifier that is null in every row is a grain mismatch, not a column.
The staging models split the same way and the national mart reads the new
pair instead of filtering `where urn is null`.

Verified against the live API: the school stream yields 135,240 rows over
4,508 schools with no duplicate keys, no null key columns and all 31,382
suppression sentinels intact; the national streams yield 30 and 33 rows
with no urn column.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 21:36:43 +01:00
tudor fb5a0928bd Merge pull request 'fix(airflow): a fixed admin password from the environment, and a tap that missed #137' (#138) from fix/airflow-fixed-admin-password into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 55s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m36s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m14s
Reviewed-on: #138
2026-08-31 19:50:00 +00:00
TudorandClaude Opus 5 264edd2e3a fix(airflow): a login that survives a container restart
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m7s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 12s
PR Checks / Build Frontend (no push) (pull_request) Successful in 50s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 51s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m58s
The simple auth manager generates a random password on first start and
writes it to a file, so every restart of the api-server invalidated the
last one and the password had to be dug out of the container logs again.

The stack now writes that file itself from AIRFLOW_ADMIN_PASSWORD before
exec'ing the api-server. Airflow generates nothing when the file already
exists, so the login is whatever the stack environment says it is.

Written with python rather than echo, so json.dumps escapes a password
containing quotes, backslashes or non-ASCII correctly — verified against
`p@ss "wo\rd' £5`, which round-trips intact.

An unset AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits.
Falling back to a generated password would silently undo the point of the
change, and a compose-level `:?` gives the same refusal a readable reason.
This does mean the variable MUST be set in Portainer before the next
deploy of either stack.

Not affected by the two Docker gotchas in the upstream docs: this image
has no USER directive so it runs as root, and the file is rewritten from
the environment on every start rather than persisted on a volume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 17:38:42 +01:00
TudorandClaude Opus 5 cd2cbe7be6 build(pipeline): install the destinations tap in the image
meltano install would resolve it from pip_url, but five of the six custom
taps are also installed explicitly and a new plugin failing to appear is
not something you want to debug from a deploy log.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 17:37:46 +01:00
tudor 73182d0c0c Merge pull request 'feat(destinations): say what happened to a school's leavers, without republishing what DfE withheld' (#137) from feat/ks4-destinations into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m6s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m2s
Reviewed-on: #137
2026-08-30 20:49:14 +00:00
TudorandClaude Opus 5 cbe3a9a772 fix(destinations): the table said 'withheld' for a category that just doesn't apply
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m14s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
The Share column keyed off `percentage === null`, which is true for
not_applicable as well as suppressed, so a destination that does not apply
to the school was labelled as one DfE withheld — while the Pupils column
in the same row rendered blank. Two columns, one row, disagreeing about
what the row was, and one of them making a claim about DfE that wasn't
true.

Both columns now derive from `status`, which is the distinction the mart,
the SQLAlchemy model and the serialiser all preserve deliberately:
published shows the figure, suppressed shows the withheld badge,
not_applicable shows an em-dash with a title saying so.

A published count with no published percentage now derives its share from
the cohort rather than falling through to a marker — both halves are
published, so nothing withheld is involved, and it is the same derivation
the bar widths already use.

Verified the new tests fail against the old logic before keeping them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-30 21:48:39 +01:00
TudorandClaude Opus 5 2e9b5c83c5 fix(destinations): the masking pass can no longer exit unsafely
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 7m23s
Review found _mask_for_disclosure could return with its invariant broken
and say nothing. add_companion only ever withheld a *published* cell, so a
group with one suppressed category and every other one not_applicable —
routine in special schools and AP, where few categories apply — left the
loop with the lone suppressed cell still solvable. Reproduced on a
nine-pupil cohort: one hidden cell, cohort served, residual intact.

A disclosure-control pass that fails silently is worse than none, because
everything downstream trusts it. The loop now runs until the invariant
holds and escalates when no companion exists: the pupil group is dropped
from the payload, and an empty block serialises as None so the section is
absent rather than an empty shell. disclosure_invariant_holds() is exported
so tests assert it directly instead of re-deriving it, and an exhaustive
test sweeps all 81 suppression patterns of a four-category group.

Also fixes a test that set up six measures and checked one: the loop was
`for measure in ["school_sixth_form"]`. It now checks every measure, and
against the real invariant — none hidden, or at least two, rather than
"at least two", which the five published measures would have failed.

No regression on real data: 262 mainstream secondaries, all-pupils bar
still drawable on 94%, zero invariant violations, one disadvantaged group
dropped by the new escalation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-30 21:35:03 +01:00
TudorandClaude Opus 5 102397fe69 fix(destinations): withhold at the API, not just in the chart
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m14s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m49s
Code review found the disclosure the whole design was meant to prevent.
R1 was written as a rendering rule and implemented as one: canRenderBar
stopped the bar being drawn, but GET /api/schools/{urn} still carried the
cohort and every published category. cohort - sum(published) returned
Whitley Bay's withheld further-education figure exactly — 18 pupils — to
any caller, and the RSC payload put it in the browser too.

app.py already stated the principle for admission_distance: this endpoint
is public and unauthenticated, so a field left in the payload is a
published field. The same reasoning applies here and did not get applied.

_mask_for_disclosure now closes both identities before serialisation —
categories sum to the cohort, and disadvantaged + other = all — by adding
secondary suppression until every row and column hides none or at least
two. My first attempt picked the smallest published cell as the companion
and a new test caught it choosing a zero, which protects nothing: the
residual still resolved to 18. The companion must carry pupils.

DfE's own aggregates are no longer served. Nothing rendered them, and one
spanning a single suppressed component names it.

Cost, measured over 262 mainstream secondaries: the all-pupils bar
survives on 94% rather than 100%. Zero lone-suppressed groups remain.

The e2e helper now tells a missing feature apart from missing data: it
fails if the API serves no destinations key at all, and skips if the key
is served but the annual DAG has not populated the marts. Failing on the
second would redden the staging gate for unrelated commits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 19:58:02 +01:00
Tudor 68a192e430 Merge remote-tracking branch 'origin/main' into feat/ks4-destinations
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m12s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 4m4s
# Conflicts:
#	nextjs-app/__tests__/components/darkThemeSafety.test.ts
2026-08-28 18:45:27 +01:00
TudorandClaude Opus 5 ccd5074c90 test(e2e): destination journeys, including the no-bar rule
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m10s
The helper throws rather than skipping when no school returns a
destinations block: a silent skip would let a real regression in the
sections ride along unnoticed, which is why the distance journeys were
changed the same way in 4f01fbd.

The disadvantaged journey computes the residual itself and asserts it
appears nowhere on the page — the one number the section must never state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:16:25 +01:00
TudorandClaude Opus 5 2b4cf20d75 feat(destinations): the post-16 section, replacing the placeholder
The 'Post-16 destination data coming soon' note is deleted rather than
reworded: for a school with no sixth form the truthful statement is that
the question does not apply, and a placeholder there implies something is
missing. hasSixthForm and .sixthFormNote go with it — nothing else used them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:15:21 +01:00
TudorandClaude Opus 5 cef2f77149 feat(destinations): the After Year 11 section
Question cards over one bar, with the cards acting as a lens on the bar
rather than a summary beside it — focusing a card dims everything it is
not made of, so the grouping we chose is inspectable rather than asserted.

The bar renders only when canRenderBar allows it. Where a category is
withheld the section says so and shows the table instead: the categories
sum to the cohort, so a bar drawn from the published segments leaves a
gap whose width is the withheld figure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:14:13 +01:00
TudorandClaude Opus 5 68b6417149 feat(destinations): types and secondary section flags
Destinations are secondary-only, so the flags go on computeSecondaryFlags
rather than computeSchoolFlags. A phase counts as present only when some
pupil group carries categories — an empty block would otherwise open a nav
entry pointing at a section that never renders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:11:37 +01:00
TudorandClaude Opus 5 c5719ef362 feat(destinations): serve destinations without closing the gaps
The serialiser carries status through and computes no totals of its own.
The only aggregates in the payload are ones DfE published itself; whether
showing one is safe depends on how many of its components are suppressed,
which the frontend decides.

The batch guard grows from six tables to eight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:10:29 +01:00
TudorandClaude Opus 5 5e5b61987a feat(destinations): marts, with R3 masking applied at the boundary
Disadvantaged and other-pupils partition the whole and the all-pupils
figure is published, so publishing both halves recovers the suppressed
one. The mask is applied in the mart rather than the API so no consumer
added later can reach an unmasked combination.

The R1 test is a warn, not an error: DfE publishes the recoverable
combination and the mart's job is to carry it faithfully. Refusing to
close the gap is the API's job and the frontend's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:08:20 +01:00
TudorandClaude Opus 5 c564566432 feat(destinations): staging models that keep 'withheld' distinct from 'absent'
safe_numeric maps every EES sentinel to NULL, which is right for attainment
and wrong here: one of those states has to print 'withheld' and the other
has to print nothing. A status column carries the difference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:07:17 +01:00
TudorandClaude Opus 5 9188626051 feat(destinations): a tap that preserves the suppression sentinel
EES writes 'c' where a figure is withheld and the categories sum to the
cohort, so counts and percentages are emitted as text with the sentinel
intact. safe_numeric must never be pointed at them.

School rows and the England reference need different establishment pins:
at national level selective schools, studios and UTCs are separate
populations rather than labels, so leaving establishment open multiplies
30 rows into 190. Two queries per period, each keeping its own level.

Verified against the live API for 2022/23: 135,240 school records over
4,508 schools, exactly 30 each, no duplicate keys, 31,382 sentinels kept.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:06:37 +01:00
TudorandClaude Opus 5 7ae9ecdc36 feat(destinations): colour tokens, with the absence hatched not coloured
Activity not captured includes independent schools and moving abroad, so a
red segment would be a factual error. The hatch doubles as the secondary
encoding that rescues the neutral/blue pair, which separates at only dE 7.6
as flat fills.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:58:25 +01:00
TudorandClaude Opus 5 1980d79eee feat(destinations): the disclosure rules, as executable guards
The destination categories sum to the cohort and DfE publishes the cohort
total, so a lone suppressed cell is recoverable by subtraction. canAggregate,
canRenderPublishedAggregate and canRenderBar are what stop a consumer doing
that; toBarSegments throws rather than leaving a readable gap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:57:46 +01:00
TudorandClaude Opus 5 9423f11567 docs(destinations): implementation plan, ten tasks
Ordered so the disclosure guards land first and everything downstream
consumes them: lib/destinations.ts, tokens, tap, staging, marts, API,
then the two sections and the journeys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:56:36 +01:00
TudorandClaude Opus 5 576013d627 docs(destinations): design for KS4 and post-16 destination measures
The published files suppress individual cells, not whole cohorts, and the
categories sum to the cohort — so on 22% of mainstream secondaries the
withheld figure can be recovered by subtraction. Three disclosure rules
fall out of that, and the rest of the design is downstream of them.

Verified against the EES API rather than assumed: both datasets carry
school-level rows keyed by URN, with the disadvantage split and every
destination category the display needs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:28:09 +01:00
tudor 7c08138fe4 Merge pull request 'fix(map): the popup never took the dark theme' (#136) from fix/dark-mode-map-popup-contrast into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m41s
Reviewed-on: #136
2026-08-27 21:57:53 +00:00
TudorandClaude Opus 5 a7829d591a fix(map): the popup never took the dark theme
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 15s
leaflet.css paints `background: white; color: #333` on the popup card and its
tip. LeafletMapInner binds themed content into it — the school name and the
headline figure are var(--text-primary) — so in dark mode #E9EEF0 landed on
#FFFFFF at 1.17:1. The two things the popup exists to say were the two least
readable things on the page.

Every other foreground in that popup failed too, from the same cause: the
muted phase line at 2.90:1, the vs-national delta at 1.94:1, the Ofsted badge
at 1.74:1. Moving the surface onto --bg-card fixes all of them at once —
13.52, 5.45, 8.14 and 9.11:1 respectively. In light mode --bg-card is #FFFFFF,
so the popup renders exactly as it did.

globals.css already pulls the rest of Leaflet's chrome onto the tokens, and
says why: "this matters most in dark mode, where Leaflet's white attribution
bar would otherwise sit on a near-black page." The popup was simply missed.

The View Details button needed its own fix. It pairs background:var(--status-
above) with a literal white label, which theming the card does not reach:
--status-above is #36743F in light but #7FCB8A in dark, taking the label from
5.63:1 to 1.94:1. --text-inverse is the token for ink on a saturated fill, and
the popup's own Ofsted badge already uses it.

darkThemeSafety already guards this defect class, but only inside .module.css.
Neither half of this one lives there — the surface is a third party's, the
text is inline in a TSX template — so it scanned clean throughout. Two rules
added for the layer it could not see. Fixing the grouped-selector blind spot
in its rules() helper was needed to write them: taking only a selector's last
line discarded every selector in a grouped rule but the final one, which makes
a safety guard fail open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 22:34:18 +01:00
tudor 1ed4470fc2 Merge pull request 'fix(admissions): flag-off pages must not speak for the council' (#135) from fix/distance-flag-off-absence-copy into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 52s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m43s
Reviewed-on: #135
2026-08-27 20:30:25 +00:00
TudorandClaude Opus 5 7a16b1b52f fix(admissions): flag-off pages must not speak for the council
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m8s
The secondary admissions section words the absence of a cut-off distance:
"<LA> has not published a cut-off distance for this school." That sentence
is true when the authority publishes nothing. It is false when the
authority does publish and the admission_distance flag is simply off — and
off is the current state, so every secondary page with an EES admissions
row has been making a claim about a council on our behalf.

The backend already draws the distinction the copy needs. /api/schools/{urn}
omits the admission_distance key entirely while the flag is dark rather than
sending null, precisely so that "we are not publishing cut-offs" stays
distinguishable from "this school has no cut-off"; lib/types.ts says so in
as many words. The page then collapsed the two with `?? null` before the
section ever saw them.

So stop collapsing it: thread the raw field to SecondarySchoolSections and
word the absence only when the feature is on. Null still gets the sentence
naming the authority — that case is unchanged and still tested.

Primary pages are unaffected: AdmissionsSection carries no absence copy and
renders nothing when there is no figure. DistanceSection already treated
absent and null alike; only its type widens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 21:17:16 +01:00
tudor cf9d41b476 Merge pull request 'fix(analytics): the funnel source read a referrer that never changes' (#134) from fix/navigation-source-soft-nav into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 52s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m40s
Reviewed-on: #134
2026-08-27 08:26:45 +00:00
TudorandClaude Opus 5 e820e7fecd fix(analytics): the funnel source read a referrer that never changes
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m13s
The staging E2E gate has been red since #132 merged (run 1064, and
1066 after it): "a school reached from a location page is attributed
to it, not to direct" expects `place`, receives `direct`.

#132 fixed a real bug — `/schools/` had no case and fell through to
`direct` — but the mechanism underneath it never worked.
getNavigationSource read document.referrer, which the browser writes
only when a *document* loads. Every internal navigation here is an App
Router soft navigation: history.pushState, no new document, so
document.referrer goes on naming whatever opened the tab for the whole
session.

Verified on staging: load /schools/brentwood, click a school, the URL
becomes /school/… and document.referrer is still "".

So `from` reported `direct` for essentially every in-app journey, not
just the ones through the location layer — search, rankings, compare
and detail were all being counted as "typed the URL". The unit suite
passed throughout because every case set document.referrer directly,
which only happens on a full page load.

The fix is a module-level trail written by RouteTrail, a render-nothing
client component in the root layout. Its lifetime is exactly right: it
survives soft navigation, and it dies on a real document load — which
is precisely when document.referrer becomes meaningful again, so the
two cover each other with no overlap.

Reading it skips entries equal to the current path rather than taking
the second-to-last. That makes the answer independent of whether the
layout effect or the page effect ran first — React orders those by
tree position, which is not a contract worth resting a measurement on
— and it gives the right answer both when the user returns to a page
they came from and on a hard load of a school page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 09:22:37 +01:00
tudor 4fdeb70a93 Merge pull request 'feat(places): say what each school is, not only how it scored' (#133) from feat/place-school-attributes into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m42s
Reviewed-on: #133
2026-08-27 07:59:10 +00:00
TudorandClaude Opus 5 9a1f56c431 feat(places): say what each school is, not only how it scored
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m44s
The location tables carried one column: a percentage. A parent
shortlisting from a town page is asking a different question first —
does it take my child's age, is it a faith school, does it have a
nursery — and the page could not answer any of it.

Primary tables gain Ages, Religious character, Nursery and
Constituency; secondary tables the same minus Nursery, which is a
question about a different intake. An all-through school renders in
both groups, so its nursery shows under primary alone.

The measure moves to the second column rather than the last. Six
columns overflow a phone and .tableWrap turns that into a horizontal
swipe; with the measure last, the one number the page exists for is
the one scrolled off the screen.

Cell rules are the ones the school page already uses, so the two
surfaces cannot disagree about the same school: "Does not apply",
"None" and "Not applicable" all read as no religious character, and
the en-dash age normalisation moves into formatAgeSpan, which
formatAgeRange now delegates to.

Backend: nursery_provision and parliamentary_constituency were not in
the place response. Both are optional GIAS mart columns that
data_loader degrades to NULL, and the `in rows.columns` guard keeps a
mart the pipeline has not rebuilt working.

Also fixes a live bug on the same line: SCHOOL_COLUMNS already ends
with latitude and longitude, and the endpoint concatenated them again,
so pandas dropped one of every duplicated pair and warned "columns are
not unique" on each request. Ordered de-duplication removes the
warning and the silent drop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 08:49:14 +01:00
tudor ade9dbb3ba Merge pull request 'feat(analytics): measure the location layer, and stop calling it direct' (#132) from feat/place-analytics into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m44s
2026-08-27 07:29:48 +00:00
TudorandClaude Opus 5 d1a8596208 feat(analytics): measure the location layer, and stop calling it direct
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 40s
The location pages were only half-tracked. Umami counts a pageview for
each of the ~3,900 URLs automatically, but nothing else: components/
places contained no track() call, and place_viewed was not even a
declared event name.

The part that mattered was worse than a gap. getNavigationSource mapped
a same-origin referrer to a funnel source and had no case for /schools/,
so every school view arriving through the location layer fell through to
'direct' — the bucket you read as "typed the URL, no referrer". W2's
whole purpose is funnelling search traffic onto school pages, so the one
measurement that says whether it worked was reporting the wrong answer,
and reporting it confidently. Verified live against staging: expected
"place", received "direct".

/schools/ is checked before /school/. They differ by one letter and mean
different things — the location layer versus a single school — and a
prefix test in the wrong order silently merges them.

place_viewed carries kind, slug, phase and school_count. kind is the
reason it exists: whether to keep investing in these pages turns on
which sort earns engagement, and a pageview cannot say, because all four
families share the /schools/ prefix and only the registry knows which is
which. It is a client component because PlaceView is a server component;
one line in PlaceView covers all four families, since they all render
through it.

Both E2E journeys were verified failing against staging first — one
because place_viewed does not exist there, the other on the exact
"place" vs "direct" mismatch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-27 08:23:21 +01:00
tudor a3c09d9b67 Merge pull request 'fix(suggest): the dropdown reopened on top of the search results' (#131) from fix/suggest-reopens-over-results into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m37s
Reviewed-on: #131
2026-08-26 21:07:55 +00:00
tudor a7f4c86464 Merge pull request 'fix(search): the mobile hero search was indented by a card's padding' (#130) from fix/mobile-hero-search into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m50s
Reviewed-on: #130
2026-08-26 20:51:20 +00:00
TudorandClaude Opus 5 0804566736 fix(test): drop a committed scratch probe, and close a hole in the guard
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
Code review, all three findings valid.

e2e/tests/__m.spec.ts was a throwaway probe used to measure the mobile
hero geometry. It asserts nothing, so it could never fail; it carried a
leftover `pick('form').constructor === Object ? null : null` that is
null either way and throws if no form matches; and it should never have
been committed. Deleted.

It survived because `rm -f e2e/tests/__m.spec.ts` ran with the shell
already inside e2e/, so the path resolved to e2e/e2e/tests/... — which
does not exist, and rm -f is silent about that. `git add -A` then swept
it in. I checked `git diff --stat` before committing, which lists only
tracked modifications and never shows an untracked file; `git status
--short` would have.

The scoping guard compared the last line of a rule's prelude against the
literal '.filterBar', so a regression written as a selector list —
`.filterBar, .other { padding }`, or the same split across two lines —
would have walked straight past the test meant to catch it. Selectors
are now split on commas and matched individually, and comments are
stripped first so a brace inside one cannot desynchronise the parse.

Verified against all three shapes: bare, inline comma list, and
multi-line comma list. Each is caught; each passes again once reverted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:49:48 +01:00
TudorandClaude Opus 5 55363cbd18 fix(suggest): the dropdown reopened on top of the search results
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 55s
Three staging-gate failures, two of them one real bug.

After a search, the results-page bar still holds the term in its input,
so on every render the query was >= 2 characters and the suggestion list
opened again — on top of the very results the search had just produced.
Playwright reported it as "<li role=option ...> intercepts pointer
events" while trying to click the first result; a reader would simply
have found their first result unclickable. Both the school-detail and
hero-map journeys failed on it, and neither is about autosuggest.

Suggestions now answer typing, not the mere presence of a value:
`hasTyped` gates the hook, is set on change, and is cleared when a
search is submitted or a suggestion is chosen. A pre-filled input makes
no request and shows no list.

Third failure was my test, not the product. An unphased place page
renders one table per phase, and an all-through school legitimately
appears in both — so the page's school links were never one alphabetical
run. The assertion collected them all together and only passed because
no town it picked had held an all-through school. When the data gave
Abbots Langley one, Breakspeare School appeared in the primary table and
again in the secondary, and the test failed on correct behaviour. It now
checks each table separately, and passes against the data that broke it.

Guards: a jest test that a pre-filled input neither fetches nor opens
(verified by reverting — it is the only one that fails), and an E2E
journey that submits a search and then requires the first result to be
clickable, which is the reader-facing version of the same thing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:38:16 +01:00
tudor 868eb344f5 Merge pull request 'fix(map): the hero map's fade to the header was hardcoded white' (#129) from fix/dark-map-fade into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m49s
Reviewed-on: #129
2026-08-26 20:32:45 +00:00
TudorandClaude Opus 5 0b15497c09 fix(search): the mobile hero search was indented by a card's padding
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m3s
Measured at 390px: the headline and lede sit at x=34, while the search
box, the hint and the location link all sat at x=48 and the field was
28px narrower than the copy above it.

The 14px came from `@media (max-width: 768px) { .filterBar { padding:
0.875rem } }`. That rule is for the results filter bar, which is a card
— background, border, shadow — and needs inner padding. The hero search
is not a card: .heroMode strips all of it, padding included.

Both selectors are specificity (0,1,0), so source order decides, and
.heroMode only wins because it is declared right after .filterBar. A
bare .filterBar rule inside a media query comes later and silently wins
instead. The two rules directly below this one in the same block were
already written as `.filterBar:not(.heroMode)`; this one was missed.

Scoping it aligns the search box, hint and location link to the same
left edge as the headline and gives the field back its 28px.

The location link also carried its own 6px of button padding, so its
label started further right than the hint even once the boxes agreed.
Pulled back with a negative margin, which keeps the tap target.

The guard is a stylesheet test: the failure is a plausible-looking
layout rather than a broken one, so nothing short of measuring or
looking would catch it. Verified by reverting: it names ".filterBar sets
padding".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:32:41 +01:00
tudor d55f6cce23 Merge pull request 'fix(suggest): let the dropdown out of the hero panel' (#128) from fix/hero-dropdown-clipping into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m50s
Reviewed-on: #128
2026-08-26 20:23:33 +00:00
TudorandClaude Opus 5 3236efa846 fix(map): the hero map's fade to the header was hardcoded white
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m7s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 49s
The fade between the map band and the school header ramped through
rgba(255,255,255,...) and landed on var(--bg-card). In the light theme
that is white into white and invisible, as designed. In the dark theme
it climbed to 95% WHITE and then met a near-black card, putting a bright
band across the full width exactly where the map should dissolve into
the title.

Fading to the colour the gradient lands on is the whole trick, and it
only works if that colour is a token — so --bg-card-rgb now exists in
both theme blocks, matching the --hero-ground-rgb precedent.

Two more defects in the same file, same cause, found while in there:

The controls floating over the map paired a hardcoded white background
with color: var(--text-primary), which resolves to #E9EEF0 in dark —
near-white text on a near-white button. These deliberately do NOT follow
the theme, because the map tiles are light in both, so the ink is now
literal too and says why. A themed token is the wrong tool for a surface
that never changes.

The loading skeleton swept 50% white across var(--bg-secondary), which
is a bright flash every 1.4s on a dark page. It now sweeps toward the
card colour, a shade lighter than the ground in both themes.

The guard is a stylesheet test rather than a render test, because the
bug is invisible in the theme it was written for. Verified by reverting
each fix in turn: it names .fade and .openHint exactly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:20:57 +01:00
TudorandClaude Opus 5 d5a6db289d fix(suggest): let the dropdown out of the hero panel
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 54s
.heroPanel had overflow: hidden to clip its artwork and scrim to the
rounded corners. It clipped the suggestion dropdown too. Measured on
staging with the flag on: the list runs 482 to 802, the panel ends at
624 — so 178px of 320 was cut off, about half the options, with nothing
on screen to say anything was missing.

The two things that actually needed clipping now round themselves:
.heroArt gets border-radius: inherit plus its own overflow, and the
::before scrim inherits the radius. Below 860px the artwork is a band
flush with the top of the panel rather than a layer covering it, so it
takes the top two corners only — inheriting all four would leave it
floating with rounded corners against the copy.

Nothing else depended on the panel clipping: .valueProps below it is
entirely static, so a positioned dropdown paints above it without a
z-index fight.

The regression test asserts the LAST option is the element actually
painted at its own coordinates. toBeVisible() would not have caught
this — it checks for a non-empty box and visibility, and an ancestor's
overflow clips neither. elementFromPoint catches clipping and occlusion
alike.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:10:32 +01:00
tudor d8ccb5b733 Merge pull request 'feat(suggest): school autosuggest, and the rate-limit fix it needed first' (#127) from feat/school-autosuggest into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m49s
Reviewed-on: #127
2026-08-26 19:58:57 +00:00
TudorandClaude Opus 5 0fa1a292c7 fix(api): bound what a forged CF-Connecting-IP can buy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Code review, both findings valid.

The design doc claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the X-Forwarded-For fallback was
forgeable. That is true only for traffic that actually passed through
Cloudflare, and nothing in this process can verify that it did. Reaching
the origin directly, both headers are equally attacker-controlled — and
rotating CF-Connecting-IP mints a fresh rate-limit bucket per request,
defeating per-client limits on every endpoint including the
DataFrame-heavy /api/schools. Against abuse that is worse than the
shared bucket it replaced, which at least capped everyone together.

So the ceiling comes back. I dropped it earlier arguing it belonged at
Cloudflare; that argument assumed the keying was sound, and it is not.
GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window
against a total, independent of client identity, outermost so it refuses
before any work happens. Written by hand because slowapi cannot express
a global cap: default_limits and application_limits are both keyed by
key_func, and the latter needs middleware this app does not install.

It does not make the header trustworthy — it makes trusting it
survivable. The real fix is Authenticated Origin Pulls or an origin
firewall, now documented in DEPLOY.md as the open gap it is.

127.0.0.1 is exempt: the healthcheck curls localhost from inside the
container, and starving it would restart the container and turn a load
spike into an outage loop. Keyed on the peer address, never the Host
header, which the caller sets.

Second finding: suggest_schools_typesense promised "never raises" while
the parsing loop sat outside the try, so int(None) on a malformed
document would have made a keystroke a 500. The loop now skips bad rows
rather than dropping the whole list — and a hit with no document no
longer becomes a suggestion pointing at /school/0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:56:32 +01:00
tudor c3f044bd65 Merge pull request 'docs(flags): Unleash does not create flags by itself' (#126) from fix/flags-runbook into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 44s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 53s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m10s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m46s
Reviewed-on: #126
2026-08-26 19:48:58 +00:00
TudorandClaude Opus 5 d2115364ae test(e2e): autosuggest journeys, gated on the flag
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m13s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m40s
Feature state is read from its observable effect — whether the search
box is a combobox — because /api/flags is denied to the public on
purpose. Same approach as the distance journeys.

The three endpoint tests are ungated: /api/suggest is live whether or
not the UI is, which is what lets it be smoke-tested in an environment
where the feature is still dark.

The flag-off journey asserts the plain search still works, not just that
the combobox is absent. Verified against staging, where the flag is off:
it passes and the flag-on journey correctly skips.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:39:33 +01:00
TudorandClaude Opus 5 28cf0a342c feat(suggest): wire autosuggest into the search box behind a flag
Off means off — no combobox role, no listener, no fetch. A test asserts
the absence of the request, not just the absence of the dropdown,
because a hidden-but-fetching control would still be spending the rate
limit on a feature nobody can see.

Enter with no active option falls through to the form's submit handler
and searches the typed text exactly as before. The existing behaviour is
preserved, not replaced, and that has its own test.

Suppressed once the value parses as a postcode: the box takes a name OR
a postcode, and suggesting schools during postcode entry fights the user.

.omniBoxContainer gains position: relative — the dropdown is absolutely
positioned and without it would have anchored to the page instead.

Four render sites, all wired: page.tsx renders HomeView in the success
path AND the catch fallback, and HomeView renders FilterBar as hero AND
sticky. Missing any one would make the flag silently do nothing
somewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:38:52 +01:00
TudorandClaude Opus 5 d88e77f459 feat(suggest): the dropdown, with combobox ARIA
Presentational only — it fetches nothing and owns no state, so the
fetching rules and the ARIA rules can be read separately.

onMouseDown, not onClick. The input's blur handler closes the list and
blur fires before click, so a click handler never runs: the classic bug
where a dropdown works perfectly by keyboard and is dead to the mouse.

The plan's CSS guessed at token names like --color-surface. The real
tokens are --bg-card, --border, --text-muted, --bg-secondary and
--shadow-soft, and all five are redefined in the dark theme — invented
names would have silently fallen back to hardcoded light values and
broken dark mode.

Local authority is rendered because there are many schools called
'St Mary's'; a list without it is unusable for exactly the query
autosuggest exists to serve, which is what the test asserts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:36:39 +01:00
TudorandClaude Opus 5 06eb433db5 feat(suggest): debounced, abortable suggestion hook
The AbortController is correctness, not economy. Without it a slow
response for 'st' can land after the fast one for 'st marys' and replace
a correct list with a stale one — the classic autosuggest race.

No cache: 'no-store', unlike the compare modal's search. This is the one
endpoint where prefix queries repeat most across users, so discarding
the browser cache and the backend's ETag 304s would be throwing away the
cheapest win available.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:35:27 +01:00
TudorandClaude Opus 5 1a6d349dad feat(suggest): GET /api/suggest, cacheable and DataFrame-free
A dedicated endpoint rather than a mode of /api/schools, because that
path filters and sorts 25,000 pandas rows per query while holding the
GIL — affordable once per search, not once per keystroke. A test asserts
the distinction directly by making load_school_data raise and requiring
the endpoint to answer anyway.

Nothing errors on ordinary input: a short query, no matches, or
Typesense being down are all 200 with an empty list.

Cached deliberately. Prefix queries repeat enormously across users and
school names change once a year, so s-maxage plus the existing ETag
middleware turns most keystrokes into 304s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:34:34 +01:00
TudorandClaude Opus 5 75d3534d82 feat(suggest): Typesense rows for autosuggest, no DataFrame
search_schools_typesense returns URNs, which forces the caller to
hydrate from the 25,000-row in-memory frame. Every field a suggestion
needs is already in the Typesense document, so this returns documents
and the caller needs no pandas at all — the difference between a query
that can run per keystroke and one that cannot.

Never raises. Typesense unreachable or erroring gives an empty list,
because a dropdown that quietly stops appearing is the right failure for
a keystroke path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:33:42 +01:00
TudorandClaude Opus 5 ff041544f2 fix(api): rate-limit per caller, not per proxy
The limiter keyed on request.client.host, which in staging and prod is
the Next container — the backend has no published ports and nothing else
can reach it. So every browser user on the site shared one 60/minute
bucket per route. Measured against staging: 70 concurrent requests to
/api/schools returned exactly 60 OK and 10 refused, from one machine.

CF-Connecting-IP first. Cloudflare fronts both environments and
overwrites any client-supplied value, which a parsed X-Forwarded-For
chain does not guarantee. The XFF fallback is forgeable only from inside
the Docker network.

Named rather than hidden: the shared bucket was an accidental global
throttle on a single-process backend, and correct per-user keying
removes it. A real global ceiling belongs at Cloudflare, which is
already in the path; slowapi cannot express one without a second Limiter
and middleware this app does not install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:32:58 +01:00
TudorandClaude Opus 5 59265f78b6 docs(flags): Unleash does not create flags by itself
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m6s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 35s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 35s
The runbook said a flag 'appears in the Unleash UI after the backend has
evaluated it once'. That is wrong. SDKs read definitions from the server
and never register anything, and metrics for an unknown flag are
discarded — so a declared flag is evaluated on every request, stays
False forever, and never shows up until someone creates it by hand.

Found the way these things usually are: staging had been running the
flag code for a while and the UI was still empty.

Also names the environment trap while here — each stack's token is
scoped to one environment, so toggling the other does nothing visible
and looks like the flag is broken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:11:45 +01:00
TudorandClaude Opus 5 6e0a278340 docs(suggest): implementation plan, eight tasks
Self-review caught three defects in the plan. .omniBoxContainer, the
wrapper the dropdown positions against, does not declare position:
relative — without it the list anchors to the page. The postcode
suppression test typed character by character, so it would have asserted
no request while 'NW1' legitimately fires one; it now sets the value in
one go. And the Enter-submits-search assertion needed waitFor, because
updateURL pushes inside startTransition.

Task 1 is the one to review hardest: it is the only unflagged change and
it alters rate limiting for every endpoint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:48:39 +01:00
TudorandClaude Opus 5 e651dd0d65 docs(suggest): the in-app global ceiling would not have worked
Reading slowapi rather than assuming: default_limits and
application_limits are both evaluated with the same key_func, so they
are per-client across routes, not global. And application_limits only
apply 'if in_middleware' — this app installs no SlowAPIMiddleware, so
they would never have fired at all.

A genuine global cap would need a second Limiter with a constant key
plus that middleware. Cloudflare is already in the path on both
environments and does this at the right layer, so the ceiling is named
as a follow-up there rather than built badly here.

The risk that leaves is stated plainly in the risks section instead of
being papered over with a mechanism that does not do the job.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:44:21 +01:00
TudorandClaude Opus 5 22c113fc29 docs(suggest): design for school autosuggest
The load-bearing finding is not about autosuggest. The rate limiter keys
on request.client.host, which in staging and prod is the Next container
— so all browser users share one 60/min bucket per route. Measured
against staging: 70 concurrent requests gave exactly 60 x 200 and
10 x 429. Eight concurrent searchers would 429 the site once each
keystroke costs a request, so the keying fix is part of this work.

Both environments are behind Cloudflare, which sets CF-Connecting-IP and
overwrites any client-supplied value — trustworthy in a way a parsed
X-Forwarded-For chain is not, and the backend is unreachable except
through the Next proxy.

Named honestly: the shared bucket has been an accidental global throttle
on a single-process backend, so correct per-user keying removes a
protection. A global ceiling ships with it rather than instead of it.

Suggestions come from Typesense alone. The existing search path filters
a 25,000-row DataFrame per query, which is exactly the cost a keystroke
endpoint cannot pay, so there is deliberately no DataFrame fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:41:00 +01:00
tudor e953ee7c5f Merge pull request 'feat(flags): ship-dark feature flags, with last-distance-offered behind the first one' (#125) from feat/feature-flags into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 40s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m34s
Reviewed-on: #125
2026-08-23 11:34:56 +00:00
TudorandClaude Opus 5 413d86cc3c chore(flags): wire UNLEASH_URL and the SDK cache volume into the stacks
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 27s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m59s
Both variables default to empty, so an environment without Unleash has
every flag off — the correct dark state rather than a boot failure.

The cache volume is the mitigation for the one real regression risk in
this design: the SDK evaluates everything False until it syncs, so a
backend cold-starting with an empty cache while Unleash is unreachable
would make a *released* feature disappear. On a named volume the disk
cache survives a restart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:58:29 +01:00
TudorandClaude Opus 5 4f01fbdedb test(e2e): make the distance journeys fail loudly, not skip quietly
The existing distance journeys all skip when no school has a published
figure, which is right when the feature is off — and wrong when it is
supposed to be on and is silently broken, because that shows up as a
green run full of skips. The new gate fails in exactly that case.

Feature state is read from the data, not from /api/flags: the public
proxy denies that path on purpose, since it names unreleased features.
Presence of the admission_distance key is the observable effect.

Verified against staging, where the feature is currently on: the on-gate
passes, the off-gate skips, the existing eight distance journeys are
unaffected.

One honest caveat — the /api/flags check passes on staging today because
that image predates the endpoint, not because the denylist works. The
denylist itself is covered by the jest unit test; this is defence in
depth and becomes a real assertion once deployed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:57:34 +01:00
TudorandClaude Opus 5 c3ba7aae0d feat(flags): ship last-distance-offered dark behind a flag
One gate, at the source. The frontend needs no change: DistanceSection
already returns null when distance_m is missing, and the admissions
block already conditions on (admissions || admissionDistance). Only 57
local authorities publish cut-offs, so the off-path is the commonest
path on the site and is well covered already.

Absent, not null. /api/schools/ is public and unauthenticated, so a
field left in the payload is a published field — the reasoning already
recorded in c9a1892 when history was withheld. The two are also
different claims: null says this school has no cut-off, absent says
cut-offs are not being published at all. The frontend type now says so.

The feature is on main and live on staging and has never reached
production, which is what makes it the right first consumer: the flag
lets the code promote without the feature appearing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:56:33 +01:00
TudorandClaude Opus 5 54a30de0d8 feat(flags): server-side getFlags for the frontend
Ships without a consumer, deliberately. The first flag needs none — the
backend withholds the field and the page follows — but 'UI elements on
existing pages' is one of the three surfaces this capability exists for,
and a flag layer that cannot gate one is incomplete.

Never throws: an unreadable flag is a dark one, which matches the
backend's fail-closed default. A page that 500s because the flags
endpoint blinked would be a worse outcome than a hidden feature.

Reading flags pins the calling route to a 300s ISR floor, since Next
takes the lowest revalidate among a route's fetches. That matches what
/school/[slug] already sits at, and it is the same property that makes a
flip propagate without a webhook.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:55:32 +01:00
TudorandClaude Opus 5 c30ad1db07 feat(flags): serve /api/flags, and keep the public proxy off it
The endpoint and its exposure control ship together on purpose. The
moment /api/flags exists, app/api/[...path] forwards it — and the
response names every unreleased feature the codebase knows about, along
with whether it is on. Publishing that is the opposite of shipping dark.

Denied on an exact first-segment match, not a prefix, so /api/flagship
does not go down with /api/flags. Next reads the endpoint server-side
over the Docker network, which never transits the public proxy.

jest.setup.js now guards its browser globals. It runs for every suite,
including the one that declares @jest-environment node to exercise the
route handler — NextRequest needs Fetch API globals jsdom lacks, and
there is no window there to define matchMedia on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:54:26 +01:00
TudorandClaude Opus 5 7424cef7c6 feat(flags): the registry and a fail-closed Unleash client
Unleash holds flag state; it does not hold the list of flags. REGISTRY
is that list, because the SDK evaluates an unknown flag to False and
without a registry that is an undeclared False — indistinguishable from
a typo in a flag name.

Fail-closed throughout, and never raises: an unset UNLEASH_URL, an
unreachable server, a client that throws, an undeclared name — all
False. A flag layer that can 500 a request path or stop the API booting
is worse than one that is switched off.

Every flag defaults to False, with no per-flag override, because a flag
that defaults on is a kill switch and this is deliberately not one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:52:55 +01:00
TudorandClaude Opus 5 01ccbb8e82 feat(flags): add the Unleash stack and its runbook
Its own Portainer stack, belonging to neither application stack: a
staging redeploy must not be able to disturb production's flag state.

One instance serves both. OSS Unleash ships development and production
environments with environment-scoped client tokens, so the same flag
holds independent state in each — which is what lets a feature be on in
staging, where the E2E journeys exercise it, while production stays dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:51:35 +01:00
TudorandClaude Opus 5 c339c2f1a1 docs(flags): implementation plan, eight tasks
Task 1 is the Unleash stack and ends with a human step — the Portainer
deploy and the token generation cannot be automated from here. Nothing
else blocks on it: an unset UNLEASH_URL means every flag is False, which
is what local development and CI get, so the whole suite runs without a
flag server existing.

Self-review caught three defects in the plan itself. get_supplementary_data
takes (db, urn), not (urn), and the test DataFrame was minimised to the
point where the endpoint would have failed for reasons unrelated to
flags — both now copy the known-good shape from test_school_details.py.
The proxy test needs the node jest environment, since NextRequest wants
Fetch API globals jsdom does not provide. And the e2e off-state check
hardcoded a URN, so a 404 page would have satisfied it without proving
anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:46:15 +01:00
TudorandClaude Opus 5 e2ca3d79f9 docs(flags): drop the webhook — the seven-day premise was wrong
Next uses the LOWEST revalidate among a route's fetches, not the segment
value. School pages fetch school details at 300s and place pages fetch
national averages at 3600s, so the effective ISR period is five minutes
and one hour respectively — not the seven days the segment declares.

A flag flip therefore propagates on its own, well inside the monthly,
by-hand cadence these flags are for. That deletes two webhook
integrations, a revalidate route, a secret-in-query-string scheme, an
idempotency requirement, and the rule that every fetch carry a cache
tag — which was the part most likely to rot as fetches are added.

Two constraints survive: a flag must never gate content on a
force-static page, because app/admissions never revalidates; and a
route-family flag must rebuild the sitemap, deferred with the route case
since no flag in scope touches it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:40:35 +01:00
TudorandClaude Opus 5 c2364bf09e docs(flags): design for a ship-dark feature flag layer
Unleash self-hosted in its own Portainer stack, with FastAPI holding the
only SDK and Next reading flags through a tagged fetch.

The two hard parts are consequences of putting flag state in a service
rather than the repo: main stops being the whole truth about what is on,
and a flag can now change without the deploy that would have cleared the
caches. A code-declared registry bounds the first; webhook-driven
revalidateTag handles the second.

Cache tagging is deliberately coarse — every server fetch carries the
flags tag, not just the flags fetch itself. The first consumer proves
why: admission_distance changes the shape of /api/schools/{urn}, so a
narrow purge would leave ~25,000 school pages serving the pre-flip
render for a week, invisibly.

First consumer is the last-distance-offered feature, which is on main
and staging and has never reached production. It needs one gate, at the
API, because the frontend already no-ops on a missing field.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 09:47:58 +01:00
tudor 43e0621728 Merge pull request 'fix(places): phase links must stay in their own namespace' (#124) from fix/place-phase-links into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m31s
Reviewed-on: #124
2026-08-22 17:33:12 +00:00
TudorandClaude Opus 5 d1358cc00f fix(places): phase links must stay in their own namespace
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m37s
Every place page built its phase links as /schools/[slug]/[phase], the
shape that belongs to towns alone.

On an authority page that pointed into the town namespace. For 87 of the
151 authorities the target does not exist and the link 404s; for the
other 64 it resolves to the town of the same name — a different set of
schools, which is precisely the near-duplicate the two namespaces were
introduced to prevent. On an outcode page it 404s outright.

Two causes behind it, both a rule written twice and inherited by only
one of the places that needed it.

The authority phase route was in the spec and dropped by the plan, which
built the three bare routes and no fourth. The sitemap is generated from
the place registry, which was right about them all along, so 302
authority phase URLs have been submitted to Google and every one 404s.
Adding the route makes the sitemap true and serves a real query —
admissions are authority-run, so "primary schools in Kent" is how a
parent searches before they have settled on a town.

The outcode variants were the opposite: the registry computed phases for
outcodes although the spec gives them no route, and the sitemap knew to
skip them while the API did not. The registry now decides alone, and the
sitemap's duplicate of that rule is gone.

Also: an authority under the five-school threshold has no page, so the
API sends a null slug for it and the page names it without linking.
Two English authorities are in that position. It was unreachable in
today's data — verified across the EC and TR outcodes — but the thin
place redirect would have sent a reader to a 404 the year it isn't.

The e2e journey now walks every /schools link a page of each family
emits and requires a 200, which is the check that was missing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 17:22:11 +01:00
tudor 865a69b54d Merge pull request 'feat(places): list schools alphabetically on place pages' (#123) from feat/place-alphabetical-sort into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m27s
Reviewed-on: #123
2026-08-21 23:22:49 +00:00
TudorandClaude Opus 5 9cc87c41bb fix(places): a phase page needs results, not merely publishable schools
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m4s
Asked where schools with no results should sit in an alphabetical list, and
found that some pages were almost entirely made of them.

The per-phase threshold counted schools that were publishable — a result OR
an Ofsted grade — while a phase page exists for its results column.
/schools/kent/primary published with none of its five rows carrying a result;
Minehead had one of seven, Buntingford one of five. Forty-four phase pages
were majority-blank.

It is the same rule as "no page without a local average", which was written
into the spec as a thin-page control and never extended per phase.

The threshold now counts schools with a result for that phase. It gates
whether the page exists; it does not filter rows — a page that publishes still
lists every school of the phase, because someone looking up a school by name
has to find it whether or not it published results.

126 of 1,012 variant pages stop publishing: 62 primary, 64 secondary. Every
one of them was a table with too little in it to be worth a page.

The ordering itself is unchanged: pure A-Z, blanks interleaved. A school sits
where its name says it does, and at roughly a tenth of rows that reads fine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:14:31 +01:00
TudorandClaude Opus 5 8967966eef feat(places): list schools alphabetically on place pages
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m21s
Someone on a place page is usually looking for a school they can name, so the
order should serve scanning for it rather than ranking. /api/rankings keeps
its league-table ordering; this is a place-page decision, not a site-wide one.
Sorted case-insensitively, or a capitalised name would sort ahead of every
lowercase one.

The change made five pieces of copy untrue, so they go with it. The phase
variant titled itself "— Ranked", and all four route families described
themselves as "ranked by SATs and GCSE results". A page that opens by claiming
an order it does not keep is worse than one that claims nothing.

The ItemList markup carried `position` with no declared order, which reads as
a ranking. It now declares ItemListOrderAscending, so the structured data says
what the table does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:09:58 +01:00
tudor 4a9a5c734b Merge pull request 'fix(e2e): three assertions that were wrong about correct behaviour' (#122) from fix/e2e-canonical-and-robots into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m26s
Reviewed-on: #122
2026-08-21 23:07:57 +00:00
TudorandClaude Opus 5 4e82e6c916 fix(e2e): three assertions that were wrong about correct behaviour
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 41s
The staging gate was red on three journeys. All three were faults in the
tests; the site was behaving correctly in each case.

Next normalises canonical URLs against trailingSlash:false, so the homepage
ships "https://www.schoolcompare.co.uk" with no slash while every other route
keeps its path. Both address the same document. The test hardcoded the slash
and so failed only on the root — /rankings and /admissions passed throughout,
which is what made it look like a homepage bug rather than a test bug.
Compared with trailing slashes stripped from both sides.

The robots.txt assertion matched "Disallow: /" anywhere in the file and
tripped over the AI-crawler groups Cloudflare injects — ClaudeBot, GPTBot,
Amazonbot and six others all carry a blanket disallow, deliberately, and none
of them is Googlebot. It now parses the file into user-agent groups and checks
only the "*" group, which is also the thing the test was always trying to say:
Google may crawl the page, so it can see the noindex header.

Both were the same mistake as the doubled brand: asserting a naive string
rather than the semantics, and asserting against what the code assembles
rather than what the page renders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 23:51:34 +01:00
tudor d4340a8fdd Merge pull request 'feat(places): name every authority a place sits in' (#121) from feat/place-multiple-authorities into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #121
2026-08-21 21:56:50 +00:00
TudorandClaude Opus 5 bb2f7a5841 fix(places): address review, and merge places GIAS spells more than one way
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m34s
Two findings from review on #121, plus a third the review prompted.

The cap at three authorities silently dropped the fourth in exactly the case
where the information matters most — a genuinely fragmented place — and
contradicted the stated goal of naming every authority a place sits in. It is
gone. The share rule was always the real limit and already bounds the list at
ten. Measured against the live corpus, one town would have been truncated
today: LONDON, split evenly between Hackney, Lambeth, Westminster and
Lewisham.

parent_authority used mode() while authorities used value_counts(), and on an
exact tie pandas does not guarantee the two pick the same name, so the 301
could have pointed somewhere other than the authority named first on the page.
The parent is now derived from authorities[0]: one computation, one answer.
It also inherits the sentinel filter, so a place can no longer redirect to
/schools/authority/does-not-apply.

Chasing the truncation case surfaced a worse bug. Places were grouped by raw
town value, but the registry is keyed by slug, and GIAS spells the same place
several ways. Five town slugs come from more than one spelling: "London"
(1,819 schools) and "LONDON" (12) both slugify to `london`, so the later group
simply overwrote the earlier one — /schools/london could have shown twelve
schools, silently, depending on row order. Weston-super-Mare was split 14/19
across two spellings and Newcastle-under-Lyme across three. Grouping is now by
slug, and the display name is the most common spelling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:42:18 +01:00
TudorandClaude Opus 5 1cb5314c53 feat(places): name every authority a place sits in
SW19 is mostly Merton but partly Wandsworth, and the page said only Merton.
The cause was one field doing two jobs: _parent_authority takes the modal
authority, which is right for a 301 target and wrong as a statement about
where a place is.

This is not a corner case. A quarter of viable outcodes (425 of 1,760) and a
third of viable towns (263 of 783) cross an authority boundary — Bedford the
town spans Bedford and Central Bedfordshire.

Place now carries `authorities`, every authority holding at least a tenth of
the schools and at least two of them, largest first. parent_authority stays
single and unchanged, because a redirect still needs one target.

The share threshold exists because GIAS carries postcode errors: EN6 lists two
Shropshire schools among fourteen in Hertfordshire, and a bare "any authority
present" rule would print those as though they were real. A place too small or
too fragmented to clear the threshold still names its largest, so the page
never goes silent about where it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:40:58 +01:00
tudor 4cea26b813 Merge pull request 'fix(places): align the measure column's heading with its values' (#120) from fix/place-table-alignment into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #120
2026-08-21 21:21:03 +00:00
TudorandClaude Opus 5 dbb74d9b60 fix(places): align the measure column's heading with its values
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 14s
The heading sat on the right edge of the column and every value on the left.
A specificity collision, not a layout problem: the two were aligned by
different selectors and only one of them won.

  .table td            (0,1,1)  text-align: left    <- won for the value
  .num                 (0,1,0)  text-align: right   <- lost
  .table th:last-child (0,2,1)  text-align: right   <- won for the heading

The heading and the value cell now share one class and one rule, so they
cannot drift apart again whatever else changes around them.

The column also stretched to half the table. It now hugs its content with
width:1% and nowrap, so the school name takes the remaining width — which is
what made the gap read as misalignment on a wide screen, and what crowded the
name column on a narrow one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:15:02 +01:00
tudor 9545aec7f4 Merge pull request 'fix(places): phase-grouped tables, plain-English measures, styled links' (#119) from fix/place-presentation-to-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 56s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #119
2026-08-21 20:57:13 +00:00
Tudor 3365ebcb3a fix(places): phase-grouped tables, plain-English measures, styled links
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Three presentation faults on the place pages, all found by looking at a
rendered page rather than at a test.

An unphased place page showed one primary-only measure for a list holding both
phases: 8 of 27 rows on /schools/brentwood were blank, because secondaries
have no reading-writing-maths score. Picking the other measure would only have
inverted which rows were empty, and putting both in one column would have
mixed a percentage with a 0-90 score. Each phase now gets its own table, so a
blank cell means the school genuinely has no published result — which is worth
saying, and now says "Not published" rather than a bare dash.

"RWM expected" was invented here. The site already names the measure in
METRIC_DEFINITIONS, surfaced at /api/metrics: "Reading, Writing & Maths
Combined %". The heading now reads "Reading, writing & maths" with the full
definition in the tooltip.

Links carried no class at all, so they rendered as default blue underlined
browser links beside a site that styles table links as body colour with a
brand hover. They now follow RankingsView's convention, and running-copy links
take the brand colour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 21:56:06 +01:00
tudor 6d79bd3331 Merge pull request 'fix(places): stop the place titles doubling the brand' (#117) from fix/place-title-brand-doubling into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #117
2026-08-21 20:53:14 +00:00
TudorandClaude Opus 5 24e114dee7 fix(places): stop the place titles doubling the brand
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 27s
Every place page shipped as 'Schools in Brentwood - Compare 27 Schools |
schoolcompare | schoolcompare'. The root layout's title template appends
'| schoolcompare' to any plain-string title, and all four place routes already
carried the brand. W8 opted the other routes out with an absolute title; the
place routes were written afterwards and did not inherit the lesson.

~2,600 titles affected, and the repetition pushed them past Google's
truncation point, so the doubled brand displaced real words in the result.

An e2e journey now asserts no title repeats the brand, across the static
routes and a place page, so this cannot come back on a route added later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 21:43:32 +01:00
tudor 6c5db0c266 Merge pull request 'fix(places): submit and link the phase variants' (#116) from fix/place-phase-variants into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m29s
Reviewed-on: #116
2026-08-21 19:46:55 +00:00
TudorandClaude Opus 5 6f749ed21f fix(places): submit and link the phase variants
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m39s
/schools/[place]/[phase] shipped as routes but reached nothing. The sitemap
emitted one URL per registry entry and the registry had no phase dimension, so
~950 pages were absent from every sitemap — and PlaceView did not link them
either, leaving them reachable by nothing at all.

That is the query shape the baseline actually showed: 'primary schools in
beccles', 'secondary schools in brentwood'. Publishing the routes without a
path in meant building for the demand and then hiding from it.

Place now carries phase_urns so the per-phase threshold can be applied without
re-querying, the sitemap emits a variant wherever a phase clears the threshold
on its own, and the API exposes the qualifying phases so the place page links
only variants that exist. Outcodes are excluded: nobody searches 'primary
schools in SW11' and those routes do not exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 20:43:07 +01:00
tudor d423826840 Merge pull request 'fix(places): a locality collision must not break the sitemap' (#115) from fix/locality-collision-skip into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m38s
Reviewed-on: #115
2026-08-21 19:26:02 +00:00
TudorandClaude Opus 5 d3c63ccc6d fix(places): a locality collision must not break the sitemap
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 34s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 34s
Sitemap regeneration failed on staging. 'richmond' in the curated locality
list collides with the GIAS town Richmond in North Yorkshire (37 schools), the
registry raised, and the admin endpoint 500d — taking down sitemap generation
for all 25,000 school pages over one bad row of curated data.

The guard now skips the colliding locality and logs an error. Skipping still
achieves what the guard was for — a locality never silently shadows a town —
without letting curated data break the site. That matters beyond this bug:
GIAS town names change with no code change here, so a raise could fire
spontaneously in production later.

Also removes four localities that were London boroughs rather than districts.
Hackney, Islington, Greenwich and Ealing are local authorities with 104, 72,
108 and 115 schools and already have authority pages; a locality defined by
two or three outcodes would have been a partial near-duplicate of one — the
thin-content failure the two-namespace design exists to avoid. A test now
guards the whole borough list.

Validated against the live corpus: 15 localities, no town collisions, no
authority duplicates, all 15 clear the threshold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 20:20:52 +01:00
tudor b93eb3a691 Merge pull request 'feat(seo): the location layer — town, locality, authority and outcode pages (W2)' (#114) from feat/w2-location-layer into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m18s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 4s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 11m8s
Reviewed-on: #114
2026-08-21 17:56:05 +00:00
TudorandClaude Opus 5 6b871ce1e9 feat(places): ItemList and BreadcrumbList, and the e2e gate
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 4m33s
ItemList tells Google the page is a ranked set rather than prose;
BreadcrumbList puts the place in a hierarchy. School URLs in the markup are
absolute on the canonical host, since a relative URL in JSON-LD is ambiguous.

Eight journeys covering all four families, the two-namespace guarantee, the
threshold, the canonical, the sitemap and the local-versus-England line — the
last because that comparison is the reason these pages are not a name dropped
into a template.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:19:08 +01:00
TudorandClaude Opus 5 c981d89137 feat(places): town, locality, authority and outcode routes
Every generateStaticParams is gated behind PRERENDER_PLACES and wrapped in the
same try/catch the school route uses. The plan claimed authority pages were
'few enough to always prebuild' — but few enough still means the API must be
reachable at build time, and in CI it is not: the build failed with
ECONNREFUSED rather than degrading to ISR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:17:57 +01:00
TudorandClaude Opus 5 de5e790112 feat(places): place page client and view component
One component for all four families: they differ in what fills the registry,
not in what the page shows, so a second would be a second place to forget the
same change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:15:10 +01:00
TudorandClaude Opus 5 42138fc402 feat(places): submit place and outcode sitemaps
Separate children per family so Search Console reports the location layer's
indexation apart from the school pages' — which is the point of the index
built in W1, and the number the stop condition watches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:13:55 +01:00
TudorandClaude Opus 5 c5af476213 feat(places): /api/places registry and place detail endpoints
The registry is cached for the process and reset by the same admin endpoint
that rebuilds the sitemaps, so places and sitemap always describe the same
corpus rather than drifting apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:13:05 +01:00
TudorandClaude Opus 5 de853b90b3 feat(places): London localities and postcode districts
The GIAS town field puts 1,819 London schools under the single value
'London', so it cannot answer 'schools in Battersea' — a query that appears in
the baseline. No single field can: parliamentary constituency gives Battersea
but not Canary Wharf, admin_ward gives Canary Wharf but not Battersea, and
neither gives Clapham or Shoreditch. So a locality is curated, defined by the
postcode districts it covers, which needs no new ingestion.

A locality may not shadow a published town: the registry raises rather than
silently costing a page that carries real demand. One below the threshold is
logged rather than raising, because a locality can legitimately be too small.

The pipeline seed mirrors the module, with a test guarding the drift — the
same arrangement gias_codes has, and for the same reason: the backend image
does not contain pipeline/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:11:56 +01:00
TudorandClaude Opus 5 759d9f5cea feat(places): registry of towns and authorities
Two namespaces because 67 town names collide with an authority name and
neither set contains the other — postal towns cross authority boundaries, so
Bedford the town holds 104 schools against the authority's 86.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:10:37 +01:00
TudorandClaude Opus 5 555d3f0a7d docs(seo): implementation plan for the W2 location layer
Seven tasks: the place registry, London localities and outcodes, the places
API, per-family sitemaps, the shared place view, the four route families, and
structured data plus the e2e gate.

Two things the plan corrects against the spec. The backend image does not
contain pipeline/, so the curated locality list cannot live only in a dbt
seed — it follows the gias_codes.py precedent instead, canonical in backend
with the seed as a mirror. And NationalAverages is nested by phase rather than
flat, which the first draft read wrongly and would have rendered every page
without its England comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:09:35 +01:00
TudorandClaude Opus 5 ecc847091c docs(seo): design for the W2 location layer
Supersedes the original spec's W2. The Search Console baseline inverted its
ordering: every measured location query is town or district level, none is an
administrative area, and phase is part of the query rather than a filter.

Two problems the original design did not anticipate. 67 viable towns share a
name with a local authority, and the authority is the larger set in only 43 of
them — postal towns cross authority boundaries, so neither can absorb the
other. Two namespaces resolve it by construction. And the GIAS town field
collapses 1,819 London schools into one value, which a curated
locality-to-outcode seed solves without new ingestion.

Sizing is measured against the live 25,185-school corpus rather than
estimated: 783 viable towns, 1,760 outcodes, 154 authorities.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:09:35 +01:00
tudor b187a478c9 Merge pull request 'feat(seo): rewrite the C1 snippets to earn the click (W8)' (#113) from feat/seo-metadata-c1 into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #113
2026-08-20 23:26:12 +00:00
TudorandClaude Opus 5 c0547c45e5 feat(seo): rewrite the C1 snippets to earn the click (W8)
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 11s
The baseline says these pages already rank and are not clicked. 'compare
school performance' sits at position 6.1 with 0.43% CTR; 'compare schools' at
7.2 with 0.87%. The brand query 'school compare' draws 9.16% from the same
neighbourhood of the same results page, which rules out a ranking explanation
— when the snippet gives a reason to click, it gets clicked.

These SERPs are owned by the DfE's own 'Compare school performance' service.
The old title put a lowercase brand nobody searches for in the most valuable
pixels, then a near-paraphrase of that service's name. Beside the government's
own result it read as a lookalike.

Intent in the title, differentiator in the description. Titles now match what
people type, and the descriptions carry the one fact gov.uk does not publish:
how close you had to live to get a place.

/compare deliberately takes the tool phrasing rather than the homepage's, so
the two pages stop competing for one phrase. The root layout's default and
Open Graph copy were saying something different again; they now agree.

No hard school counts in any of it. The corpus moves with every data refresh
and this repo has already shipped one copy bug of that kind.

Tests guard the mechanics — SERP length, intent keyword, the differentiator,
no brand-first title — and leave the wording free to iterate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:21:08 +01:00
tudor 4a3928df9f Merge pull request 'fix(seo): a school is publishable on any year's results, not the latest' (#112) from fix/sitemap-any-year-data into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #112
2026-08-20 23:15:19 +00:00
TudorandClaude Opus 5 07c97a46c5 fix(seo): a school is publishable on any year's results, not the latest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 53s
_school_sitemap_rows tested only the latest year's row, which quietly dropped
every school with results in its history but a null row for the most recent
year — a school that stopped reporting, or whose figures were suppressed for
small-cohort disclosure.

The Mallard Academy (150367) is the case that caught it: real KS2 results for
2015-16 through 2018-19, then null rows from 2022-23 on. Its detail page shows
all four years; the sitemap omitted it. Sampling 40 of the 2,206 excluded
schools found 4 like this, so roughly 220 real pages were being withheld.

Publishable is now a property of the school, computed across every row, while
lastmod still comes from the latest row so the most recent Ofsted date wins.
The field list is a module constant shared with _has_publishable_data so the
two checks cannot drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:04:32 +01:00
tudor bb81337aba Merge pull request 'fix(seo): keep staging out of the search index' (#111) from fix/staging-noindex into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m1s
Reviewed-on: #111
2026-08-20 22:39:38 +00:00
TudorandClaude Opus 5 b34511e459 chore: record the branch cleanup manifest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m30s
79 remote branches deleted: 77 fully merged into main, plus
feat/seo-crawl-hygiene and feat/england-only-corpus, whose content is
preserved on feat/seo-crawl-hygiene-main (PR #110).

Each line carries the SHA, so any branch can be restored with
  git push origin <sha>:refs/heads/<name>

The 14 branches left standing all carry content that differs from main and
none of them is mine to judge abandoned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:23:30 +01:00
tudor f928a15c1e Merge pull request 'fix(seo): crawl hygiene and a per-family sitemap index (W1)' (#110) from feat/seo-crawl-hygiene-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m21s
Reviewed-on: #110
2026-08-20 22:23:10 +00:00
TudorandClaude Opus 5 1fc1e07d21 fix(seo): keep staging out of the search index
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 13s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
Staging serves the same image as production off stx., with robots.txt saying
Allow: / and no noindex — a fully crawlable duplicate of the site. Nothing
appears indexed today, most likely because the pages canonicalise across to
production, but that is a side effect rather than a control.

X-Robots-Tag, not a robots.txt Disallow. Disallow blocks crawling, which is
not the same as blocking indexing: a disallowed URL can still be indexed from
external links, and blocking the crawl means Google never fetches the page and
so never sees a noindex at all. Staging stays crawlable and answers noindex.

Matched on the staging host explicitly rather than 'any host that is not
production'. The inverted form would cover future environments automatically,
but its failure mode is deindexing production if the Host header ever arrives
rewritten by a proxy — which cannot be verified from here. This form's failure
mode is a new environment being indexable until someone adds it, which is
recoverable. Any new non-production hostname must be added.

The journeys only ever run against staging (deploy.yml passes
STAGING_BASE_URL; promote.yml smoke-polls production without Playwright), so
asserting the header there is safe. The two assertions live in one test
because the halves only work together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:21:58 +01:00
TudorandClaude Opus 5 2208ad93c1 feat(seo): split the sitemap into a per-family index
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m2s
Search Console reports coverage per submitted sitemap, so one file per page
family is what will make W2's location pages measurable when they land. The
index's lastmod is generation time, which is the correct semantic there —
unlike on a <url>, where it would be a claim we cannot support.

Children sit under /sitemaps/ because Next only treats a whole bracketed path
segment as dynamic; a route folder named sitemap-[...parts] would be read as a
literal static segment and never match. Confirmed by the build output, which
lists /sitemaps/[...parts] as a dynamic route.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
TudorandClaude Opus 5 24f3cb4c65 fix(seo): submit only school pages that have something to show
Drops the schools with neither results nor an Ofsted grade, adds /admissions
which was never listed, replaces the invented priority and changefreq with a
lastmod taken from each school's Ofsted date.

lastmod is omitted where no date is known rather than defaulted to now. An
always-now lastmod is a claim Google learns to distrust; absent honestly
means unknown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
TudorandClaude Opus 5 3816b92d06 fix(seo): noindex parameterised comparisons, keep bare /compare
25,193 schools make ~317 million pairs. The bare page stays indexable as the
landing page for the head term; the parameter space goes noindex, follow so
its outbound links still count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
TudorandClaude Opus 5 51ce2d6373 fix(seo): declare a canonical on every route
The homepage read eleven search params and declared no canonical, so every
filter combination was a crawlable near-duplicate of the page we most want to
rank. Rankings and admissions declared none either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
TudorandClaude Opus 5 5ddc7314fd fix(seo): canonicalise on the www host, which is the one that serves 200
The apex 301s to www at Cloudflare, but metadataBase, the school-page
canonical, robots.txt's Sitemap: line and the sitemap's own <loc> entries all
named the apex. Every one of those pointed Google at a redirect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
tudor 8ad2070e76 Merge pull request 'fix(data): drop the Welsh school GIAS does not type as Welsh' (#109) from fix/welsh-establishment-leak into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m18s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m25s
Reviewed-on: #109
2026-08-20 21:59:38 +00:00
TudorandClaude Opus 5 3aad5101a8 fix(data): drop the Welsh school GIAS does not type as Welsh
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 38s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 54s
Beechwood College (URN 142458, Sully, CF64 5SE) survived the England-only
filter. GIAS types it a Special post 16 institution (32), not a Welsh
establishment (30), so filtering on establishment type alone left it behind —
the last Welsh school on the site, and the reason Vale of Glamorgan was still
in the authority list.

The earlier verification claimed the type codes mapped onto the Welsh
authorities in both directions. That was checked exhaustively for Cardiff and
by count for three others; Vale of Glamorgan was never checked, and it was the
one that did not hold.

LA code is the reliable discriminator: GIAS gives the 22 Welsh unitary
authorities the contiguous block 660-681, which English authorities never use.
Filtering on postcode would have been wrong — Redbrook, Tutshill, Wyedean and
two other Gloucestershire schools carry NP16/NP25 postcodes because Royal Mail
areas straddle the border, and they are English schools with English data. An
e2e test now pins those five so the fix cannot be simplified into a postcode
filter later.

The 660-681 range is documented GIAS structure this project cannot verify from
its own data, so the range filters and the authority NAME checks: if the range
is ever wrong, a Welsh authority reappears in assert_england_only_schools and
the pipeline fails loudly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 22:54:24 +01:00
tudor c47fe38971 Merge pull request 'feat(data): publish England only, dropping Welsh and overseas establishments' (#107) from feat/england-only-corpus into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m15s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m27s
Reviewed-on: #107
2026-08-20 21:20:13 +00:00
tudor 8f211577c8 Merge pull request 'fix(e2e): scope the Distance-tab assertion, and separate the tile figures' (#106) from fix/e2e-distance-locator into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m26s
Reviewed-on: #106
2026-08-20 21:18:28 +00:00
TudorandClaude Opus 5 69f2201244 docs(seo): correct W1 plan's sitemap child routes
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 35s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
Next only treats a whole bracketed path segment as dynamic, so the planned
app/sitemap-[...parts]/route.ts would have been read as a literal static
folder and never matched. Children move under /sitemaps/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 22:09:00 +01:00
TudorandClaude Opus 5 e6048c9ca6 docs(seo): implementation plan for W1, crawl hygiene and sitemap
Five tasks: one canonical host, canonicals on every route, noindex on
parameterised comparisons, a sitemap that drops dataless schools and invented
priorities, and a per-family sitemap index.

Planning turned up a fault the spec had missed: the apex 301s to www, but
metadataBase, the school-page canonical, robots.txt's Sitemap: line and every
sitemap <loc> named the apex. Task 1 fixes it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 22:05:00 +01:00
TudorandClaude Opus 5 7650b16f62 feat(data): publish England only, dropping Welsh and overseas establishments
GIAS ships the whole UK plus overseas and offshore establishments. None of
them carry comparable DfE performance data — Wales does not publish on the
English measures at all — so every one of these pages rendered with null
results, null Ofsted and null phase. There were 2,036 of them: 1,569 Welsh,
123 offshore (Jersey, Guernsey, Isle of Man, Gibraltar), 316 British schools
overseas and 28 service children's schools. All 2,036 were being submitted to
search engines, alongside 29 local authorities that existed in the filters
purely to list them.

Filter at the mart boundary rather than the view layer. dim_school and
dim_location both exclude TypeOfEstablishment in {25, 26, 30, 37}, listed once
as vars.non_england_school_type_codes. Everything downstream reads those two
marts — search, the school page, /api/filters, rankings, Typesense and
build_sitemap() — so one filter removes them from the site and the sitemap
together, and Typesense drops them on its next rebuild since it recreates the
collection and swaps the alias rather than upserting in place.

coalesce rather than a bare NOT IN: a null type code would make the predicate
null and drop the row silently, and an unknown type is not grounds for
exclusion. No establishment has a null type today, but a future GIAS refresh
could ship one and the loss would be invisible.

assert_england_only_schools guards both directions: no excluded type survives
in dim_school, and dim_location holds no URN dim_school lacks — the API
inner-joins them, so the two filters drifting apart would silently shrink the
corpus.

Corpus goes from 27,229 schools to 25,193, and the authority list from 182 to
153. The 1,569 Welsh URLs now 404.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 21:46:28 +01:00
TudorandClaude Opus 5 9abd020967 fix(e2e): scope the Distance-tab assertion, and separate the tile figures
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 59s
The failing journey was wrong; the page was correct.

  Locator: getByRole('button', { name: 'Distance' })
  Expected: 0   Received: 1

getByRole matches accessible names by case-insensitive SUBSTRING, so
{ name: 'Distance' } matched <button>Show this distance on a map</button> —
the map toggle added by the same feature. The assertion was meant to say "no
Distance tab in the admissions segmented control" and instead said "no button
anywhere whose label contains the word distance".

Now scoped to the control it is about, via its own aria-label, and read
positively: the tab list must contain "This year" and must not contain
"Distance". An absence check against an unscoped locator passes for the wrong
reason the moment the selector stops matching, which is exactly how the
regression this test guards would return unnoticed.

Two sibling locators had the same weakness and are tightened: 'Check' is a
prefix of the button's own busy label "Checking…", and the figure matcher
accepted `(miles|m)` — a leftover from the mixed-unit era that would have kept
passing if the headline regressed to metres, which is the thing #105 just
fixed.

Tightening that matcher surfaced a real defect behind it. The cut-off figure
and its metric support are flex children with the gap drawn by CSS and nothing
between them in the text layer, so the element read "0.88 miles1.4 km" —
what a screen reader announces, and why a `miles\b` boundary could never
match. Both templates now carry an explicit space. Whitespace text nodes are
not rendered as flex items, so the reading changes and the layout does not.

Verified against staging: 54/54.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
2026-08-20 21:01:28 +01:00
tudor 228eb214f5 Merge pull request 'fix(admissions): state distance in miles throughout, never mixed' (#105) from fix/standardise-distance-units into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m28s
Reviewed-on: #105
2026-08-20 19:39:18 +00:00
TudorandClaude Opus 5 ea5249a2ea fix(admissions): state distance in miles throughout, never mixed
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 58s
The postcode check read "69 m away — inside the September 2026 cut-off of
0.17 miles". Both numbers are right and the sentence is still useless: the
reader has to convert one of them to check a comparison we had already made
for them.

The cause was a readability rule of mine in formatCutoffDistance, which swapped
to metres below 100m on the grounds that "0.04 miles" carries less than "69 m".
Taken one figure at a time that holds. Taken in a sentence containing two
figures it guarantees a mismatch whenever they fall either side of the
threshold — and a 270m cut-off with a nearby home does exactly that.

Miles now lead everywhere. It is the unit UK school admissions runs on:
councils publish cut-offs in miles (90% of the collected source rows), and it
is what a parent has already been quoted in their booklet and offer letter.
The metric figure survives only as support beside the miles figure on the
Admissions tile, where it converts the same value rather than presenting a
second one to compare.

Below 0.01 miles the decimal places run out rather than the unit being wrong,
so a very short distance is described — "under 0.01 miles" — instead of
rounding to a flat "0.00 miles", which would read as no distance at all.

Both figures in the verdict now go through one formatter with no fallback that
could reach for another unit. The old `?? "N m"` fallbacks on that line were a
second route to the same defect and are gone.

Covered by a sweep over sixty home/cut-off combinations spanning the old
switch point, asserting no verdict contains a metric reading and that exactly
two miles figures appear; plus the reported case pinned verbatim, and an e2e
guard on the rendered verdict.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
2026-08-20 20:35:05 +01:00
tudor ffe7e04951 Merge pull request 'feat(admissions): publish the latest cut-off only, holding history back' (#104) from feat/latest-cutoff-only into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m29s
Reviewed-on: #104
2026-08-20 17:51:18 +00:00
TudorandClaude Opus 5 c9a1892bfb feat(admissions): publish the latest cut-off only, holding history back
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m17s
Earlier years are to become a paid feature, so they stop being published.

The load-bearing part is that this is a change to the API, not only to the
page. /api/schools/{urn} is public and unauthenticated: leaving
admission_distance_history in the payload while declining to render it would
have handed the whole record to anyone who opened the network tab. It is
withheld at the source, and the page follows.

Nothing changes upstream. The tap, the plausibility band and
fact_admission_distance are untouched and still load every published year, so
restoring history for entitled callers is a change to one function in
data_loader rather than a re-collection.

What the reader now gets is the latest figure on the Admissions tile, and a
Distance section that answers the question the number alone cannot: whether
their own address falls inside it. Retitled to "How far away are you?", which
is what it now does — the previous title described a record that is no longer
there.

Removed with the history: the trend chart, the year-by-year table, the
per-year verdict strip, the trend summary and the coverage note, along with
their CSS. The section goes from 743px to 417px.

One consequence worth naming. A run of years used to soften a single close
call — a home just outside one year's cut-off was usually inside another. With
one year published, the "too close to call" band is the entire safety margin
between a parent and a place they do not have, so the verdict now names its
year, and the three outcomes are tinted apart rather than distinguished by
wording alone.

The existing stylesheet test earned its keep here: the three verdict classes
were referenced before they were written, and it caught them. Unstyled, a
"beyond the cut-off" result would have been indistinguishable from an "inside"
one — the exact failure the longhand class map was written to prevent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
2026-08-20 18:44:57 +01:00
tudor 50b599a09b Merge pull request 'fix(admissions): move the cut-off detail into its own section' (#103) from fix/admissions-section-height into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 52s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m29s
Reviewed-on: #103
2026-08-20 14:03:01 +00:00
TudorandClaude Opus 5 94151c58ea fix(admissions): move the cut-off detail into its own section
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 13s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m59s
The admissions card measured 1503px on a live school page — half the height
of every section put together, and nearly three times the next largest — with
its default view rendering as four tiles adrift in about 1080px of blank card.

The cause was a layout trick meeting content it was never sized for. The
admissions views are stacked in one grid cell so switching them never shifts
layout, and the hidden ones keep their box: only visibility is dropped. That
works while the views are comparable. The distance view added in #102 carries
a chart, a table and a map, came to 1402px against the tile grid's 316px, and
pinned every other view to its height — including the one that renders by
default, which nobody had clicked.

Rather than only unpinning it, the detail moves out. Every other topic on the
page is a section with a nav entry, and "how close did we need to live, and
would we have got in?" is a topic, not a variant reading of the intake
figures. The headline number stays on the Admissions tile where the intake
story is; the record behind it now lives in a Distance section directly below.

  admissions   1503px -> 554px
  distance        new -> 743px   (median section on the page is ~528px)

Three further changes, each of which also makes the content better rather
than only shorter:

  * The map renders on request. Before a postcode is entered it is a circle
    drawn round a school, and it costs a Leaflet bundle and 240px to say so;
    a successful check opens it automatically, which is the point at which it
    starts answering something. Map height 320px -> 240px.
  * The chart appears only at the four published years that let the summary
    state a direction. Below that we already refuse to call the series a
    trend, and a line through three points asserts one regardless of what the
    sentence beneath it admits. The table carries those years anyway, with
    the reasons a line cannot show.
  * Three caveat paragraphs become one. They said walking-route twice and
    made the same point about priorities in two voices. It now sits in
    CutoffDistanceDetail rather than inside the check, so it still renders
    for a school with coordinates missing, where there is a table but no map
    and no check.

The new e2e guard asserts no section exceeds 2.5x the median section height.
Measuring the card's internals cannot catch this: the tile grid is flex: 1,
so it absorbs the stretch and every box still looks full. The first version
of this test targeted an arbitrary primary, passed against the live bug, and
proved nothing; pointed at a school that actually holds cut-off history it
fails on staging with "#admissions is 1459px against a 526px median".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
2026-08-20 14:54:02 +01:00
tudor 8bf6145a73 Merge pull request 'feat(admissions): add the cut-off history, map and postcode check' (#102) from feat/last-distance-offered-full into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m14s
Reviewed-on: #102
2026-08-17 07:51:14 +00:00
TudorandClaude Opus 5 a72323874f feat(admissions): add the cut-off history, map and postcode check
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m39s
Completes the last-distance-offered feature against the mockup: the
year-by-year record, the same numbers drawn over real streets, and the
reader's own address measured against them.

Serving the history
  The first cut deliberately served only the latest year, because a plain
  series would draw a trend line straight through gaps that are absences of
  publication, not of a cut-off. That reasoning is answered rather than
  abandoned: cutoffYearRows classifies every year in the span, and the chart
  breaks the line rather than interpolating across it.

  A missing year is not one fact but three. It may be unpublished; it may be
  a year the school was not oversubscribed; or there may be no record at all.
  Collapsing them into "no data" throws away the reassuring case and hides
  the important caveat, so each is stated in words in the table.

  The claim is held to what the data supports. fact_admissions.oversubscribed
  compares FIRST PREFERENCES against places, which does not establish that
  every applicant was offered one — so the copy says "places available on
  first preferences" and a test asserts the stronger claim never appears.

The trend summary is not a verdict
  It names both endpoints and their years and lets the reader conclude. It is
  withheld below four published points, and a swing under a tenth of the
  earlier figure is reported as "broadly the same" rather than dressed up as
  a direction.

The postcode check
  This is the only place on the site that answers a question about a family
  rather than a school, so most of the care went into what it refuses to say.
  postcodes.io returns a centroid covering roughly fifteen addresses, which
  against a 500 m cut-off is a fifth of the whole distance — so a margin
  inside 100 m returns "too close to call" rather than a place a family does
  not have. Unpublished years count as unknown, never as a pass. The limits
  are stated before the check is used, not revealed with the answer.

  The postcode is geocoded in the browser and never stored.

Both templates
  Banded and selective secondaries are exactly where this matters most, so
  the detail is shared. The primary page gives it a third tab; the secondary
  page is one flat panel by design and renders it inline.

Absence is explained rather than reported. A selective school's missing
figure is explained by how it admits; a consistently undersubscribed school
reads as good news.

Also makes the batch loader's test double honour ORDER BY. It was a no-op,
so "latest row per URN" was really "first row in the fixture" and the test
would have passed with the sort reversed or removed.

Verified: 214 frontend tests, 54 backend, 45/53 e2e green against staging
(the 8 cut-off journeys skip until the DAG runs). Rendered offline against
the real compiled CSS in both themes and at 390px; every new surface clears
WCAG AA, measured on composited pixels.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
2026-08-16 13:42:06 +01:00
tudor 5a71f54d94 Merge pull request 'feat(admissions): show the last distance offered where councils publish it' (#101) from feat/last-distance-offered into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 43s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 3m15s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m53s
Reviewed-on: #101
2026-08-16 12:21:24 +00:00
TudorandClaude Opus 5 88c653215d feat(admissions): show the last distance offered where councils publish it
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m12s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m7s
Adds the cut-off distance a parent actually asks about — "how close do we
need to live?" — end to end: a Singer tap, dbt staging and mart models, an
Airflow DAG, and a tile on both detail templates. 3,597 schools across 57
local authorities carry a figure; the rest are unchanged.

There is no national source for this. Each LA publishes its own cut-offs in
its own format, and the collected CSV is transcribed from PDFs, spreadsheets
and web pages — so most of the work here is deciding what is safe to show.

Data
  * tap-uk-school-distance loads the CSV verbatim into raw. Keyed on
    (urn, year, school_name), because school_name carries the admission
    route: (urn, year) alone collides on 118 keys and a reload would have
    silently dropped every band but one.
  * stg_school_distance applies a 25 m – 25 km plausibility band. The source
    contains 0.0-mile rows (published where a school filled on a higher
    criterion), 1-metre cut-offs, and one reading 533 miles — ~4% of rows,
    all of which would put a visibly wrong number on a live page.
  * fact_admission_distance collapses routes to one row per school per year
    using the furthest, and keeps route_count so the page can say the figure
    is the widest of several bands rather than the one for a given child.

Serving
  * Kept out of fact_admissions: that mart is EES-derived and near-complete
    for England, this one covers 57 LAs, and the two refresh independently.
  * Latest year only. Coverage is ragged — a school may have 2021 and 2026
    and nothing between — so a history array would invite a trend line drawn
    through gaps that are absences of publication, not of a cut-off.
  * The Admissions section now renders on either source. 3% of the schools
    that render have a cut-off and no EES admissions row, and gating on
    admissions alone would have hidden the figure on those pages.

Interface
  * The year travels with the figure everywhere it appears; a cut-off
    detached from its admissions round is not a fact about anything.
  * "Not a fixed catchment — it moves every year" sits under every instance,
    because that is the inference a parent will otherwise draw.
  * Replaces a hardcoded "Historical distance cut-off data is not available
    for this school" that appeared on every secondary page, including the
    ones whose council does publish it. The absence is now stated only when
    it is real, and names the authority that would hold it.

The tint costs the muted tokens their AA margin: measured on the composited
backdrop (not the computed one, which reports the untinted card), --text-muted
falls to 4.09:1 in dark theme. The tile uses --text-secondary instead — 6.50:1
dark, 6.60:1 light.

The DAG is manual, like the other annual ones: councils publish on allocation
day, each on its own timetable, so there is no date worth scheduling against.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WDvkyqqHABm4bmth2kjAxE
2026-08-15 22:48:30 +01:00
tudor 5156a85bd1 Merge pull request 'chore: drop the hero byline, and refresh a figure in the UX audit notes' (#100) from chore/byline-removal-and-audit-figure into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m5s
Reviewed-on: #100
2026-08-15 09:25:59 +00:00
TudorandClaude Opus 5 fa1abff642 chore: drop the hero byline, and refresh a figure in the UX audit notes
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 6s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 14s
Two unrelated working-tree changes, committed as one at the owner's request.

Removes "Built for parents, by a parent." and its comment from the hero. It
was added two commits ago; taking it out is the owner's call, and the reasons
it was placed under the search rather than in the footer no longer apply.

The .heroByline rules in HomeView.module.css are deliberately left in place.
They are now unreferenced, but the class was purpose-built for this one line
and keeping it makes restoring the byline a one-line change. If the removal is
permanent, that block (and its 640px media query) should go with it.

Also updates two quoted figures in the 2026-07-02 UX audit notes from
"24,000+" to "27,000+".

Worth noting for the record: that file documents what the page said when it was
audited, and at that time it genuinely did say "24,000+" — which was itself the
bug later fixed by reading unique_schools instead of a field the API never
sent. Editing the quoted evidence makes the note read consistently with the
current site, at the cost of no longer being a verbatim record of what was
observed. Left as the owner edited it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 10:21:29 +01:00
tudor 38acc76555 Merge pull request 'fix(charts): make the national-average marker visible on both templates' (#99) from fix/chart-marker-contrast into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 16s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m3s
Reviewed-on: #99
2026-08-15 08:56:40 +00:00
tudor c7e0c2eac7 Merge pull request 'feat(home): swap in the higher-fidelity hero artwork' (#98) from feat/hero-artwork-v2 into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 59s
Reviewed-on: #98
2026-08-15 08:56:33 +00:00
TudorandClaude Opus 5 59ac9c10b9 fix(charts): make the national-average marker visible on both templates
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m1s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m5s
The marker was var(--brand) — the identical value to the bar fill it sits on.
Measured by pixel-sampling every bar on both templates, in both themes:

  primary   SatsChart .natTick        1.00:1  (6 bars)
  secondary .att8VizNatLine           1.00:1

Not low contrast. The same colour. It scored 5.47:1 only against the empty
track, which means it was visible precisely when a school was BELOW the
national average and vanished for every school at or above it — failing for
exactly the schools people are looking for.

No single colour fixes this, because the marker's position is data-driven: it
can land on the bar, on the empty track, or across the boundary. It is now a
knockout — a light core carrying a dark edge, both from tokens that flip with
the theme, so one part or the other always separates:

              on the bar        on the track
  light       5.47 / 9.70:1     15.17:1
  dark        7.81 / 10.85:1    13.52:1

THE LEGEND DESCRIBED A CHART THAT DID NOT EXIST

Both data swatches were var(--status-above) green while their bars were
var(--brand) teal, and the two were identical to each other — one swatch for
two series. Worse, the only swatch matching the bar colour was the one
labelled "National average", so reading the chart by matching colours told you
the teal bars were the benchmark. Each swatch now carries its bar's exact
value, and the marker swatch mirrors the knockout.

TWO SERIES, ONE COLOUR

Expected and Exceeding were both var(--brand), distinguished only by row.
They are a sequential pair — exceeding is the same cohort at a harder bar — so
they take two steps of one hue, the harder measure being the step further from
the ground in each theme.

They measure 1.77:1 (light) and 1.39:1 (dark) against each other, and that is
accepted rather than overlooked: two fills that must EACH clear 3:1 against
the same white track are geometrically forced close together. The distinction
is carried by the row labels and printed values; colour is redundant here, not
load-bearing.

THE TEST THAT SHOULD HAVE CAUGHT THIS

The WCAG journey composites backgrounds by walking the ancestor chain, but
these markers are absolutely positioned over a sibling — ancestor-walking is
structurally blind to overlap, which is why this shipped in two templates and
passed every gate. The new journey asks the stacking order instead, via
document.elementsFromPoint.

Writing it surfaced a second trap worth recording: elementsFromPoint takes
viewport coordinates and returns an empty stack off-screen, and these markers
sit ~1200px down. The first version defaulted an empty stack to white, which
made teal-on-teal look like teal-on-white and PASS in the light theme. It now
scrolls each marker into view and counts any marker it cannot resolve a
backdrop for as a failure rather than a pass.

Verified both directions: the new journey fails against current staging with
"1.00:1 over rgb(15,118,110)" in both themes, and passes against this build
with 6/6 markers measured at 5.47–15.17:1.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 09:48:41 +01:00
TudorandClaude Opus 5 eddf74745f feat(home): swap in the higher-fidelity hero artwork
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 43s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 13s
Same 16:9 frame, but not a drop-in replacement — the composition, palette and
copy area all moved, and each one drives a change here.

COPY AREA

The old artwork was a flat #FDF9F3 down its whole left side. This one runs
peach #FEE8D2 at the top to cream #FDF3E7 around 43% height, and below ~48%
the left edge is foliage rather than cream at all. --hero-ground is resampled
to #FEF2E1, the middle of the pale run where the headline and search actually
sit; the scrim covers the foliage further down.

Measured on rendered pixels with the text hidden, sampling the real glyph runs
(via Range, not the element boxes) plus a 120px growth margin:

           title      body      byline
  light    13.61:1    5.74:1    6.66:1
  dark      9.49:1    5.45:1    8.35:1

Body drops from 6.52:1 to 5.74:1 as the foliage comes closer, still clear.

BAND CROP

The schoolhouse moved to ~82% across the frame, leaving only ~140px of artwork
to its right, so the band crop is anchored to the right edge and takes 1344px
back — putting the school at 73%, which the build script now derives and
prints rather than leaving it to drift from the CSS.

The crop is also 3.2:1 rather than matching the phone. The band is not one
ratio: it runs 2.4:1 on a phone to about 3.8:1 under the one-column
breakpoint. Cropping at the narrow end means the wide end throws away height,
which cut the flag off the roof and the base off the building. Sitting above
the middle costs a little width on phones — where the crop's left is hillside —
and keeps the building whole where it matters.

BAND HEIGHTS

Raised to 13rem (861–860px) and 10rem (≤640px). The band was widest-per-height
at exactly 640px, where an 8.5rem band measured 4.2:1 — worse than any wider
viewport, because the height steps down at that breakpoint while the width does
not. Ratios across the range are now 2.0 / 2.6 / 3.6 / 3.1 / 3.8 rather than
spiking. The search still lands at 442px on a 667px viewport.

WEIGHT

Source is 3.1MB; a browser fetches one file — 29–59kB on desktop, 10–14kB on a
phone. Widths re-cut for the larger master: 2000/1400/1000 wide, 1344/900/600
band.

Verified: 0 AA failures across both themes at 1440 / 860 / 390, every srcset
entry present on disk with no orphans, tsc clean, 159/159, build green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 09:22:37 +01:00
tudor d5f3fdf0f9 Merge pull request 'feat(home): add a first-person byline under the hero search' (#97) from feat/hero-byline into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m1s
Reviewed-on: #97
2026-08-15 08:00:47 +00:00
TudorandClaude Opus 5 3015c37bac feat(home): add a first-person byline under the hero search
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 12s
PR Checks / Build Frontend (no push) (pull_request) Successful in 43s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
"Built for parents, by a parent." — the one line on the page written by a
person rather than by a product.

Placed under the search rather than in the footer, which is where it would
have been true and unread. It does its work at the moment someone is deciding
whether to trust a page full of government statistics, which is the moment
they are looking at the search box.

Set apart without shouting: display face, a step down in size, and a short
brand rule in place of a bullet. No italic — Manrope ships none in the loaded
weights, so font-style would be synthesised into a slant, the same reason
.heroEmph resets it.

Deliberately the only claim of its kind on the page. Every other trust signal
here is about the data's provenance and is checkable against the DfE and
Ofsted; this one is about who built it, and is not. So it is stated once,
plainly, and never repeated — the opposite of the "Official & trusted" line it
now sits near, which claimed something about the site that was not true.

Measured 7.26:1 in the light theme and 8.86:1 in dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-15 08:59:43 +01:00
tudor 153b26a32f Merge pull request 'fix(home): stop implying we are official, and fix the mobile hero and search' (#96) from fix/hero-mobile-and-wording into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 56s
Reviewed-on: #96
2026-08-14 22:28:33 +00:00
tudor e373241d51 Merge pull request 'fix(e2e): assert the brand lockup and touch icon as they are actually built' (#95) from fix/e2e-brand-assertions into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 11s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m2s
Reviewed-on: #95
2026-08-14 22:28:10 +00:00
TudorandClaude Opus 5 4043270a77 fix(home): stop implying we are official, and fix the mobile hero and search
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 48s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m6s
Three things, all reported from the live mobile view.

WE ARE NOT AN OFFICIAL SERVICE

The value prop read "Official & trusted", whose grammatical subject is this
site — it reads as a claim that schoolcompare is an official service. It is
not; it is an independent site that republishes official figures. Now "Built
on official data", which describes the data instead.

Checked every other use of the word: all nine describe the data ("official DfE
figures", "the official figure isn't in") and are correct as they stand. That
one title was the only place the site described itself.

The footer now says it outright — "An independent site. Not affiliated with
the Department for Education or Ofsted." — so it appears on every page rather
than being left to inference. An e2e test asserts both halves: that the
statement is present, and that no text presents the site itself as official.

The disclaimer first measured 4.71:1 against a 4.5 floor. That is a fine
margin for decoration and the wrong one for a line whose job is to be legible
to someone checking whether this is a government site; it is now 5.59:1,
quieter than the copy around it by size rather than by contrast.

THE ARTWORK SAT UNDER THE SEARCH ON PHONES

The DOM keeps .heroContent first so the desktop overlay does not depend on
source order, which left the band stranded at the bottom of the panel, reading
as a strip stuck on the end rather than a hero image. `order` moves it above
the copy on phones; it is decorative and aria-hidden, so no reading order
changes. Measured on a 667px viewport — the shortest phone still in use — the
search button lands at 418px, comfortably inside the fold.

THE SEARCH INPUT WAS UNUSABLE ON PHONES

"Search schools" is a fixed 134px with white-space: nowrap, so it took 48% of
the row. Measured on staging:

  390px   input text space 102px   placeholder needs 194px
  360px                     72px
  320px                     32px

At 320px you could not see what you were typing. The button now wraps to its
own full-width row below 480px, which fixes 360px and up — 390px goes from
102px to 226px of text space. The pill stays a single element, so it keeps its
border, shadow and :focus-within ring, and the button gains a full-width tap
target.

320px is still ~30px short. Closing it needs a shorter placeholder, which three
other tests match on by exact string; left alone deliberately rather than
churn them for a width that is effectively gone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 22:59:12 +01:00
TudorandClaude Opus 5 e65688d600 fix(e2e): assert the brand lockup and touch icon as they are actually built
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 48s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 21s
The staging gate has been failing these two since the supplied logo artwork
replaced the reconstruction. Both tests were still asserting the previous
implementation, and both were right to fail — they were just describing
something the site no longer does.

  the header carries the schoolcompare lockup
    Looked for an <svg> inside the header link. The mark is raster now:
    <picture><source><img src="/brand/mark.png">. Nothing matched, so the
    locator timed out.

  the brand asset set is complete and served
    Requested /apple-icon, which 404s. The route moved when the generated
    app/apple-icon.tsx became a static app/apple-icon.png — generated icons
    serve at /apple-icon, static ones at /apple-icon.png with a content hash.
    The icon was present and correctly linked the whole time.

Both now read from the page instead of hardcoding the shape of the answer: the
touch icon is fetched from its own <link rel="apple-touch-icon"> href, the way
this test already handles og:image, so it follows whatever Next emits.

The lockup assertion also got stronger rather than merely corrected. A
<picture> whose sources all 404 still lays out and still satisfies
toBeVisible(), so that alone would go green on a broken lockup; it now asserts
naturalWidth, which only a decoded image can satisfy.

Verified against staging directly: 41/41 pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 22:34:22 +01:00
tudor 7f4dfa2748 Merge pull request 'fix(home): art-direct the hero's fallback path, and declare sharp' (#94) from fix/hero-fallback-and-sharp into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m12s
Reviewed-on: #94
2026-08-14 21:28:13 +00:00
TudorandClaude Opus 5 bdaa05cd54 style(home): lift the hero artwork's dark-theme brightness to 0.75
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 48s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
At 0.52 the scene was legible but heavily suppressed; 0.75 lets the hills,
path and schoolhouse read while the white H1 stays comfortably the brightest
thing on the panel.

Re-measured rather than assumed, because in the dark theme the text is light
and the artwork is behind it — brightening the image lowers text contrast
rather than raising it. Off rendered pixels, sampling background up to 120px
past each line's right edge:

  brightness   title      body
  0.52         11.01:1    6.44:1
  0.75          9.48:1    5.21:1

Both still clear the 4.5:1 floor, body being the binding one. The trade is
recorded next to the value so the next person to reach for it knows it has a
floor and not just a taste range.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 22:26:23 +01:00
TudorandClaude Opus 5 043506cb6b fix(home): art-direct the hero's fallback path, and declare sharp
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m32s
Two findings from review on #93, both verified before fixing.

<img src> cannot vary by viewport, so it was always the wide desktop crop. A
browser taking neither AVIF nor WebP therefore fell through to the desktop
frame on a phone and lost the schoolhouse — the exact failure the two-crop
<picture> exists to prevent, surviving in the one path nobody looks at. The
band JPEG the build script already emitted was never referenced, which was the
tell. It now backs a <source media> placed after the modern formats, so they
still win wherever they are supported.

Verified by stripping the AVIF and WebP <source>s at runtime and letting
<picture> re-resolve, which is what an old browser actually sees:

  phone    hero-band-500.avif  →  hero-band-700.jpg   (band crop, school kept)
  desktop  hero-wide-1672.avif →  hero-wide-1200.jpg

sharp was not declared: it arrives transitively from next@16.1.6, so the
documented regeneration command works today and breaks on a Next upgrade or a
clean install that resolves differently. Declared in devDependencies for the
same reason next.config.js already declares its traced font files rather than
trusting the tracer to keep finding them.

The third finding — that the hero's licence is marked unconfirmed in
CREDITS.md while the artwork ships — is accurate and deliberate. It is the
owner's to answer; recording it as unknown is the point of the file.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 22:07:01 +01:00
tudor 0466986195 Merge pull request 'feat(home): use the supplied hero artwork instead of a drawn one' (#93) from feat/hero-artwork into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 53s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m14s
Reviewed-on: #93
2026-08-14 20:41:46 +00:00
TudorandClaude Opus 5 3f0e05cc99 feat(home): use the supplied hero artwork instead of a drawn one
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m29s
Replaces the hand-drawn SVG landscape with the illustration supplied by the
project owner. Master is assets/hero-source.png; everything under
public/brand/hero-* comes from scripts/build-hero-images.js and should never be
hand-edited.

LAYOUT

The artwork is composed as a full-bleed hero: it reserves an empty cream area
down its left side for the headline. The old hero was a two-column grid with
art in the right 0.88fr, which would have cropped that reserved area off and
shrunk the scene into a thumbnail — the one thing the picture is built not to
be. So the panel is now a single layered block: artwork behind, copy on top,
and below the one-column breakpoint the artwork leaves the background and
becomes a band under the search.

Overlaying does not survive down to phone widths — the panel gets too narrow
for the copy to stay inside the cream, so the scrim would have to cover nearly
the whole image and you would be left with a tinted rectangle. Hence the
switch at 860px rather than a single treatment stretched across every width.

CONTRAST

The copy sits on a gradient of --hero-ground, a new token sampled from the
artwork's own cream (#FDF9F3) rather than from Sand. Sand is eight to thirteen
points darker per channel, which leaves a visible seam straight down the hero.
The scrim exists because the artwork is a fixed image on a fluid panel: past
some width the headline would otherwise land on hillside green.

Measured on rendered pixels, sampling background up to 120px beyond the right
edge of each line, so a longer line still has margin:

  light  title 14.36:1   body 6.52:1
  dark   title 11.01:1   body 6.44:1

The worst light case is hillside green showing through the scrim at 6.52:1.

TWO CROPS

The slot is two shapes: roughly 2.1:1–2.7:1 behind the desktop panel, and
2.6:1–4.9:1 as the band. A single file under object-fit: cover centre-crops,
and at the band's extreme that slices a strip through the scene and loses the
schoolhouse — exactly how the drawn hero failed on phones. <picture> switches
crop, not just resolution: the band file is pre-cropped around the school and
is already near 2.6:1, with the subject at ~65% across so squeezing toward
4.9:1 crops the empty sides instead.

WEIGHT

1.6 MB PNG in, AVIF and WebP out; a browser fetches exactly one file — about
23–33 kB on desktop, 9–18 kB on a phone. The JPEG is only the <img> fallback.
`sizes` describes the real panel box (max-width 1400 minus padding) rather than
100vw, which over-requested a candidate at every width on the LCP element.

DARK THEME

A raster cannot be re-graded token by token the way the drawing was, but it
would reproduce the same failure — a bright illustration is the brightest
object on a near-black page. It is dimmed in CSS to read as dusk, and the
scrim fades it into the dark panel rather than into cream.

TESTS

The two illustration journeys are replaced by three, each covering a failure
that still renders a valid-looking page: the artwork actually loading
(naturalWidth, not src), the crop switching at the breakpoint, a modern format
winning negotiation, and the dark theme dimming it.

public/brand/CREDITS.md records provenance for every asset in that folder. The
hero's licence line is marked unconfirmed — that one needs the owner.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 21:36:40 +01:00
tudor f1819e9c4d Merge pull request 'fix(home): correct what the landing page claims, and give it one rhythm' (#92) from fix/homepage-truth-and-rhythm-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m13s
Reviewed-on: #92
2026-08-14 18:18:53 +00:00
TudorandClaude Opus 5 97ac5c9cef fix(a11y): clear the four AA failures blocking the staging E2E gate
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 49s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 5m29s
Main's staging E2E has been red since #90. Two of the four failures were the
staging container serving a stale image at the moment the gate ran — apple-icon
and the header lockup both check out against staging now. The other two were
real, and both are the same mistake: a colour pairing verified against one
ground while the element sits on another.

  * .heroEyebrow  --brand on its own 10% tint over Sand is 4.19:1. The token
    clears AA on Sand, but the tint darkens the ground under the label, and
    that composite was never the thing measured. --brand-strong is 5.85:1.

  * .hiwVisual    The preview panel was grounded in Sand while everything
    inside it is a translucent status tint — and those tints are specced to
    clear AA over --bg-card. Over Sand the report-card chips landed at 4.22:1
    and 4.29:1. The ratio depended on a background two levels up. Moving the
    panel to the card surface puts them at 5.69:1 and 5.81:1; a border keeps
    it reading as an inset frame now that panel and card share a colour.

  * Footer .sectionTitle  --sage flips with the theme; the footer band does
    not (it is teal in both). So the pairing only held in one of them — the
    dark sage measured 4.08:1, on every page. Now --on-sunken-muted, which is
    what the --on-sunken-* family exists for, as Footer.tsx's own header
    comment already says.

  * .compareHeadLabel  --text-muted on the header row's tint is 4.45:1 in the
    dark theme. Under by a hair, same cause. Now --text-secondary.

The harness gained the footer, because it had no footer and so could not have
caught the one failure that appeared on all three pages. Rewriting fragments
now uses each CSS module's own hash — the footer rewritten with HomeView's
prefix renders unstyled, which would have made any contrast measured on it
meaningless while still reporting a number.

Verified: 0 AA failures in both themes across the assembled page, measured with
the same probe the e2e test uses, against the real compiled CSS.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 18:08:38 +01:00
TudorandClaude Opus 5 dc22fd2853 fix(home): correct what the landing page claims, and give it one rhythm
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m25s
The homepage made four statements that were not true, carried elements that
asked nothing of anyone, and had a hero illustration that broke in both the
places it had to work.

Claims, all verified against the code or the API:

  * "24,000+ schools" (three places) against a real 27,230. The fact box meant
    to show the live figure rendered its own fallback on every request, because
    DataInfoResponse declared a `total_schools` field the API has never sent —
    it sends `unique_schools`. The fetch succeeded; only that field was
    undefined, so nothing threw and nothing failed. The interface, not the
    code, was the thing that was wrong.
  * "Up to three schools side by side" against MAX_SCHOOLS = 5, contradicting a
    card 400px below it that correctly said five.
  * "Class sizes" — data the codebase has never held. That copy line was the
    only hit in a full-repo grep.
  * Invented results and Ofsted grades attributed to two real, named schools
    in the compare preview.

Also one feature, three words: Compare (nav), shortlist (footer), pin (cards).
Settled on Compare everywhere. And <title> was the bare string "Home".

Cut: the trust line (repeated the coverage figure one paragraph after the hero
gave it, behind three decorative dots), the "Start exploring" row (three links
to two destinations already in the nav), the six-row coverage table, and three
of the four countdown cards — which gave the page's largest numeral to dates up
to 245 days away, two of them offer days, which cannot be missed. All four
dates remain, at proportionate weight. Value-prop titles drop from <h2> to <p>;
they were outranking the page's real headings in the document outline.

Rhythm: the gaps between the seven landing bands were 24/32/24/16/48/32/16px,
each band setting its own margin, with four different section-header
treatments between them. The page container now owns one gap, and there is one
header pattern. An e2e test asserts the gaps are identical.

Illustration: it kept a fixed light palette in both themes, which left a pale
sky slab as the brightest object on a near-black page, out-shouting the H1 and
the search box. It now reads from --ill-* tokens with a dark re-grade. And the
hero slot ranges from 1.34:1 to 4.9:1 across breakpoints, which no single
composition survives under `slice` — at 860x176 a 540x520 scene shows only its
bottom 110 units, so the schoolhouse was cropped away entirely on phones,
leaving hills and a pin pointing at nothing. There are now two compositions,
each drawn against the crop window its own breakpoint produces, with CSS
showing one. Both are static server-rendered SVG.

The deadline bar renders on the server rather than on hydrate. The effect-based
version needed a reserved height, and one guessed number cannot cover a block
whose supporting line wraps differently at every width — measured, it was short
at all four, shifting the page up to 108px on a phone.

Verified on the built output through an offline render harness (no local
server): real compiled CSS, real rendered markup, four widths, both themes.
tsc clean, 159/159 unit tests, build green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 17:29:47 +01:00
tudor 6b117dd26c Merge pull request 'feat(brand): use the supplied logo artwork instead of a reconstruction' (#90) from feat/brand-logo-artwork into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 42s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 53s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m0s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m42s
Reviewed-on: #90
2026-08-08 21:41:47 +00:00
299 changed files with 44198 additions and 3226 deletions

No files matched your search

+13 -5
View File
@@ -20,7 +20,7 @@ PORT=80
# =============================================================================
# CORS
# =============================================================================
# Comma-separated list of allowed origins
# JSON array of allowed origins (pydantic-settings format)
# In production, only include your actual domain
ALLOWED_ORIGINS=["https://schoolcompare.co.uk"]
@@ -33,13 +33,21 @@ ADMIN_API_KEY=CHANGE_THIS_TO_A_SECURE_RANDOM_KEY
# Rate limiting (requests per minute per IP)
RATE_LIMIT_PER_MINUTE=60
RATE_LIMIT_BURST=10
GLOBAL_RATE_LIMIT_PER_MINUTE=3000
# Maximum request body size in bytes (default 1MB)
MAX_REQUEST_SIZE=1048576
# =============================================================================
# API
# SEARCH AND OPTIONAL FEATURE FLAGS
# =============================================================================
DEFAULT_PAGE_SIZE=50
MAX_PAGE_SIZE=100
TYPESENSE_URL=http://localhost:8108
TYPESENSE_API_KEY=CHANGE_THIS_TO_YOUR_TYPESENSE_KEY
# Empty URL disables Unleash-backed flags. Match the managed environment when used.
UNLEASH_URL=
UNLEASH_API_TOKEN=
# Page-size limits are currently declared by route Query parameters.
# DEFAULT_PAGE_SIZE, MAX_PAGE_SIZE and RATE_LIMIT_BURST are not reliable tuning
# controls in the current routes; see docs/LEGACY_CODE.md.
+73 -17
View File
@@ -5,6 +5,11 @@ on:
branches:
- main
# Serialise the entire build/deploy/test cycle: no other run can move staging tags.
concurrency:
group: staging-release
cancel-in-progress: false
env:
REGISTRY: privaterepo.sitaru.org
BACKEND_IMAGE_NAME: ${{ gitea.repository }}-backend
@@ -12,7 +17,18 @@ env:
PIPELINE_IMAGE_NAME: ${{ gitea.repository }}-pipeline
jobs:
prepare:
runs-on: ubuntu-latest
outputs:
build_id: ${{ steps.identity.outputs.build_id }}
steps:
- id: identity
run: python3 -c 'import uuid; print("build_id=" + uuid.uuid4().hex)' >> "$GITHUB_OUTPUT"
build-backend:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Backend (FastAPI)
runs-on: ubuntu-latest
steps:
@@ -46,17 +62,24 @@ jobs:
type=raw,value=staging
- name: Build and push Backend Docker image
id: build
uses: docker/build-push-action@v5
with:
context: .
file: ./Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-backend.outputs.tags }}
labels: ${{ steps.meta-backend.outputs.labels }}
cache-from: type=registry,ref=${{ env.REGISTRY }}/${{ env.BACKEND_IMAGE_NAME }}:buildcache
cache-to: type=registry,ref=${{ env.REGISTRY }}/${{ env.BACKEND_IMAGE_NAME }}:buildcache,mode=max
build-frontend:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Frontend (Next.js)
runs-on: ubuntu-latest
steps:
@@ -90,18 +113,23 @@ jobs:
type=raw,value=staging
- name: Build and push Frontend Docker image
id: build
uses: docker/build-push-action@v5
with:
context: ./nextjs-app
file: ./nextjs-app/Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-frontend.outputs.tags }}
labels: ${{ steps.meta-frontend.outputs.labels }}
build-args: |
FASTAPI_URL=http://backend:80/api
# Cache disabled due to registry size limits
build-pipeline:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Pipeline (Meltano + dbt + Airflow)
runs-on: ubuntu-latest
steps:
@@ -135,11 +163,15 @@ jobs:
type=raw,value=staging
- name: Build and push Pipeline Docker image
id: build
uses: docker/build-push-action@v5
with:
context: ./pipeline
file: ./pipeline/Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-pipeline.outputs.tags }}
labels: ${{ steps.meta-pipeline.outputs.labels }}
cache-from: type=registry,ref=${{ env.REGISTRY }}/${{ env.PIPELINE_IMAGE_NAME }}:buildcache
@@ -148,30 +180,23 @@ jobs:
deploy-staging:
name: Deploy to Staging
runs-on: ubuntu-latest
needs: [build-backend, build-frontend, build-pipeline]
needs: [prepare, build-backend, build-frontend, build-pipeline]
steps:
- name: Trigger staging stack update
run: curl -fsSk -X POST "${{ secrets.PORTAINER_STAGING_WEBHOOK }}"
- name: Wait for staging to become healthy
run: |
echo "Polling ${STAGING_BASE_URL} for up to 5 minutes..."
for i in $(seq 1 60); do
if curl -fsS -o /dev/null --max-time 10 "${STAGING_BASE_URL}/"; then
echo "Staging is up (attempt $i)"
exit 0
fi
sleep 5
done
echo "Staging did not become healthy in time" >&2
exit 1
- uses: actions/checkout@v4
- name: Verify deployed release identity
run: python3 scripts/ci/release.py wait
env:
STAGING_BASE_URL: ${{ secrets.STAGING_BASE_URL }}
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
e2e-staging:
name: E2E Journeys against Staging
runs-on: ubuntu-latest
needs: [deploy-staging]
needs: [prepare, deploy-staging, build-backend, build-frontend, build-pipeline]
steps:
- name: Checkout repository
uses: actions/checkout@v4
@@ -187,11 +212,42 @@ jobs:
npm ci
npx playwright install --with-deps chromium
- name: Verify release before journeys
run: python3 scripts/ci/release.py wait --timeout 10
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- name: Run E2E journeys
working-directory: e2e
run: npx playwright test
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- name: Verify release after journeys
run: python3 scripts/ci/release.py wait --timeout 10
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- uses: docker/setup-buildx-action@v3
- uses: docker/login-action@v3
with:
registry: ${{ env.REGISTRY }}
username: ${{ gitea.actor }}
password: ${{ secrets.REGISTRY_TOKEN }}
- name: Mark tested image digests as verified
run: python3 scripts/ci/release.py verify
env:
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
BACKEND_DIGEST: ${{ needs.build-backend.outputs.digest }}
FRONTEND_DIGEST: ${{ needs.build-frontend.outputs.digest }}
PIPELINE_DIGEST: ${{ needs.build-pipeline.outputs.digest }}
# Production deployment is a second, manual approval: see promote.yml
# ("Promote to Production (manual)") and docs/DEPLOY.md.
+2 -2
View File
@@ -68,13 +68,13 @@ jobs:
python-version: "3.12"
- name: Install dependencies
run: pip install -r requirements.txt pytest "httpx<0.28"
run: pip install -r requirements.txt pytest "httpx<0.28" pyyaml
- name: Import smoke test
run: python -c "from backend.app import app; print('backend imports OK')"
- name: Backend unit tests
run: python -m pytest backend/tests -q
run: python -m pytest backend/tests pipeline/tests scripts/ci/tests -q
build-backend:
name: Build Backend (no push)
+7 -25
View File
@@ -97,33 +97,15 @@ jobs:
username: ${{ gitea.actor }}
password: ${{ secrets.REGISTRY_TOKEN }}
- name: Retag approved images as prod (keeping rollback pointer)
run: |
SHORT_SHA="${{ steps.resolve.outputs.short }}"
for IMAGE in \
"${REGISTRY}/${BACKEND_IMAGE_NAME}" \
"${REGISTRY}/${FRONTEND_IMAGE_NAME}" \
"${REGISTRY}/${PIPELINE_IMAGE_NAME}"; do
# Keep a rollback pointer before moving :prod
docker buildx imagetools create -t "${IMAGE}:prod-previous" "${IMAGE}:prod" || true
docker buildx imagetools create -t "${IMAGE}:prod" "${IMAGE}:${SHORT_SHA}"
echo "Promoted ${IMAGE}:${SHORT_SHA} -> :prod"
done
- name: Resolve verified digests and promote the complete image set
run: python3 scripts/ci/release.py promote --output release.json
env:
EXPECTED_SHA: ${{ steps.resolve.outputs.full }}
- name: Trigger production stack update
run: curl -fsSk -X POST "${{ secrets.PORTAINER_PROD_WEBHOOK }}"
- name: Wait for production to become healthy
run: |
echo "Polling ${PROD_BASE_URL} for up to 5 minutes..."
for i in $(seq 1 60); do
if curl -fsS -o /dev/null --max-time 10 "${PROD_BASE_URL}/"; then
echo "Production is up (attempt $i)"
exit 0
fi
sleep 5
done
echo "Production did not become healthy in time" >&2
exit 1
- name: Verify production release identity
run: python3 scripts/ci/release.py wait --release release.json
env:
PROD_BASE_URL: ${{ secrets.PROD_BASE_URL }}
BASE_URL: ${{ secrets.PROD_BASE_URL }}
+6
View File
@@ -5,3 +5,9 @@ __pycache__/
pipeline/transform/target/
pipeline/transform/logs/
pipeline/transform/.user.yml
# Playwright MCP scratch output (screenshots, console logs, page snapshots)
.playwright-mcp/
# Playwright run artefacts written when the suite is run from the repo root
test-results/
+13 -187
View File
@@ -1,191 +1,17 @@
# Docker Deployment Guide
# Docker deployment
## Quick Start
The maintained deployment runbook is [docs/DEPLOY.md](docs/DEPLOY.md).
Deploy the complete SchoolCompare stack (PostgreSQL + FastAPI + Next.js) with one command:
- Production: `docker-compose.portainer.yml`, using `:prod` images.
- Staging: `docker-compose.portainer.staging.yml`, using `:staging` images.
- Builds and deployment: `.gitea/workflows/deploy.yml`.
- Human-approved production promotion: `.gitea/workflows/promote.yml`.
```bash
docker-compose up -d
```
The generic `docker-compose.yml` is not a supported one-command onboarding path:
it still uses `:latest` tags that the release workflow no longer publishes and
lacks the full current CMS setup. Review the [legacy inventory](docs/LEGACY_CODE.md)
before using old compose examples. Starting an empty database does not populate
school marts.
This will start:
- **PostgreSQL** on port 5432 (database)
- **FastAPI** on port 8000 (backend API)
- **Next.js** on port 3000 (frontend)
## Service Details
### PostgreSQL Database
- **Port**: 5432
- **Container**: `schoolcompare_db`
- **Credentials**:
- User: `schoolcompare`
- Password: `schoolcompare`
- Database: `schoolcompare`
- **Volume**: `postgres_data` (persistent storage)
### FastAPI Backend
- **Port**: 8000 → 80 (container)
- **Container**: `schoolcompare_backend`
- **Built from**: Root `Dockerfile`
- **API Endpoint**: http://localhost:8000/api
- **Health Check**: http://localhost:8000/api/data-info
### Next.js Frontend
- **Port**: 3000
- **Container**: `schoolcompare_nextjs`
- **Built from**: `nextjs-app/Dockerfile`
- **URL**: http://localhost:3000
- **Connects to**: Backend via internal network
## Commands
### Start all services
```bash
docker-compose up -d
```
### View logs
```bash
# All services
docker-compose logs -f
# Specific service
docker-compose logs -f nextjs
docker-compose logs -f backend
docker-compose logs -f db
```
### Check status
```bash
docker-compose ps
```
### Stop all services
```bash
docker-compose down
```
### Rebuild after code changes
```bash
# Rebuild and restart specific service
docker-compose up -d --build nextjs
# Rebuild all services
docker-compose up -d --build
```
### Clean restart (remove volumes)
```bash
docker-compose down -v
docker-compose up -d
```
## Initial Database Setup
After first start, you may need to initialize the database:
```bash
# Enter the backend container
docker exec -it schoolcompare_backend bash
# Run migrations or data loading
python -m backend.data_loader
```
## Accessing Services
Once running:
- **Frontend**: http://localhost:3000
- **Backend API**: http://localhost:8000/api
- **API Docs**: http://localhost:8000/docs (Swagger UI)
- **Database**: localhost:5432 (use any PostgreSQL client)
## Environment Variables
Create a `.env` file in the root directory to customize:
```env
# Database
POSTGRES_USER=schoolcompare
POSTGRES_PASSWORD=your_secure_password
POSTGRES_DB=schoolcompare
# Backend
DATABASE_URL=postgresql://schoolcompare:your_secure_password@db:5432/schoolcompare
# Frontend (for client-side access)
NEXT_PUBLIC_API_URL=http://localhost:8000/api
```
Then run:
```bash
docker-compose up -d
```
## Troubleshooting
### Backend not connecting to database
```bash
# Check database health
docker-compose ps
# View backend logs
docker-compose logs backend
# Restart backend
docker-compose restart backend
```
### Frontend not connecting to backend
```bash
# Check backend health
curl http://localhost:8000/api/data-info
# Check Next.js environment variables
docker exec schoolcompare_nextjs env | grep API
```
### Port already in use
```bash
# Change ports in docker-compose.yml
# For example, change "3000:3000" to "3001:3000"
```
### Rebuild from scratch
```bash
docker-compose down -v
docker system prune -a
docker-compose up -d --build
```
## Production Deployment
For production, update the following:
1. **Use secure passwords** in `.env` file
2. **Configure reverse proxy** (Nginx) in front of Next.js
3. **Enable HTTPS** with SSL certificates
4. **Set production environment variables**:
```env
NODE_ENV=production
POSTGRES_PASSWORD=<strong-password>
```
5. **Backup database** regularly:
```bash
docker exec schoolcompare_db pg_dump -U schoolcompare schoolcompare > backup.sql
```
## Network Architecture
```
Internet
↓
Next.js (port 3000) ← User browsers
↓ (internal network)
FastAPI (port 8000) ← API calls
↓ (internal network)
PostgreSQL (port 5432) ← Data queries
```
All services communicate via the `schoolcompare-network` Docker network.
For architecture, configuration and test commands, see
[ARCHITECTURE.md](docs/ARCHITECTURE.md) and [DEVELOPMENT.md](docs/DEVELOPMENT.md).
+6
View File
@@ -24,6 +24,12 @@ RUN pip install --no-cache-dir -r requirements.txt
COPY backend/ ./backend/
COPY scripts/ ./scripts/
ARG BUILD_SHA=development
ARG BUILD_ID=development
LABEL io.schoolcompare.build-id=$BUILD_ID
LABEL io.schoolcompare.commit=$BUILD_SHA
RUN python -c 'import json,sys; open("backend/build-info.json", "w").write(json.dumps({"sha":sys.argv[1],"build_id":sys.argv[2]}))' "$BUILD_SHA" "$BUILD_ID"
# Expose the application port
EXPOSE 80
+2
View File
@@ -1,3 +1,5 @@
> Historical migration record, retained for context. Setup and architecture claims below may be obsolete. Use [README.md](README.md), [architecture](docs/ARCHITECTURE.md) and [deployment](docs/DEPLOY.md) for current guidance.
# SchoolCompare: Vanilla JS → Next.js Migration Summary
## Overview
+52 -199
View File
@@ -1,214 +1,67 @@
# Primary School Compass 🧒📚
# SchoolCompare
A modern web application for comparing **primary school (KS2)** performance data in **Wandsworth and Merton** over the last 5 years. Built with FastAPI and vanilla JavaScript with Chart.js visualizations.
SchoolCompare compares schools across England: primary (KS2), secondary (KS4),
all-through and post-16 provision, with coverage depending on the source dataset.
It provides school search, postcode maps, comparisons, rankings, place pages,
Ofsted information, admissions and destination measures. Editorial content lives
in a Payload CMS blog.
![Python](https://img.shields.io/badge/Python-3.9+-blue)
![FastAPI](https://img.shields.io/badge/FastAPI-0.109-green)
![License](https://img.shields.io/badge/License-MIT-yellow)
## Start here
## Features
- [Architecture and data flow](docs/ARCHITECTURE.md)
- [Development and validation](docs/DEVELOPMENT.md)
- [Deployment and promotion](docs/DEPLOY.md)
- [Legacy and unused-code inventory](docs/LEGACY_CODE.md)
- [Frontend conventions](nextjs-app/README.md)
- [CMS publishing](nextjs-app/docs/PUBLISHING.md)
- 📊 **Interactive Charts** - Visualize KS2 performance trends over time
- 🔍 **Smart Search** - Find primary schools by name in Wandsworth & Merton
- ⚖️ **Side-by-Side Comparison** - Compare up to 5 schools simultaneously
- 🏆 **Rankings** - View top-performing primary schools by various KS2 metrics
- 📱 **Responsive Design** - Works beautifully on desktop and mobile
## Repository map
## Key Metrics (KS2)
| Path | Responsibility |
|---|---|
| `backend/` | FastAPI routes, cached school data, read-only SQLAlchemy mappings, feature flags |
| `nextjs-app/` | Next.js App Router, React UI, Payload CMS, frontend tests |
| `pipeline/plugins/extractors/` | Custom Singer taps for GIAS, EES, Ofsted and other datasets |
| `pipeline/transform/` | dbt staging/intermediate models, marts, seeds and data tests |
| `pipeline/dags/` | Airflow extraction, transformation and publication workflows |
| `pipeline/scripts/` | Search indexing, code generation and operational diagnostics |
| `e2e/` | Playwright journeys against a running environment |
| `.gitea/workflows/` | PR checks, staging deployment and manual production promotion |
| `scripts/` | CI review tooling and historical data utilities; see the legacy inventory |
| `docs/superpowers/`, `mockups/` | Design history and prototypes, not application entry points |
The application tracks these Key Stage 2 performance indicators:
## Runtime
| Metric | Description |
|--------|-------------|
| **Reading Progress** | Progress in reading from KS1 to KS2 |
| **Writing Progress** | Progress in writing from KS1 to KS2 |
| **Maths Progress** | Progress in maths from KS1 to KS2 |
| **Reading Expected %** | Percentage meeting expected standard in reading |
| **Writing Expected %** | Percentage meeting expected standard in writing |
| **Maths Expected %** | Percentage meeting expected standard in maths |
| **Reading, Writing & Maths Combined %** | Percentage meeting expected standard in all three subjects |
The public site is **Next.js**, not the FastAPI root page. Browser `/api/*`
requests pass through a Next.js route handler to FastAPI. Server-rendered pages
call FastAPI directly using `FASTAPI_URL`, including its `/api` suffix.
## Quick Start
PostgreSQL/PostGIS stores school data. Meltano/Singer extracts source data;
dbt builds `marts.*`; FastAPI reads those tables. Typesense serves text search
and autocomplete. Payload runs inside Next.js and owns a separate `payload`
database schema and uploaded media.
### 1. Clone and Setup
There is **no automatic CSV import or sample dataset on startup**. A working
school-data environment needs populated marts from the pipeline or an approved
database snapshot. See [development](docs/DEVELOPMENT.md) before choosing a setup.
```bash
cd school_results
## Validation
# Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
```sh
cd nextjs-app
npm ci
npm run typecheck
npm test -- --runInBand
```
### 2. Run the Application
Backend checks, pipeline validation, runtime versions and E2E requirements are
listed in [DEVELOPMENT.md](docs/DEVELOPMENT.md). No `npm run lint` script is
currently defined.
```bash
# Start the server
python -m uvicorn backend.app:app --reload --port 8000
```
Then open http://localhost:8000 in your browser.
The app will run with **sample data** by default, showing **110 primary schools** (66 in Wandsworth, 44 in Merton) with 5 years of KS2 performance data.
### 3. (Optional) Use Real Data
To use real UK school performance data:
1. Visit [Compare School Performance - Download Data](https://www.compare-school-performance.service.gov.uk/download-data)
2. Download **Key Stage 2** data for the years you want (2019-2024)
- Select "Key Stage 2" as the data type
3. Place the CSV files in the `data/` folder
4. Restart the server - it will automatically load and filter to Wandsworth & Merton schools
**Note:** The app only displays schools in Wandsworth and Merton. Data from other areas will be filtered out.
See the helper script for more details:
```bash
python scripts/download_data.py
```
## Project Structure
```
school_results/
├── backend/
│ └── app.py # FastAPI application with all API endpoints
├── frontend/
│ ├── index.html # Main HTML page
│ ├── styles.css # Styling (warm, editorial design)
│ └── app.js # Frontend JavaScript
├── data/
│ └── .gitkeep # Place CSV data files here
├── scripts/
│ └── download_data.py # Helper for downloading/processing data
├── requirements.txt # Python dependencies
└── README.md
```
## API Endpoints
| Endpoint | Description |
|----------|-------------|
| `GET /api/schools` | List schools with optional search/filter |
| `GET /api/schools/{urn}` | Get detailed data for a specific school |
| `GET /api/compare?urns=...` | Compare multiple schools |
| `GET /api/rankings` | Get school rankings by metric |
| `GET /api/filters` | Get available filter options |
| `GET /api/metrics` | Get available performance metrics |
### Example API Usage
```bash
# Search for schools
curl "http://localhost:8000/api/schools?search=academy"
# Get school details
curl "http://localhost:8000/api/schools/100001"
# Compare schools
curl "http://localhost:8000/api/compare?urns=100001,100002,100003"
# Get rankings
curl "http://localhost:8000/api/rankings?metric=rwm_expected_pct&year=2024"
```
## Data Format
If using your own CSV data, ensure it includes these columns (or similar):
| Column | Type | Description |
|--------|------|-------------|
| URN | Integer | Unique Reference Number |
| SCHNAME | String | School name |
| LA | String | Local Authority (must be Wandsworth or Merton) |
| READPROG | Float | Reading progress score |
| WRITPROG | Float | Writing progress score |
| MATPROG | Float | Maths progress score |
| PTRWM_EXP | Float | % meeting expected standard in reading, writing & maths |
| PTREAD_EXP | Float | % meeting expected standard in reading |
| PTWRIT_EXP | Float | % meeting expected standard in writing |
| PTMAT_EXP | Float | % meeting expected standard in maths |
The application normalizes column names automatically and filters to only show Wandsworth and Merton schools.
## Technology Stack
- **Backend**: FastAPI (Python) - High-performance async API framework
- **Frontend**: Vanilla JavaScript with Chart.js
- **Styling**: Custom CSS with CSS variables for theming
- **Data**: Pandas for CSV processing
## Design Philosophy
The UI features a warm, editorial design inspired by quality publications:
- **Typography**: DM Sans for body text, Playfair Display for headings
- **Color Palette**: Warm cream background with coral and teal accents
- **Interactions**: Smooth animations and hover effects
- **Charts**: Clean, readable data visualizations
## Development
```bash
# Run with auto-reload
python -m uvicorn backend.app:app --reload --port 8000
# Or run directly
python backend/app.py
```
## Coverage
This application is specifically designed for:
- **School Phase**: Primary schools only (Key Stage 2)
- **Geographic Area**: Wandsworth and Merton (London boroughs)
- **Time Period**: Last 5 years of data (2020-2024)
Note: 2021 data shows as unavailable because SATs were cancelled due to COVID-19.
## Data Source
Data is sourced from the UK Government's [Compare School Performance](https://www.compare-school-performance.service.gov.uk/) service, which provides official school performance data for England.
**Important**: When using real data, please comply with the [terms of use](https://www.compare-school-performance.service.gov.uk/download-data) and data protection regulations.
## Scheduled Jobs
### Geocoding Schools (Cron Job)
School postcodes are geocoded by a scheduled job, not on-demand. This improves performance and reduces API calls.
**Setup the cron job** (runs weekly on Sunday at 2am):
```bash
# Edit crontab
crontab -e
# Add this line (adjust paths as needed):
0 2 * * 0 cd /path/to/school_compare && /path/to/venv/bin/python scripts/geocode_schools.py >> /var/log/geocode_schools.log 2>&1
```
**Manual run:**
```bash
# Geocode only schools missing coordinates
python scripts/geocode_schools.py
# Force re-geocode all schools
python scripts/geocode_schools.py --force
```
## License
MIT License - feel free to use this project for educational purposes.
---
Built with ❤️ for Wandsworth & Merton families
## Deployment
Work on a feature branch and open a PR. Merging to `main` builds images and
deploys staging. Production promotion is a separate, human-triggered Gitea
workflow. Use [DEPLOY.md](docs/DEPLOY.md) and the Portainer compose files as the
operational references. The generic compose examples still reference `:latest`,
which the current release workflow does not publish.
+651 -81
View File
@@ -6,7 +6,9 @@ Uses real data from UK Government Compare School Performance downloads.
import hashlib
import re
import time
from contextlib import asynccontextmanager
from datetime import datetime, timezone
from typing import Optional
import numpy as np
@@ -14,7 +16,7 @@ import pandas as pd
from fastapi import FastAPI, HTTPException, Query, Request, Depends, Header
from fastapi.middleware.cors import CORSMiddleware
from fastapi.middleware.gzip import GZipMiddleware
from fastapi.responses import FileResponse, Response
from fastapi.responses import FileResponse, JSONResponse, Response
from fastapi.staticfiles import StaticFiles
from slowapi import Limiter, _rate_limit_exceeded_handler
from slowapi.util import get_remote_address
@@ -24,7 +26,8 @@ from starlette.middleware.base import BaseHTTPMiddleware
import asyncio
from .config import settings
from .data_loader import (
clear_cache,
build_latest_school_data,
load_school_data_as_dataframe,
compute_benchmarks,
load_school_data,
load_latest_school_data,
@@ -32,27 +35,36 @@ from .data_loader import (
get_supplementary_data,
get_supplementary_data_batch,
search_schools_typesense,
suggest_schools_typesense,
)
from .data_loader import get_data_info as get_db_info
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from . import flags
from .places import build_place_index, build_place_registry, places_for_urn
from .schemas import METRIC_DEFINITIONS, PHASE_GROUPS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .nearby_schools import select_nearby
from .utils import clean_for_json, convert_to_native
# Values to exclude from filter dropdowns (empty strings, non-applicable labels)
EXCLUDED_FILTER_VALUES = {"", "Not applicable", "Does not apply"}
# Maps user-facing phase filter values to the GIAS PhaseOfEducation values they include.
# All-through schools appear in both primary and secondary results.
PHASE_GROUPS: dict[str, set[str]] = {
"primary": {"primary", "middle deemed primary", "all-through"},
"secondary": {"secondary", "middle deemed secondary", "all-through", "16 plus"},
"all-through": {"all-through"},
}
BASE_URL = "https://schoolcompare.co.uk"
# Must match SITE_URL in nextjs-app/lib/site.ts. The apex 301s to www, and a
# sitemap <loc> that redirects wastes a crawl on every URL it lists.
BASE_URL = "https://www.schoolcompare.co.uk"
MAX_SLUG_LENGTH = 60
# In-memory sitemap cache
_sitemap_xml: str | None = None
# In-memory sitemap cache: name -> XML. Populated on startup and by the admin
# regenerate endpoint after a pipeline run.
_sitemaps: dict[str, str] | None = None
# Built from the same DataFrame the sitemap uses, so places and sitemap can
# never describe different corpora. Reset by the same admin endpoint.
_place_registry: dict | None = None
# Cached beside the registry, and invalidated by identity against it — see
# get_place_index. Never cleared independently.
_place_index: dict | None = None
_place_index_source: dict | None = None
VALID_PLACE_KINDS = ("town", "locality", "authority", "outcode")
def _slugify(text: str) -> str:
@@ -70,43 +82,287 @@ def _school_url(urn: int, school_name: str) -> str:
return f"/school/{urn}-{slug}"
def build_sitemap() -> str:
"""Generate sitemap XML from in-memory school data. Returns the XML string."""
df = load_school_data()
# Routes worth submitting that are not a school page. /admissions was missing
# from the sitemap entirely despite being a static, indexable guide.
STATIC_SITEMAP_PATHS = ("/", "/rankings", "/compare", "/admissions")
static_urls = [
(BASE_URL + "/", "daily", "1.0"),
(BASE_URL + "/rankings", "weekly", "0.8"),
(BASE_URL + "/compare", "weekly", "0.8"),
]
lines = ['<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">']
# A page has something a search result could state if any of these is present
# in any year. Shared by _has_publishable_data and the per-school check in
# _school_sitemap_rows so the two can never drift.
_PUBLISHABLE_FIELDS = ("rwm_expected_pct", "attainment_8_score", "ofsted_grade")
for url, freq, priority in static_urls:
lines.append(
f" <url><loc>{url}</loc>"
f"<changefreq>{freq}</changefreq>"
f"<priority>{priority}</priority></url>"
def _has_publishable_data(row) -> bool:
"""True when a school page has something a search result could state.
A school with no results in any year and no Ofsted grade renders an empty
page. Submitting it spends crawl budget and drags the corpus-wide quality
signal down, so it stays out of the sitemap. The page itself still resolves
for anyone who has the URL.
"""
for field in _PUBLISHABLE_FIELDS:
value = row.get(field)
if value is not None and not pd.isna(value):
return True
return False
def _url_element(loc: str, lastmod: str | None = None) -> str:
"""One <url> entry. No priority or changefreq — Google ignores both."""
body = f"<loc>{loc}</loc>"
if lastmod:
body += f"<lastmod>{lastmod}</lastmod>"
return f" <url>{body}</url>"
def _school_sitemap_rows(df) -> list[str]:
"""A <url> element per school that has something to show.
lastmod comes from the school's Ofsted date where there is one and is
omitted otherwise. An always-now lastmod is a claim Google learns to
distrust; an absent one honestly means "unknown".
"""
if df.empty or "urn" not in df.columns or "school_name" not in df.columns:
return []
rows: list[str] = []
seen: set[int] = set()
# Publishable is a property of the SCHOOL, not of its latest row.
#
# The first cut tested the latest year's row alone, which quietly dropped
# every school that has results in its history but a null row for the most
# recent year — a school that stopped reporting, or whose figures were
# suppressed for small-cohort disclosure. The Mallard Academy (150367) is
# the case that caught it: real KS2 results for 2015-16 through 2018-19,
# then null rows for 2022-23 onward. Its page shows all four years; the
# sitemap omitted it. Roughly 220 schools were affected.
publishable_cols = [c for c in _PUBLISHABLE_FIELDS if c in df.columns]
publishable: set[int] = (
set(df.loc[df[publishable_cols].notna().any(axis=1), "urn"].astype(int))
if publishable_cols else set()
)
# Latest row per URN first, so a school's most recent Ofsted date wins.
ordered = df.sort_values("year", ascending=False) if "year" in df.columns else df
for _, row in ordered.iterrows():
urn = int(row["urn"])
if urn in seen:
continue
seen.add(urn)
if urn not in publishable:
continue
lastmod = None
ofsted_date = row.get("ofsted_date")
if ofsted_date is not None and not pd.isna(ofsted_date):
lastmod = pd.Timestamp(ofsted_date).date().isoformat()
rows.append(_url_element(
BASE_URL + _school_url(urn, str(row["school_name"])), lastmod))
return rows
# Sitemaps cap at 50,000 URLs per file. 10,000 keeps a child small enough to
# scan by eye in Search Console, which is the point of splitting at all:
# coverage is reported per submitted sitemap, so one file per page family is
# what makes an indexation problem attributable to a family.
SITEMAP_CHUNK_SIZE = 10_000
# Children are served under /sitemaps/ because Next.js only treats a whole
# bracketed path segment as dynamic — a route folder named "sitemap-[...parts]"
# is read as a literal static segment and never matches.
SITEMAP_CHILD_PREFIX = "/sitemaps"
def get_place_registry() -> dict:
"""The place registry, built once and cached for the process."""
global _place_registry
if _place_registry is None:
_place_registry = build_place_registry(load_school_data())
return _place_registry
def get_place_index() -> dict:
"""URN → its published places, cached against the registry it came from.
Invalidation is an identity check rather than a second flag to remember to
clear. Anything that drops `_place_registry` — the tests all do — gets a
fresh registry object here, which no longer matches the one the index was
built from, so the index rebuilds with it. A separate `_place_index = None`
would be one more thing to forget, and a stale reverse index is exactly the
bug that would put links to another dataset's places on a school page.
"""
global _place_index, _place_index_source
registry = get_place_registry()
if _place_index is None or _place_index_source is not registry:
_place_index = build_place_index(registry)
_place_index_source = registry
return _place_index
def _urlset(rows: list[str]) -> str:
return "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
*rows,
"</urlset>",
])
def _place_url(place) -> str:
"""The canonical path for a place. Two namespaces, per the spec.
Towns and localities share /schools/[place]; authorities take their own
prefix because 67 town names collide with an authority name and neither
set contains the other.
"""
if place.kind == "authority":
return f"/schools/authority/{place.slug}"
if place.kind == "outcode":
return f"/schools/near/{place.slug}"
return f"/schools/{place.slug}"
def _places_payload(urn: int) -> list[dict]:
"""The published places containing this school, as the school page needs
them: a name to write in the link, a count so the anchor can say what it
leads to, and the canonical path.
`phases` carries the phase variants this school actually appears on, which
is usually one and is two for an all-through school — it is listed on both
pages, so there is no tie to break.
Membership is read straight from the registry's own `phase_urns` rather
than re-derived from the school's phase string. The registry is the one
place that decides which phases a place publishes and who is on them;
computing it a second time here is how a page comes to link a school to a
phase page that does not list it, or to a route that does not exist. That
is also why outcodes need no special case: they carry empty `phase_urns`,
so they report no phase links on their own.
"""
payload = []
for place in places_for_urn(get_place_index(), int(urn)):
phases = [
{
"phase": phase,
"count": len(phase_urns),
"url": f"{_place_url(place)}/{phase}",
}
for phase, phase_urns in sorted(place.phase_urns.items())
if int(urn) in phase_urns
]
payload.append({
"kind": place.kind,
"slug": place.slug,
"name": place.name,
"count": len(place.urns),
"url": _place_url(place),
"phases": phases,
})
return payload
def _nearby_schools_payload(urn: int) -> list[dict]:
"""The nearest eligible schools this page may offer, closest first.
Phase and reach are read from the school's own row inside select_nearby,
so nothing here can hand it a phase that disagrees with the data.
Wrapped: a failure in selection must never 500 a page that is otherwise
complete, which is the posture get_supplementary_data already takes. The
section simply does not render.
"""
try:
return select_nearby(load_latest_school_data(), int(urn))
except Exception:
import logging
logging.getLogger(__name__).exception(
"Nearby schools selection failed for urn=%s", urn
)
return []
if not df.empty and "urn" in df.columns and "school_name" in df.columns:
seen = set()
for _, row in df[["urn", "school_name"]].drop_duplicates(subset="urn").iterrows():
urn = int(row["urn"])
name = str(row["school_name"])
if urn in seen:
continue
seen.add(urn)
path = _school_url(urn, name)
lines.append(
f" <url><loc>{BASE_URL}{path}</loc>"
f"<changefreq>monthly</changefreq>"
f"<priority>0.6</priority></url>"
)
lines.append("</urlset>")
return "\n".join(lines)
def _place_sitemap_rows(kinds: tuple[str, ...], registry=None) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
Phase is part of the query — "primary schools in beccles" — so each
variant is its own indexable page. Submitting only the bare place URL left
~950 of them reachable by nothing: absent from every sitemap, and not
linked from the place page either.
"""
rows: list[str] = []
if registry is None:
registry = get_place_registry()
for p in sorted(registry.values(), key=lambda p: (p.kind, p.slug)):
if p.kind not in kinds:
continue
rows.append(_url_element(BASE_URL + _place_url(p)))
# Which phases a place publishes is the registry's decision alone —
# outcodes report none, because the spec gives them no phase route.
# Repeating that rule here was how the page and the sitemap came to
# disagree about which URLs exist.
for phase in ("primary", "secondary"):
if p.publishes_phase(phase):
rows.append(_url_element(f"{BASE_URL}{_place_url(p)}/{phase}"))
return rows
def build_sitemaps(df=None, registry=None) -> dict[str, str]:
"""Build the sitemap index and every child, keyed by name."""
if df is None:
df = load_school_data()
children: dict[str, str] = {
"static.xml": _urlset(
[_url_element(BASE_URL + path) for path in STATIC_SITEMAP_PATHS]),
}
school_rows = _school_sitemap_rows(df)
# Always emit at least one school child, so the index shape is stable even
# on an empty database.
chunks = [school_rows[i:i + SITEMAP_CHUNK_SIZE]
for i in range(0, len(school_rows), SITEMAP_CHUNK_SIZE)] or [[]]
for n, chunk in enumerate(chunks, start=1):
children[f"schools-{n}.xml"] = _urlset(chunk)
# Separate children per family: Search Console reports coverage per
# submitted sitemap, which is how the location layer's indexation is
# measured apart from the school pages'.
for label, kinds in (("places", ("town", "locality", "authority")),
("outcodes", ("outcode",))):
rows = _place_sitemap_rows(kinds, registry)
chunks = [rows[i:i + SITEMAP_CHUNK_SIZE]
for i in range(0, len(rows), SITEMAP_CHUNK_SIZE)] or [[]]
for n, chunk in enumerate(chunks, start=1):
children[f"{label}-{n}.xml"] = _urlset(chunk)
# On a sitemap index, lastmod means "when this sitemap file last changed",
# so generation time is the correct value here — unlike on a <url>, where
# it would be a claim about content we cannot support.
generated = datetime.now(timezone.utc).date().isoformat()
index_rows = [
f" <sitemap><loc>{BASE_URL}{SITEMAP_CHILD_PREFIX}/{name}</loc>"
f"<lastmod>{generated}</lastmod></sitemap>"
for name in children
]
index = "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
'<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
*index_rows,
"</sitemapindex>",
])
return {**children, "sitemap.xml": index}
def build_sitemap() -> str:
"""The sitemap index. Kept for `lifespan` and the admin endpoint."""
return build_sitemaps()["sitemap.xml"]
def clean_filter_values(series: pd.Series) -> list[str]:
@@ -121,8 +377,101 @@ def clean_filter_values(series: pd.Series) -> list[str]:
# SECURITY MIDDLEWARE & HELPERS
# =============================================================================
# Rate limiter
limiter = Limiter(key_func=get_remote_address)
def client_key(request: Request) -> str:
"""The rate-limit bucket: the real caller, not the proxy in front of them.
`get_remote_address` reads request.client.host. In staging and production
the backend has no published ports and sits on the internal network, so its
only caller is the Next proxy — meaning every browser user on the site
shared one bucket. Measured before this fix: 70 concurrent requests to
/api/schools returned 60 OK and 10 refused.
CF-Connecting-IP first, because Cloudflare (in front of both environments)
sets it on every origin request and *overwrites* any client-supplied value,
which a parsed X-Forwarded-For chain does not guarantee. The XFF fallback is
forgeable, but only by a caller already inside the Docker network, which is
the one place nothing untrusted can reach.
"""
cf = request.headers.get("cf-connecting-ip")
if cf:
return cf.strip()
xff = request.headers.get("x-forwarded-for")
if xff:
return xff.split(",")[0].strip()
return get_remote_address(request)
# Per-client limiter. Paired with the global ceiling below — the two do
# different jobs and neither substitutes for the other.
limiter = Limiter(key_func=client_key)
# --- The ceiling no header can raise ----------------------------------------
#
# client_key trusts CF-Connecting-IP, and nothing in this process can tell an
# edge-set header from an attacker-set one. That distinction can only be made
# at Cloudflare, with Authenticated Origin Pulls or an origin firewall. A
# caller reaching the origin directly could otherwise mint a fresh rate-limit
# bucket per request and evade per-client limits entirely — which would make
# correct keying a net regression against abuse, since the single shared bucket
# it replaced at least capped everyone at 60/minute together.
#
# So per-client limits give fairness, and this gives the origin a hard total.
# It does not make the header trustworthy; it bounds what trusting it can cost.
# The header problem itself is closed at Cloudflare, not here.
#
# [window_start_monotonic, count], or None before the first request. A fixed
# window is crude, which is right for a backstop: it has to be obviously
# correct rather than fair.
_global_window: Optional[list] = None
# The container healthcheck runs `curl http://localhost:80/api/data-info` from
# inside the container. Starving it would fail the check, restart the
# container, and turn a load spike into an outage loop — the ceiling exists to
# protect the origin, not to kill it.
_LOCAL_HOSTS = frozenset({"127.0.0.1", "::1", "localhost"})
def exempt_from_ceiling(request: Request) -> bool:
"""Whether the ceiling should ignore this request.
Its own function so the rule is testable without standing up a server —
and so the healthcheck exemption is somewhere a reader can find it.
"""
if not request.url.path.startswith("/api/"):
return True
# The peer address, never the Host header: Host is set by the caller and
# would hand every attacker an exemption.
return (request.client.host if request.client else "") in _LOCAL_HOSTS
class GlobalRateLimitMiddleware(BaseHTTPMiddleware):
"""A cap on total /api/ traffic, independent of any client identity."""
async def dispatch(self, request: Request, call_next):
global _global_window
if exempt_from_ceiling(request):
return await call_next(request)
now = time.monotonic()
# One event loop, and no await between the read and the write, so this
# sequence is atomic without a lock.
if _global_window is None or now - _global_window[0] >= 60:
_global_window = [now, 0]
_global_window[1] += 1
if _global_window[1] > settings.global_rate_limit_per_minute:
return JSONResponse(
# Distinguishable from slowapi's per-client 429: an operator
# reading logs has to be able to tell "one noisy client" from
# "the origin is saturated".
{"detail": "The service is at capacity. Please retry shortly."},
status_code=429,
headers={"Retry-After":
str(max(1, int(60 - (now - _global_window[0]))))},
)
return await call_next(request)
class SecurityHeadersMiddleware(BaseHTTPMiddleware):
@@ -180,6 +529,7 @@ CACHE_RULES: list[tuple[str, tuple[int, int, int]]] = [
("/api/schools/", (300, 3600, 86400)), # /api/schools/{urn}
("/api/rankings", (60, 600, 3600)),
("/api/compare", (60, 600, 3600)),
("/api/suggest", (60, 3600, 86400)), # autosuggest
("/api/schools", (30, 300, 1800)), # search list
]
@@ -284,7 +634,8 @@ def validate_postcode(postcode: Optional[str]) -> Optional[str]:
@asynccontextmanager
async def lifespan(app: FastAPI):
"""Application lifespan - startup and shutdown events."""
global _sitemap_xml
global _sitemaps
flags.init()
print("Loading school data from marts...")
df = load_school_data()
if df.empty:
@@ -294,9 +645,9 @@ async def lifespan(app: FastAPI):
# Pre-compute the latest-year snapshot so the first search request is fast
await asyncio.to_thread(load_latest_school_data)
try:
_sitemap_xml = build_sitemap()
n = _sitemap_xml.count("<url>")
print(f"Sitemap built: {n} URLs.")
_sitemaps = build_sitemaps()
n = sum(x.count("<url>") for x in _sitemaps.values())
print(f"Sitemaps built: {len(_sitemaps)} files, {n} URLs.")
except Exception as e:
print(f"Warning: sitemap build failed on startup: {e}")
@@ -326,6 +677,10 @@ app.add_middleware(CacheAndETagMiddleware)
app.add_middleware(SecurityHeadersMiddleware)
app.add_middleware(RequestSizeLimitMiddleware)
app.add_middleware(GZipMiddleware, minimum_size=512)
# Added last, so it is outermost and refuses before anything downstream does
# work. A ceiling that only applies after the expensive part has run is not a
# ceiling.
app.add_middleware(GlobalRateLimitMiddleware)
# CORS middleware - restricted for production
app.add_middleware(
@@ -363,6 +718,15 @@ async def get_config():
}
@app.get("/api/release")
async def release_identity():
import json
from pathlib import Path
path = Path(__file__).with_name("build-info.json")
identity = json.loads(path.read_text()) if path.exists() else {"sha": "development", "build_id": "development"}
return JSONResponse(identity, headers={"Cache-Control": "no-store"})
@app.get("/api/schools")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_schools(
@@ -399,7 +763,7 @@ async def get_schools(
df_latest = load_latest_school_data()
if df_latest.empty:
return {"schools": [], "total": 0, "page": page, "page_size": 0}
raise HTTPException(status_code=503, detail="School data temporarily unavailable")
# Use configured default if not specified
if page_size is None:
@@ -498,8 +862,8 @@ async def get_schools(
# Apply filters
if search:
ts_urns = search_schools_typesense(search)
if ts_urns:
ts_urns = await asyncio.to_thread(search_schools_typesense, search)
if ts_urns is not None:
urn_order = {urn: i for i, urn in enumerate(ts_urns)}
schools_df = schools_df[schools_df["urn"].isin(set(ts_urns))].copy()
schools_df["_ts_rank"] = schools_df["urn"].map(urn_order)
@@ -507,9 +871,9 @@ async def get_schools(
else:
# Fallback: Typesense unavailable, use substring match
search_lower = search.lower()
mask = schools_df["school_name"].str.lower().str.contains(search_lower, na=False)
mask = schools_df["school_name"].str.lower().str.contains(search_lower, na=False, regex=False)
if "address" in schools_df.columns:
mask = mask | schools_df["address"].str.lower().str.contains(search_lower, na=False)
mask = mask | schools_df["address"].str.lower().str.contains(search_lower, na=False, regex=False)
schools_df = schools_df[mask]
if local_authority:
@@ -568,7 +932,7 @@ async def get_school_details(request: Request, urn: int):
df = load_school_data()
if df.empty:
raise HTTPException(status_code=404, detail="No data available")
raise HTTPException(status_code=503, detail="School data temporarily unavailable")
school_data = df[df["urn"] == urn]
@@ -626,16 +990,35 @@ async def get_school_details(request: Request, urn: int):
return {
"school_info": school_info,
# Where this school sits in the location layer, for the page's link
# module and breadcrumb. Derived from the same registry the place
# pages and the sitemap use, so a link is never offered for a page
# that does not exist. Empty is a valid answer: a school whose town
# and authority both fall below the publish threshold has nowhere to
# point, and the page renders without the module.
"places": _places_payload(urn),
# The nearest eligible schools, closest first. Always present on a
# build with this code; the frontend treats absent and empty
# identically, which is what lets the two images deploy independently.
"nearby_schools": _nearby_schools_payload(urn),
"yearly_data": clean_for_json(school_data),
# Supplementary data (null if not yet populated by Kestra)
"ofsted": supplementary.get("ofsted"),
"census": supplementary.get("census"),
"admissions": supplementary.get("admissions"),
"admissions_history": supplementary.get("admissions_history") or [],
# Behind a flag, and withheld at the source rather than rendered-but-
# hidden: this endpoint is public and unauthenticated, so a field left
# in the payload is a published field. The key is absent, not null —
# null would state that this school has no cut-off, which is a
# different claim from "we are not publishing cut-offs".
**({"admission_distance": supplementary.get("admission_distance")}
if flags.is_enabled("admission_distance") else {}),
"sen_detail": supplementary.get("sen_detail"),
"phonics": supplementary.get("phonics"),
"deprivation": supplementary.get("deprivation"),
"finance": supplementary.get("finance"),
"destinations": supplementary.get("destinations"),
}
@@ -1006,6 +1389,145 @@ async def get_rankings(
}
@app.get("/api/places")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def list_places(request: Request):
"""Every published place. The sitemap and the link modules read this."""
registry = get_place_registry()
return {"places": [
{"kind": p.kind, "slug": p.slug, "name": p.name, "count": len(p.urns)}
for p in sorted(registry.values(), key=lambda p: (p.kind, p.slug))
]}
@app.get("/api/places/{kind}/{slug}")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_place(request: Request, kind: str, slug: str,
phase: Optional[str] = None):
"""One place: its schools ranked, and its local averages."""
if kind not in VALID_PLACE_KINDS:
raise HTTPException(status_code=404, detail="No such place")
registry = get_place_registry()
place = registry.get(f"{kind}:{slug}")
if place is None:
raise HTTPException(status_code=404, detail="No such place")
df = load_latest_school_data()
rows = df[df["urn"].isin(place.urns)]
if phase:
wanted = PHASE_GROUPS.get(phase.lower())
if wanted and "phase" in rows.columns:
rows = rows[rows["phase"].fillna("").str.lower().isin(wanted)]
# The metric the page shows, and averages.
metric = "attainment_8_score" if phase == "secondary" else "rwm_expected_pct"
# Alphabetical, not by score. A place page is read by someone looking for
# a school they can name, and scanning for it is what the order should
# serve. /rankings is where the league-table ordering lives, and it keeps
# sorting by metric.
if "school_name" in rows.columns:
rows = rows.sort_values("school_name", key=lambda c: c.str.lower())
averages = {
m: (None if m not in rows.columns or rows[m].dropna().empty
else float(rows[m].dropna().mean()))
for m in ("rwm_expected_pct", "attainment_8_score")
}
# dict.fromkeys, not a list: SCHOOL_COLUMNS already ends with latitude and
# longitude, so concatenating them again selected each twice and pandas
# dropped one of every duplicated pair with a "columns are not unique"
# warning. Ordered de-duplication keeps the column order and the warning
# cannot come back.
#
# nursery_provision and parliamentary_constituency are not in
# SCHOOL_COLUMNS and the place table shows both. The `in rows.columns`
# guard is what keeps a mart the pipeline has not rebuilt working: those
# two are the optional GIAS columns data_loader degrades to NULL.
cols = [c for c in dict.fromkeys(
SCHOOL_COLUMNS + ["latitude", "longitude", "phase",
"nursery_provision",
"parliamentary_constituency",
"rwm_expected_pct", "attainment_8_score",
"total_pupils"])
if c in rows.columns]
return {
"place": {"kind": place.kind, "slug": place.slug, "name": place.name,
"count": len(place.urns),
"parent_authority": place.parent_authority,
# Every authority the place meaningfully sits in. SW19 is
# mostly Merton but partly Wandsworth; naming one asserts
# something false.
#
# The slug is null where that authority has no page of its
# own: City of London and the Isles of Scilly hold fewer
# schools than the threshold. Naming them is still right;
# linking them would be a 404.
"authorities": [
{"name": name,
"slug": (_slugify(name)
if f"authority:{_slugify(name)}" in registry
else None),
"count": n}
for name, n in place.authorities
],
# Only phases that clear the threshold, so the page links
# variants that exist rather than 404s.
"phases": [ph for ph in ("primary", "secondary")
if place.publishes_phase(ph)]},
"schools": clean_for_json(rows[cols]),
"averages": averages,
}
# Two characters. One is not a query — it matches thousands of schools and the
# response is useless, so it is not worth a round trip.
SUGGEST_MIN_QUERY = 2
@app.get("/api/suggest")
@limiter.limit("120/minute")
async def suggest_schools(
request: Request,
q: str = Query("", max_length=100),
limit: int = Query(8, ge=1, le=20),
):
"""School name suggestions, from Typesense alone.
Deliberately not a mode of /api/schools: that path filters and sorts the
full in-memory DataFrame, which is far too expensive to run per keystroke.
Nothing here returns an error for ordinary input. A short query, no
matches, or Typesense being unreachable are all 200 with an empty list —
a dropdown that quietly does not appear is the right failure for a
keystroke path, and there is no DataFrame fallback because the 25,000-row
substring scan is precisely what this endpoint exists to avoid.
120/minute rather than the default 60: a 200 ms debounce makes typing
legitimately bursty.
"""
query = q.strip()
if len(query) < SUGGEST_MIN_QUERY:
return {"suggestions": []}
return {"suggestions": suggest_schools_typesense(query, limit)}
@app.get("/api/flags")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_feature_flags(request: Request):
"""Every declared flag and its current value.
Internal only. The Next proxy denies this path, because the response names
every unreleased feature the codebase knows about — which is exactly what
shipping dark is meant to keep quiet.
"""
return flags.all_flags()
@app.get("/api/data-info")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_data_info(request: Request):
@@ -1051,20 +1573,51 @@ async def get_data_info(request: Request):
}
_publication_lock = asyncio.Lock()
def _prepare_publication(df):
if df.empty:
raise ValueError("Refusing to publish an empty school dataset")
if not {"urn", "year", "school_name"}.issubset(df.columns):
raise ValueError("School dataset is missing required columns")
if df["urn"].isna().any() or df.duplicated(["urn", "year"]).any():
raise ValueError("School dataset has missing URNs or duplicate school years")
latest = build_latest_school_data(df)
registry = build_place_registry(df)
index = build_place_index(registry)
sitemaps = build_sitemaps(df, registry)
return df, latest, registry, index, sitemaps
def _publish(prepared):
# Called on the event loop with no await: routes cannot observe half a swap.
# The application currently runs one worker; replicas require coordination.
from . import data_loader
global _place_registry, _place_index, _place_index_source, _sitemaps
df, latest, registry, index, sitemaps = prepared
data_loader._df_cache = df
data_loader._df_latest_cache = latest
_place_registry = registry
_place_index = index
_place_index_source = registry
_sitemaps = sitemaps
@app.post("/api/admin/reload")
@limiter.limit("5/minute")
async def reload_data(
request: Request,
_: bool = Depends(verify_admin_api_key)
):
"""
Admin endpoint to force data reload (useful after data updates).
Requires X-API-Key header with valid admin API key.
"""
clear_cache()
await asyncio.to_thread(load_school_data)
await asyncio.to_thread(load_latest_school_data)
return {"status": "reloaded"}
async def reload_data(request: Request, _: bool = Depends(verify_admin_api_key)):
"""Validate a complete replacement before publishing it; retain data on failure."""
async with _publication_lock:
try:
df = await asyncio.to_thread(load_school_data_as_dataframe)
prepared = await asyncio.to_thread(_prepare_publication, df)
except Exception as exc:
import logging
logging.getLogger(__name__).exception("Dataset reload failed")
raise HTTPException(status_code=503, detail="Dataset reload failed; previous data retained") from exc
_publish(prepared)
return {"status": "reloaded", "schools": len(prepared[1])}
@@ -1086,16 +1639,28 @@ async def robots_txt():
return FileResponse(settings.frontend_dir / "robots.txt", media_type="text/plain")
@app.get("/sitemap.xml")
async def sitemap_xml():
"""Serve sitemap.xml for search engine indexing."""
global _sitemap_xml
if _sitemap_xml is None:
def _serve_sitemap(name: str) -> Response:
global _sitemaps
if _sitemaps is None:
try:
_sitemap_xml = build_sitemap()
_sitemaps = build_sitemaps()
except Exception as e:
raise HTTPException(status_code=503, detail=f"Sitemap unavailable: {e}")
return Response(content=_sitemap_xml, media_type="application/xml")
if name not in _sitemaps:
raise HTTPException(status_code=404, detail="No such sitemap")
return Response(content=_sitemaps[name], media_type="application/xml")
@app.get("/sitemap.xml")
async def sitemap_xml():
"""Serve the sitemap index."""
return _serve_sitemap("sitemap.xml")
@app.get("/sitemaps/{name}")
async def sitemap_child(name: str):
"""Serve a child sitemap (static.xml, or schools-N.xml)."""
return _serve_sitemap(name)
@app.post("/api/admin/regenerate-sitemap")
@@ -1104,11 +1669,16 @@ async def regenerate_sitemap(
request: Request,
_: bool = Depends(verify_admin_api_key),
):
"""Rebuild and cache the sitemap from current school data. Called by Airflow after data updates."""
global _sitemap_xml
_sitemap_xml = build_sitemap()
n = _sitemap_xml.count("<url>")
return {"status": "ok", "urls": n}
"""Rebuild derived publication data without clearing the live registry."""
async with _publication_lock:
try:
prepared = await asyncio.to_thread(_prepare_publication, load_school_data())
except Exception as exc:
raise HTTPException(status_code=503, detail="Sitemap rebuild failed; previous data retained") from exc
_publish(prepared)
n = sum(x.count("<url>") for x in prepared[4].values())
return {"status": "ok", "urls": n, "sitemaps": len(prepared[4])}
# Mount static files directly (must be after all routes to avoid catching API calls)
+15
View File
@@ -35,6 +35,11 @@ class Settings(BaseSettings):
# Security
admin_api_key: str = Field(default_factory=lambda: secrets.token_urlsafe(32))
rate_limit_per_minute: int = 60 # Requests per minute per IP
# A ceiling on total /api/ traffic, independent of any client identity.
# client_key trusts headers only Cloudflare can vouch for, so a caller
# reaching the origin directly could otherwise mint a fresh bucket per
# request. See GlobalRateLimitMiddleware in backend/app.py.
global_rate_limit_per_minute: int = 3000
rate_limit_burst: int = 10 # Allow burst of requests
max_request_size: int = 1024 * 1024 # 1MB max request size
@@ -42,6 +47,16 @@ class Settings(BaseSettings):
typesense_url: str = "http://localhost:8108"
typesense_api_key: str = ""
# Feature flags (Unleash). An empty unleash_url disables flags entirely and
# every flag evaluates False — the correct behaviour for local development
# and CI, and the reason no test needs a running Unleash.
unleash_url: str = ""
unleash_api_token: str = ""
unleash_app_name: str = "schoolcompare-backend"
# On a named volume, so a restart during an Unleash outage keeps
# last-known state instead of reverting a released feature to dark.
unleash_cache_directory: str = "/app/.unleash"
# Analytics
ga_measurement_id: Optional[str] = "G-J0PCVT14NY" # Google Analytics 4 Measurement ID
+386 -19
View File
@@ -18,8 +18,9 @@ from .config import settings
from .database import SessionLocal, engine
from .models import (
DimSchool, DimLocation, KS2Performance,
FactOfstedInspection, FactAdmissions,
FactOfstedInspection, FactAdmissions, FactAdmissionDistance,
FactDeprivation, FactFinance, FactPupilCharacteristics,
FactKs4Destinations, FactKs5Destinations,
)
from .ofsted_codes import ofsted_page_url, report_card_labels
from .schemas import SCHOOL_TYPE_MAP
@@ -83,22 +84,111 @@ def _get_typesense_client():
return None
def search_schools_typesense(query: str, limit: int = 250) -> List[int]:
"""Search Typesense. Returns URNs in relevance order, or [] if unavailable."""
SEARCH_PAGE_SIZE = 250
# Search results are filtered again by the API (authority, phase, postcode,
# etc.), so one page is too small for scoped searches. Keep the candidate set
# bounded, though: a broad query must not turn into an unbounded sequence of
# Typesense requests. Four pages is enough to preserve useful scoped matches
# while putting a hard ceiling on latency and upstream load.
SEARCH_MAX_CANDIDATES = 1_000
def search_schools_typesense(query: str) -> Optional[List[int]]:
"""Return a bounded set of matching URNs in relevance order.
``None`` means Typesense is unavailable; ``[]`` is a valid zero-match
result. The API applies its remaining filters after this search, so the
first few pages are fetched rather than only the first page. Once the
candidate ceiling is reached, the relevance-ordered prefix is returned on
purpose; fetching every match would make common or adversarial queries
unbounded.
"""
client = _get_typesense_client()
if client is None:
return None
urns: list[int] = []
fetched = 0
try:
page = 1
while fetched < SEARCH_MAX_CANDIDATES:
page_size = min(SEARCH_PAGE_SIZE, SEARCH_MAX_CANDIDATES - fetched)
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority,postcode",
"per_page": page_size,
"page": page,
"typo_tokens_threshold": 1,
})
hits = result.get("hits", [])
urns.extend(int(h["document"]["urn"]) for h in hits)
fetched += len(hits)
if fetched >= result.get("found", fetched):
return list(dict.fromkeys(urns))
if not hits:
raise ValueError("Search pagination ended before all matches arrived")
page += 1
logging.getLogger(__name__).info(
"Typesense search capped at %d candidates for query %r",
SEARCH_MAX_CANDIDATES,
query,
)
return list(dict.fromkeys(urns))
except Exception:
logging.getLogger(__name__).exception("School search unavailable")
return None
# The most a public endpoint will return in one response.
SUGGEST_MAX_LIMIT = 20
# Fields a suggestion row carries, and the default when the document omits an
# optional one. phase and school_type are optional in the Typesense schema.
_SUGGEST_FIELDS = ("school_name", "local_authority", "postcode",
"phase", "school_type")
def suggest_schools_typesense(query: str, limit: int = 8) -> List[dict]:
"""Autosuggest rows straight from Typesense. Never raises.
Returns documents rather than URNs, unlike search_schools_typesense, so the
caller needs no DataFrame. Every field below is already in the index — see
pipeline/scripts/sync_typesense.py — which is what makes this cheap enough
to run per keystroke.
"""
client = _get_typesense_client()
if client is None:
return []
try:
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority,postcode",
"per_page": min(limit, 250),
"query_by": "school_name,local_authority",
"per_page": max(1, min(limit, SUGGEST_MAX_LIMIT)),
"typo_tokens_threshold": 1,
})
return [int(h["document"]["urn"]) for h in result.get("hits", [])]
except Exception:
# A dropdown that quietly stops appearing is the right failure here.
return []
rows = []
for hit in result.get("hits", []) or []:
doc = (hit or {}).get("document") or {}
try:
urn = int(doc["urn"])
except (KeyError, TypeError, ValueError):
# Skip the row, keep the rest. Typesense declares urn as int32 so
# this should be unreachable, but the index is a separate system
# that something other than this code can reindex — and "never
# raises" is a promise the keystroke path actually depends on.
# Dropping one malformed document is right; blanking the whole
# dropdown, or serving a suggestion pointing at /school/0, is not.
logging.getLogger(__name__).warning(
"skipping malformed suggestion document: %r", doc)
continue
row = {"urn": urn}
row.update({f: str(doc.get(f, "") or "") for f in _SUGGEST_FIELDS})
rows.append(row)
return rows
def normalize_school_type(school_type: Optional[str]) -> Optional[str]:
"""Convert cryptic school type codes to user-friendly names."""
@@ -135,16 +225,6 @@ def geocode_single_postcode(postcode: str) -> Optional[Tuple[float, float]]:
return None
def haversine_distance(lat1: float, lon1: float, lat2: float, lon2: float) -> float:
"""Calculate great-circle distance between two points (miles)."""
from math import radians, cos, sin, asin, sqrt
lat1, lon1, lat2, lon2 = map(radians, [lat1, lon1, lat2, lon2])
dlat = lat2 - lat1
dlon = lon2 - lon1
a = sin(dlat / 2) ** 2 + cos(lat1) * cos(lat2) * sin(dlon / 2) ** 2
return 2 * asin(sqrt(a)) * 3956
# =============================================================================
# MAIN DATA LOAD — joins dim_school + dim_location + fact_performance
# fact_performance is a merged KS2+KS4 table (one row per URN per year).
@@ -459,7 +539,12 @@ def load_latest_school_data() -> pd.DataFrame:
if _df_latest_cache is not None:
return _df_latest_cache
df = load_school_data()
_df_latest_cache = build_latest_school_data(load_school_data())
return _df_latest_cache
def build_latest_school_data(df: pd.DataFrame) -> pd.DataFrame:
"""Build a replacement snapshot without mutating the published caches."""
if df.empty:
return df
@@ -492,8 +577,7 @@ def load_latest_school_data() -> pd.DataFrame:
df_latest = pd.concat([df_latest, df_no_perf], ignore_index=True)
print(f"Latest-snapshot cache built: {len(df_latest)} schools")
_df_latest_cache = df_latest
return _df_latest_cache
return df_latest
def clear_cache():
@@ -723,6 +807,17 @@ def _admissions_row_dict(a) -> dict:
}
def _admission_distance_dict(d) -> dict:
"""Serialize one fact_admission_distance row for API responses."""
return {
"year": d.year,
"distance_m": d.distance_m,
"route_count": d.route_count,
"la_name": d.la_name,
"distance_unit_raw": d.distance_unit_raw,
}
def _census_dict(pc) -> dict:
return {
"year": pc.year,
@@ -753,16 +848,230 @@ def _finance_dict(f) -> dict:
}
# Destination measures that are totals DfE published itself, rather than one of
# the categories that partition the cohort.
_AGGREGATE_MEASURES = {"agg_sustained_education", "agg_sustained_all"}
def _format_cohort_year(year) -> str | None:
"""202223 -> '2022/23'.
The section has to date its own cohort. Destination measures run about two
GCSE years behind the results shown above them on the same page, so an
undated figure reads as stale data rather than as a different question.
"""
if not year:
return None
text = str(year)
if len(text) == 6:
return f"{text[:4]}/{text[4:6]}"
if len(text) == 8:
return f"{text[:4]}/{text[6:8]}"
return text
_PUPIL_GROUPS = ("disadvantaged", "other", "all")
def _lone_hidden_groups(groups: dict) -> list:
"""Pupil groups hiding exactly one category — solvable by subtraction."""
return [
key for key, group in groups.items()
if sum(1 for c in group["categories"] if c["status"] == "suppressed") == 1
]
def _lone_hidden_categories(groups: dict) -> list:
"""Categories hidden in exactly one of several pupil groups."""
lone = []
categories = {c["category"] for g in groups.values() for c in g["categories"]}
for category in categories:
found = [
c for g in groups.values() for c in g["categories"]
if c["category"] == category
]
hidden = [c for c in found if c["status"] == "suppressed"]
if len(hidden) == 1 and len(found) > 1:
lone.append(category)
return lone
def disclosure_invariant_holds(groups: dict) -> bool:
"""Every row and every column hides none, or at least two.
Public so the tests can assert it directly rather than re-deriving it.
"""
return not _lone_hidden_groups(groups) and not _lone_hidden_categories(groups)
def _mask_for_disclosure(groups: dict) -> None:
"""Withhold further cells until nothing suppressed can be solved for.
Not rendering a figure is not the same as not publishing it. This endpoint
is public and unauthenticated, so anything left in the payload is
published, whatever the UI chooses to draw — the same reasoning the
admission_distance field carries in app.py.
Two identities let a caller solve for a withheld cell:
* within a pupil group, the categories sum to the cohort, so a group with
exactly ONE suppressed category gives it away as cohort - sum(rest);
* across groups, disadvantaged + other = all for every category, so a
category suppressed in exactly ONE of the three gives itself away.
DfE's own answer is secondary suppression: withhold a second cell so the
residual spans two unknowns and identifies neither.
Where no companion can do that — a sparse cohort whose every other category
is `not_applicable`, which is common in special schools and alternative
provision — there is nothing left to withhold, so the pupil group is
DROPPED entirely. An earlier version simply gave up here and returned with
the violation intact and no signal, which is the one outcome this function
must never produce: a disclosure-control pass that fails silently is worse
than none, because everything downstream trusts it.
Mutates `groups` in place. Guaranteed to return with
disclosure_invariant_holds(groups) true.
"""
def suppress(cell):
if cell["status"] == "published":
cell["status"] = "suppressed"
cell["pupils"] = None
cell["percentage"] = None
return True
return False
def add_companion(candidates) -> bool:
"""Withhold a second cell so the residual spans two unknowns.
The companion must carry pupils. Suppressing a zero looks like
secondary suppression and protects nothing: the residual still equals
the original withheld figure exactly. Returns False when no cell can
do the job, which escalates to dropping the group.
"""
published = [c for c in candidates if c["status"] == "published"]
useful = sorted(
(c for c in published if (c["pupils"] or 0) > 0),
key=lambda c: c["pupils"],
)
if useful:
return suppress(useful[0])
# Every remaining cell is zero or not applicable: withholding any of
# them leaves the residual equal to the original figure.
return False
# Fixpoint: each new suppression can break the other identity. Terminates
# because every pass either adds a suppression, drops a group, or stops.
while not disclosure_invariant_holds(groups):
changed = False
for category in _lone_hidden_categories(groups):
siblings = [
c for g in groups.values() for c in g["categories"]
if c["category"] == category
]
if add_companion(siblings):
changed = True
for key in _lone_hidden_groups(groups):
if add_companion(groups[key]["categories"]):
changed = True
if changed:
continue
# Nothing left to withhold. Drop the groups that are still solvable,
# and any category still solvable across the groups that remain.
for key in _lone_hidden_groups(groups):
del groups[key]
changed = True
for category in _lone_hidden_categories(groups):
for group in groups.values():
for cell in group["categories"]:
if cell["category"] == category and suppress(cell):
changed = True
if not changed:
# Unreachable given the two escalations above, but a masking pass
# must never spin or exit unsafely. Withhold everything.
groups.clear()
return
def _destinations_block(rows: list) -> dict | None:
"""Shape destination rows for one phase into the API's block.
Applies secondary suppression before returning, so no caller of this public
endpoint can solve for a figure DfE withheld. See _mask_for_disclosure.
Aggregate measures are dropped entirely. DfE publishes them, and they would
be useful for a "what is published for this group" fallback, but nothing
renders them today and an aggregate spanning exactly one suppressed
component names that component. An unused field that leaks is not a
trade-off worth carrying — re-add them with their own guard if the fallback
is ever built.
Deliberately computes no residual, no "remaining pupils" figure, and no
total that would close a gap left by a suppressed category.
"""
if not rows:
return None
years = [r["year"] for r in rows if r.get("year") is not None]
if not years:
return None
latest_year = max(years)
rows = [r for r in rows if r.get("year") == latest_year]
groups: dict = {}
for row in rows:
group = groups.setdefault(
row["pupil_group"],
{"cohort": row.get("cohort_pupils"), "categories": []},
)
measure = row["destination_measure"]
published = row.get("status") == "published"
# Belt and braces: percentage is derived from the same source cell as
# pupils, but publishing one without the other would hand back the
# cohort (pupils / percentage) and with it the residual.
cell = {
"category": measure,
"pupils": row.get("pupils") if published else None,
"percentage": row.get("percentage") if published else None,
"status": row.get("status"),
}
if measure in _AGGREGATE_MEASURES:
continue
group["categories"].append(cell)
if not groups:
return None
_mask_for_disclosure(groups)
# Masking can empty the block entirely — a sparse cohort where no group
# could be made safe. Return None so the section is absent rather than
# rendering an empty shell.
if not groups:
return None
return {"cohort_year": _format_cohort_year(latest_year), "groups": groups}
def _empty_supplementary() -> dict:
return {
"ofsted": None,
"census": None,
"admissions": None,
"admissions_history": [],
"admission_distance": None,
"sen_detail": None,
"phonics": None,
"deprivation": None,
"finance": None,
"destinations": None,
}
@@ -837,6 +1146,32 @@ def get_supplementary_data_batch(db: Session, urns: list[int]) -> dict:
result[urn]["admissions"] = rows_for_urn[-1] if rows_for_urn else None
_safe(_admissions)
# Last distance offered — the latest year per URN, and only that.
#
# The mart holds every published year and the DAG keeps loading them; what
# changed is what leaves this process. Earlier years are being held back as
# a paid feature, and this API is public and unauthenticated — serving the
# history here would hand it to anyone who opened the network tab, whatever
# the page chose to render. Withholding it in the client would have been
# decoration, not a decision.
#
# Restoring it for entitled callers is a change to this function, not to
# the pipeline: fact_admission_distance is untouched and complete.
def _admission_distance():
rows = (
db.query(FactAdmissionDistance)
.filter(FactAdmissionDistance.urn.in_(urns))
.order_by(FactAdmissionDistance.urn, FactAdmissionDistance.year.desc())
.all()
)
seen = set()
for d in rows:
if d.urn in seen:
continue
seen.add(d.urn)
result[d.urn]["admission_distance"] = _admission_distance_dict(d)
_safe(_admission_distance)
# Deprivation — one row per URN.
def _deprivation():
rows = (
@@ -864,6 +1199,38 @@ def get_supplementary_data_batch(db: Session, urns: list[int]) -> dict:
result[f.urn]["finance"] = _finance_dict(f)
_safe(_finance)
# Destinations — KS4 and 16-18. Both marts are long-format, so every row
# for a URN is collected and _destinations_block picks the latest year and
# shapes the pupil groups. A phase with no rows serialises as null rather
# than an empty shell, so the frontend renders nothing rather than an empty
# section.
def _destinations():
from collections import defaultdict
def _collect(model):
per_urn = defaultdict(list)
for r in db.query(model).filter(model.urn.in_(urns)).all():
per_urn[r.urn].append({
"year": r.year,
"pupil_group": r.pupil_group,
"destination_measure": r.destination_measure,
"cohort_pupils": r.cohort_pupils,
"pupils": r.pupils,
"percentage": r.percentage,
"status": r.status,
})
return per_urn
ks4_rows = _collect(FactKs4Destinations)
ks5_rows = _collect(FactKs5Destinations)
for urn in urns:
ks4 = _destinations_block(ks4_rows.get(urn, []))
ks5 = _destinations_block(ks5_rows.get(urn, []))
result[urn]["destinations"] = (
{"ks4": ks4, "ks5": ks5} if (ks4 or ks5) else None
)
_safe(_destinations)
return result
+135
View File
@@ -0,0 +1,135 @@
"""Feature flags: what can be switched, and what is switched right now.
Ship-dark, not a kill switch. Flags let work merge and deploy without becoming
visible; they are expected to flip about monthly, by a person, deliberately.
Nothing here does percentage rollouts or user targeting — the site has no user
identity to target.
Unleash holds the state. It does not hold the list. REGISTRY below is that
list, and it exists for three reasons: the SDK evaluates an unknown flag to
False, so without a registry that is an *undeclared* False, indistinguishable
from a typo; /api/flags needs a key set to return when Unleash is unreachable;
and a flag in the UI but not in the registry is orphaned and should be visibly
so rather than quietly authoritative.
Every flag defaults to False. There is no per-flag default, because a flag that
defaults on is a kill switch, and this is not one.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from datetime import date
from .config import settings
logger = logging.getLogger(__name__)
# A flag is temporary scaffolding. See test_a_flag_older_than_the_limit.
MAX_FLAG_AGE_DAYS = 90
@dataclass(frozen=True)
class Flag:
# One string: the registry key, the Unleash flag name, and the JSON key in
# /api/flags. snake_case, matching the API's existing convention. No case
# transformation anywhere, so there is no mapping layer to get wrong.
name: str
description: str # one line: what turning this on reveals
added: date # for the staleness tripwire
REGISTRY: dict[str, Flag] = {
f.name: f for f in (
Flag(
name="admission_distance",
description=(
"The last-distance-offered figure on the Admissions tile and "
"the 'How far away are you?' section on school pages."
),
added=date(2026, 8, 23),
),
Flag(
name="school_autosuggest",
description=(
"School name suggestions as you type in the main search box."
),
added=date(2026, 8, 26),
),
Flag(
name="about_page",
description=(
"The /about page, its footer link, its sitemap entry, and the "
"named-author byline on every blog post."
),
added=date(2026, 9, 8),
),
Flag(
name="blog",
description=(
"The /blog index, post pages, the RSS feed, their footer link "
"and their sitemap entries. Not /admin: posts must be "
"writable before the blog is readable."
),
added=date(2026, 9, 8),
),
)
}
_client = None
def init() -> None:
"""Start the Unleash client, or log why flags are all off.
Called once from the app lifespan. Never raises: a flag system that can
stop the API from booting is worse than one that is switched off.
"""
global _client
if not settings.unleash_url or not settings.unleash_api_token:
logger.warning(
"Unleash is not configured (UNLEASH_URL / UNLEASH_API_TOKEN); "
"every feature flag evaluates to False.")
return
try:
from UnleashClient import UnleashClient
_client = UnleashClient(
url=settings.unleash_url,
app_name=settings.unleash_app_name,
custom_headers={"Authorization": settings.unleash_api_token},
cache_directory=settings.unleash_cache_directory,
refresh_interval=15,
)
_client.initialize_client()
logger.info("Unleash client initialised against %s", settings.unleash_url)
except Exception:
# Fail closed and keep serving. The SDK also evaluates everything False
# until its first successful sync, so this is the same direction.
_client = None
logger.exception("Unleash client failed to start; flags are all False.")
def is_enabled(name: str) -> bool:
"""Whether `name` is on. False for anything unknown, unreachable or broken."""
if name not in REGISTRY:
logger.error(
"undeclared feature flag %r was evaluated; returning False. "
"Add it to backend/flags.py REGISTRY or fix the name.", name)
return False
if _client is None:
return False
try:
return bool(_client.is_enabled(
name, fallback_function=lambda feature_name, context: False))
except Exception:
logger.exception("flag %r failed to evaluate; returning False", name)
return False
def all_flags() -> dict[str, bool]:
"""Every declared flag and its current value. Serves /api/flags."""
return {name: is_enabled(name) for name in REGISTRY}
+54
View File
@@ -0,0 +1,54 @@
"""Curated London localities, defined by the postcode districts they cover.
The GIAS `town` field puts 1,819 London schools under the single value
"London", so it cannot answer "schools in Battersea" — a query that appears in
the Search Console baseline. No single field can: parliamentary constituency
gives Battersea but not Canary Wharf; postcodes.io's admin_ward gives Canary
Wharf but not Battersea; neither gives Clapham or Shoreditch, which are postal
and colloquial rather than administrative.
So this is curated. Where a locality ends is a judgement, not a fact, and a
reviewable file is the honest place for a judgement. No new ingestion is
needed — the corpus already carries postcodes.
This is the canonical copy. `pipeline/transform/seeds/locality_outcodes.csv`
mirrors it for anyone querying the warehouse directly; the backend image does
not contain `pipeline/`, which is why the module rather than the seed is
canonical. Same arrangement as `backend/gias_codes.py`.
A locality whose outcodes hold fewer than MIN_SCHOOLS schools is not
published, so a typo produces no page rather than an empty one. Places that
fail that check are logged at startup, because a locality you meant to publish
quietly not appearing is the failure worth hearing about.
Two rules for anything added here.
**Sub-borough districts only.** A London borough is a local authority and
already has a page at /schools/authority/[la] covering all of its schools; a
locality defined by two or three outcodes would be a partial, near-duplicate
subset of it. Hackney, Islington, Greenwich and Ealing were all in the first
draft for that reason and have been removed.
**The slug must not match a GIAS town.** "Richmond" did — GIAS has a Richmond
in North Yorkshire with 37 schools — so the London one could never publish.
The registry skips any locality that collides and logs it.
"""
# slug -> (display name, outcodes)
LOCALITY_OUTCODES: dict[str, tuple[str, tuple[str, ...]]] = {
"battersea": ("Battersea", ("SW11",)),
"canary-wharf": ("Canary Wharf", ("E14",)),
"clapham": ("Clapham", ("SW4",)),
"shoreditch": ("Shoreditch", ("EC2A", "E1")),
"peckham": ("Peckham", ("SE15",)),
"brixton": ("Brixton", ("SW2", "SW9")),
"camden-town": ("Camden Town", ("NW1",)),
"wimbledon": ("Wimbledon", ("SW19",)),
"putney": ("Putney", ("SW15",)),
"fulham": ("Fulham", ("SW6",)),
"chiswick": ("Chiswick", ("W4",)),
"stratford": ("Stratford", ("E15",)),
"walthamstow": ("Walthamstow", ("E17",)),
"tooting": ("Tooting", ("SW17",)),
"dulwich": ("Dulwich", ("SE21", "SE22")),
}
-512
View File
@@ -1,512 +0,0 @@
"""
Database migration logic for importing CSV data.
Used by both CLI script and automatic startup migration.
"""
import re
from pathlib import Path
from typing import Dict, Optional
import numpy as np
import pandas as pd
import requests
from .config import settings
from .database import Base, engine, get_db_session
from .models import School, SchoolResult
from .schemas import (
COLUMN_MAPPINGS,
LA_CODE_TO_NAME,
NULL_VALUES,
SCHOOL_TYPE_MAP,
)
def parse_numeric(value) -> Optional[float]:
"""Parse a numeric value, handling special cases."""
if pd.isna(value):
return None
if isinstance(value, (int, float)):
return float(value) if not np.isnan(value) else None
str_val = str(value).strip().upper()
if str_val in NULL_VALUES or str_val == "":
return None
# Remove percentage signs if present
str_val = str_val.replace("%", "")
try:
return float(str_val)
except ValueError:
return None
def extract_year_from_folder(folder_name: str) -> Optional[int]:
"""Extract year from folder name like '2023-2024'."""
match = re.search(r"(\d{4})-(\d{4})", folder_name)
if match:
return int(match.group(2))
match = re.search(r"(\d{4})", folder_name)
if match:
return int(match.group(1))
return None
def geocode_postcodes_bulk(postcodes: list) -> Dict[str, tuple]:
"""
Geocode postcodes in bulk using postcodes.io API.
Returns dict of postcode -> (latitude, longitude).
"""
results = {}
valid_postcodes = [
p.strip().upper()
for p in postcodes
if p and isinstance(p, str) and len(p.strip()) >= 5
]
valid_postcodes = list(set(valid_postcodes))
if not valid_postcodes:
return results
batch_size = 100
total_batches = (len(valid_postcodes) + batch_size - 1) // batch_size
for i, batch_start in enumerate(range(0, len(valid_postcodes), batch_size)):
batch = valid_postcodes[batch_start : batch_start + batch_size]
print(
f" Geocoding batch {i + 1}/{total_batches} ({len(batch)} postcodes)..."
)
try:
response = requests.post(
"https://api.postcodes.io/postcodes",
json={"postcodes": batch},
timeout=30,
)
if response.status_code == 200:
data = response.json()
for item in data.get("result", []):
if item and item.get("result"):
pc = item["query"].upper()
lat = item["result"].get("latitude")
lon = item["result"].get("longitude")
if lat and lon:
results[pc] = (lat, lon)
except Exception as e:
print(f" Warning: Geocoding batch failed: {e}")
return results
def load_csv_data(data_dir: Path) -> pd.DataFrame:
"""Load all CSV data from data directory."""
all_data = []
for folder in sorted(data_dir.iterdir()):
if not folder.is_dir():
continue
year = extract_year_from_folder(folder.name)
if not year:
continue
# Specifically look for the KS2 results file
ks2_file = folder / "england_ks2final.csv"
if not ks2_file.exists():
continue
csv_file = ks2_file
print(f" Loading {csv_file.name} (year {year})...")
try:
df = pd.read_csv(csv_file, encoding="latin-1", low_memory=False)
except Exception as e:
print(f" Error loading {csv_file}: {e}")
continue
# Rename columns
df.rename(columns=COLUMN_MAPPINGS, inplace=True)
df["year"] = year
# Handle local authority name
la_name_cols = ["LANAME", "LA (name)", "LA_NAME", "LA NAME"]
la_name_col = next((c for c in la_name_cols if c in df.columns), None)
if la_name_col and la_name_col != "local_authority":
df["local_authority"] = df[la_name_col]
elif "LEA" in df.columns:
df["local_authority_code"] = pd.to_numeric(df["LEA"], errors="coerce")
df["local_authority"] = (
df["local_authority_code"]
.map(LA_CODE_TO_NAME)
.fillna(df["LEA"].astype(str))
)
# Store LEA code
if "LEA" in df.columns:
df["local_authority_code"] = pd.to_numeric(df["LEA"], errors="coerce")
# Map school type
if "school_type_code" in df.columns:
df["school_type"] = (
df["school_type_code"]
.map(SCHOOL_TYPE_MAP)
.fillna(df["school_type_code"])
)
# Create combined address
addr_parts = ["address1", "address2", "town", "postcode"]
for col in addr_parts:
if col not in df.columns:
df[col] = None
df["address"] = df.apply(
lambda r: ", ".join(
str(v)
for v in [
r.get("address1"),
r.get("address2"),
r.get("town"),
r.get("postcode"),
]
if pd.notna(v) and str(v).strip()
),
axis=1,
)
all_data.append(df)
print(f" Loaded {len(df)} records")
if all_data:
result = pd.concat(all_data, ignore_index=True)
print(f"\nTotal records loaded: {len(result)}")
print(f"Unique schools: {result['urn'].nunique()}")
print(f"Years: {sorted(result['year'].unique())}")
return result
return pd.DataFrame()
def migrate_data(df: pd.DataFrame, geocode: bool = False, geocode_cache: dict = None):
"""Migrate DataFrame data to database."""
if geocode_cache is None:
geocode_cache = {}
# Clean URN column - convert to integer, drop invalid values
df = df.copy()
df["urn"] = pd.to_numeric(df["urn"], errors="coerce")
df = df.dropna(subset=["urn"])
df["urn"] = df["urn"].astype(int)
# Group by URN to get unique schools (use latest year's data)
school_data = (
df.sort_values("year", ascending=False).groupby("urn").first().reset_index()
)
print(f"\nMigrating {len(school_data)} unique schools...")
# Geocode postcodes that aren't already in the cache
geocoded = dict(geocode_cache) # start with preserved coordinates
if geocode and "postcode" in df.columns:
cached_postcodes = {
str(row.get("postcode", "")).strip().upper()
for _, row in school_data.iterrows()
if int(float(str(row.get("urn", 0) or 0))) in geocode_cache
}
postcodes_needed = [
p for p in df["postcode"].dropna().unique()
if str(p).strip().upper() not in cached_postcodes
]
if postcodes_needed:
print(f"\nGeocoding {len(postcodes_needed)} postcodes ({len(geocode_cache)} restored from cache)...")
fresh = geocode_postcodes_bulk(postcodes_needed)
geocoded.update(fresh)
print(f" Successfully geocoded {len(fresh)} new postcodes")
else:
print(f"\nAll {len(geocode_cache)} postcodes restored from cache, skipping geocoding.")
with get_db_session() as db:
# Create schools
urn_to_school_id = {}
schools_created = 0
for _, row in school_data.iterrows():
# Safely parse URN - handle None, NaN, whitespace, and invalid values
urn_val = row.get("urn")
urn = None
if pd.notna(urn_val):
try:
urn_str = str(urn_val).strip()
if urn_str:
urn = int(float(urn_str)) # Handle "12345.0" format
except (ValueError, TypeError):
pass
if not urn:
continue
# Skip if we've already added this URN (handles duplicates in source data)
if urn in urn_to_school_id:
continue
# Get geocoding data
postcode = row.get("postcode")
lat, lon = None, None
if postcode and pd.notna(postcode):
coords = geocoded.get(str(postcode).strip().upper())
if coords:
lat, lon = coords
# Safely parse local_authority_code
la_code = None
la_code_val = row.get("local_authority_code")
if pd.notna(la_code_val):
try:
la_code_str = str(la_code_val).strip()
if la_code_str:
la_code = int(float(la_code_str))
except (ValueError, TypeError):
pass
school = School(
urn=urn,
school_name=row.get("school_name")
if pd.notna(row.get("school_name"))
else "Unknown",
local_authority=row.get("local_authority")
if pd.notna(row.get("local_authority"))
else None,
local_authority_code=la_code,
school_type=row.get("school_type")
if pd.notna(row.get("school_type"))
else None,
school_type_code=row.get("school_type_code")
if pd.notna(row.get("school_type_code"))
else None,
religious_denomination=row.get("religious_denomination")
if pd.notna(row.get("religious_denomination"))
else None,
age_range=row.get("age_range")
if pd.notna(row.get("age_range"))
else None,
address1=row.get("address1") if pd.notna(row.get("address1")) else None,
address2=row.get("address2") if pd.notna(row.get("address2")) else None,
town=row.get("town") if pd.notna(row.get("town")) else None,
postcode=row.get("postcode") if pd.notna(row.get("postcode")) else None,
latitude=lat,
longitude=lon,
)
db.add(school)
db.flush() # Get the ID
urn_to_school_id[urn] = school.id
schools_created += 1
if schools_created % 1000 == 0:
print(f" Created {schools_created} schools...")
print(f" Created {schools_created} schools")
# Create results
print(f"\nMigrating {len(df)} yearly results...")
results_created = 0
for _, row in df.iterrows():
# Safely parse URN
urn_val = row.get("urn")
urn = None
if pd.notna(urn_val):
try:
urn_str = str(urn_val).strip()
if urn_str:
urn = int(float(urn_str))
except (ValueError, TypeError):
pass
if not urn or urn not in urn_to_school_id:
continue
school_id = urn_to_school_id[urn]
# Safely parse year
year_val = row.get("year")
year = None
if pd.notna(year_val):
try:
year = int(float(str(year_val).strip()))
except (ValueError, TypeError):
pass
if not year:
continue
result = SchoolResult(
school_id=school_id,
year=year,
total_pupils=parse_numeric(row.get("total_pupils")),
eligible_pupils=parse_numeric(row.get("eligible_pupils")),
# Expected Standard
rwm_expected_pct=parse_numeric(row.get("rwm_expected_pct")),
reading_expected_pct=parse_numeric(row.get("reading_expected_pct")),
writing_expected_pct=parse_numeric(row.get("writing_expected_pct")),
maths_expected_pct=parse_numeric(row.get("maths_expected_pct")),
gps_expected_pct=parse_numeric(row.get("gps_expected_pct")),
science_expected_pct=parse_numeric(row.get("science_expected_pct")),
# Higher Standard
rwm_high_pct=parse_numeric(row.get("rwm_high_pct")),
reading_high_pct=parse_numeric(row.get("reading_high_pct")),
writing_high_pct=parse_numeric(row.get("writing_high_pct")),
maths_high_pct=parse_numeric(row.get("maths_high_pct")),
gps_high_pct=parse_numeric(row.get("gps_high_pct")),
# Progress
reading_progress=parse_numeric(row.get("reading_progress")),
writing_progress=parse_numeric(row.get("writing_progress")),
maths_progress=parse_numeric(row.get("maths_progress")),
# Averages
reading_avg_score=parse_numeric(row.get("reading_avg_score")),
maths_avg_score=parse_numeric(row.get("maths_avg_score")),
gps_avg_score=parse_numeric(row.get("gps_avg_score")),
# Context
disadvantaged_pct=parse_numeric(row.get("disadvantaged_pct")),
eal_pct=parse_numeric(row.get("eal_pct")),
sen_support_pct=parse_numeric(row.get("sen_support_pct")),
sen_ehcp_pct=parse_numeric(row.get("sen_ehcp_pct")),
stability_pct=parse_numeric(row.get("stability_pct")),
# Absence
reading_absence_pct=parse_numeric(row.get("reading_absence_pct")),
gps_absence_pct=parse_numeric(row.get("gps_absence_pct")),
maths_absence_pct=parse_numeric(row.get("maths_absence_pct")),
writing_absence_pct=parse_numeric(row.get("writing_absence_pct")),
science_absence_pct=parse_numeric(row.get("science_absence_pct")),
# Gender
rwm_expected_boys_pct=parse_numeric(row.get("rwm_expected_boys_pct")),
rwm_expected_girls_pct=parse_numeric(row.get("rwm_expected_girls_pct")),
rwm_high_boys_pct=parse_numeric(row.get("rwm_high_boys_pct")),
rwm_high_girls_pct=parse_numeric(row.get("rwm_high_girls_pct")),
# Disadvantaged
rwm_expected_disadvantaged_pct=parse_numeric(
row.get("rwm_expected_disadvantaged_pct")
),
rwm_expected_non_disadvantaged_pct=parse_numeric(
row.get("rwm_expected_non_disadvantaged_pct")
),
disadvantaged_gap=parse_numeric(row.get("disadvantaged_gap")),
# 3-Year
rwm_expected_3yr_pct=parse_numeric(row.get("rwm_expected_3yr_pct")),
reading_avg_3yr=parse_numeric(row.get("reading_avg_3yr")),
maths_avg_3yr=parse_numeric(row.get("maths_avg_3yr")),
)
db.add(result)
results_created += 1
if results_created % 10000 == 0:
print(f" Created {results_created} results...")
db.flush()
print(f" Created {results_created} results")
# Commit all changes
db.commit()
print("\nMigration complete!")
def _apply_schema_alterations():
"""
Add new columns to existing tables using ALTER TABLE … ADD COLUMN IF NOT EXISTS.
Safe to run on every migration — no-ops if the column already exists.
Add entries here whenever models.py gains new columns on an existing table.
"""
alterations = [
# v4: Ofsted Report Card columns
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS framework VARCHAR(20)",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_safeguarding_met BOOLEAN",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_inclusion INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_curriculum_teaching INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_achievement INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_attendance_behaviour INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_personal_development INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_leadership_governance INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_early_years INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_sixth_form INTEGER",
]
from sqlalchemy import text as sa_text
with engine.connect() as conn:
for stmt in alterations:
try:
conn.execute(sa_text(stmt))
except Exception as e:
print(f" Warning: alteration skipped ({e})")
conn.commit()
def _apply_schema_drops():
"""
Drop tables retired from the schema. Idempotent (DROP … IF EXISTS), so it's
safe to run on every migration. Add entries here when a model is removed.
"""
drops = [
# v6: Ofsted Parent View feature removed
"DROP TABLE IF EXISTS marts.fact_parent_view CASCADE",
]
from sqlalchemy import text as sa_text
with engine.connect() as conn:
for stmt in drops:
try:
conn.execute(sa_text(stmt))
except Exception as e:
print(f" Warning: drop skipped ({e})")
conn.commit()
def run_full_migration(geocode: bool = False) -> bool:
"""
Run a complete migration: drop all tables and reimport from CSV.
Returns True if successful, False if no data found.
Raises exception on error.
"""
# Preserve existing geocoding so a reimport doesn't throw away coordinates
# that took a long time to compute.
geocode_cache: dict[int, tuple[float, float]] = {}
inspector = __import__("sqlalchemy").inspect(engine)
if "schools" in inspector.get_table_names():
try:
with get_db_session() as db:
rows = db.execute(
__import__("sqlalchemy").text(
"SELECT urn, latitude, longitude FROM schools "
"WHERE latitude IS NOT NULL AND longitude IS NOT NULL"
)
).fetchall()
geocode_cache = {r.urn: (r.latitude, r.longitude) for r in rows}
print(f" Saved {len(geocode_cache)} existing geocoded coordinates.")
except Exception as e:
print(f" Warning: could not save geocode cache: {e}")
# Only drop the core KS2 tables — leave supplementary tables (ofsted, census,
# finance, etc.) intact so a reimport doesn't wipe integrator-populated data.
# schema_version is NOT dropped: it persists so restarts don't re-trigger migration.
ks2_tables = ["school_results", "schools"]
print(f"Dropping core tables: {ks2_tables} ...")
inspector = __import__("sqlalchemy").inspect(engine)
existing = set(inspector.get_table_names())
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
print("Creating all tables...")
Base.metadata.create_all(bind=engine)
# ALTER existing supplementary tables to add any new columns.
# create_all() only creates missing tables; it won't add columns to tables
# that already exist from an older schema version. These statements are
# idempotent (IF NOT EXISTS) so they're safe to run on every migration.
print("Applying column additions to supplementary tables...")
_apply_schema_alterations()
print("Dropping retired tables...")
_apply_schema_drops()
print("\nLoading CSV data...")
df = load_csv_data(settings.data_dir)
if df.empty:
print("Warning: No CSV data found to migrate!")
return False
migrate_data(df, geocode=geocode, geocode_cache=geocode_cache)
return True
+76
View File
@@ -188,6 +188,37 @@ class FactAdmissions(Base):
admissions_policy = Column(String(100))
class FactAdmissionDistance(Base):
"""Last distance offered — one row per URN per year.
Separate from FactAdmissions because the source is separate: EES publishes
admissions for the whole country, whereas cut-off distances exist only for
the local authorities that choose to publish them (57 at the time of
writing), and the two refresh on unrelated timetables.
"""
__tablename__ = "fact_admission_distance"
__table_args__ = (
Index("ix_admission_distance_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
# Straight-line distance in metres from the school to the last home offered
# a place that year.
distance_m = Column(Float)
# How many admission routes (ability bands, separate reception/junior
# intakes) were collapsed into distance_m. >1 means the figure is the
# furthest of several and the page must say so.
route_count = Column(Integer)
la_code = Column(Integer)
la_name = Column(String(100))
# The unit the council published in, so the page can lead with the unit a
# parent was given rather than always converting.
distance_unit_raw = Column(String(20))
source_file = Column(Text)
class FactPupilCharacteristics(Base):
"""School pupil composition from EES census — one row per URN per year."""
__tablename__ = "fact_pupil_characteristics"
@@ -290,3 +321,48 @@ class Ks2NationalAverage(Base):
gps_high_pct = Column(Float)
gps_avg_score = Column(Float)
science_expected_pct = Column(Float)
class FactKs4Destinations(Base):
"""KS4 leavers destinations — one row per URN, year, pupil group, measure.
Long format rather than wide because pupil_group is a real third dimension.
`status` is load-bearing: 'suppressed' means DfE withheld a figure it
considered disclosive and the page must print "withheld"; 'not_applicable'
means the measure does not apply and the page must print nothing. `pupils`
is null for both, so collapsing status to a null check loses the
difference — and the categories sum to the cohort, so a consumer that
treats a withheld cell as zero republishes what DfE hid.
"""
__tablename__ = "fact_ks4_destinations"
__table_args__ = (
Index("ix_ks4_dest_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
pupil_group = Column(String(20), primary_key=True)
destination_measure = Column(String(40), primary_key=True)
cohort_pupils = Column(Integer)
pupils = Column(Integer)
percentage = Column(Float)
status = Column(String(20))
class FactKs5Destinations(Base):
"""16-18 study leavers destinations — same grain as FactKs4Destinations."""
__tablename__ = "fact_ks5_destinations"
__table_args__ = (
Index("ix_ks5_dest_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
pupil_group = Column(String(20), primary_key=True)
destination_measure = Column(String(40), primary_key=True)
cohort_pupils = Column(Integer)
pupils = Column(Integer)
percentage = Column(Float)
status = Column(String(20))
+259
View File
@@ -0,0 +1,259 @@
"""Which nearby schools a detail page may offer as alternatives.
HARD FILTERS decide eligibility, and encode claims the section is not allowed
to make. A selective school is not an alternative to a non-selective one, a
special school is not comparable to a mainstream one, and a Girls school is not
an option for a Boys school's reader. They never relax, at any distance, even
where that means the section does not render at all.
DISTANCE decides the order, and nothing else does.
An earlier version ranked by intake similarity first and used distance only as
a tiebreak. That put a Catholic school 2.9 miles away above the community
school 0.3 miles down the road, and — because the row filled from the best tier
before widening — filled all six slots with faith matches while omitting every
school a parent could actually walk to. For a primary, a school that far is not
a weaker option; it is not an option. Distance is a constraint and intake is a
preference, and the ranking now says so.
Similarity survives as `shared`: what a candidate genuinely has in common with
this school, reported on its card, so a reader applies their own weighting
instead of having ours applied for them.
Pure functions over a DataFrame: no I/O, no FastAPI, no database.
"""
from __future__ import annotations
import re
import numpy as np
import pandas as pd
from .schemas import PHASE_GROUPS
# Three fit the row; the rest are behind the carousel arrows.
MAX_SCHOOLS = 6
MINIMUM = 2
# How far the section will reach, in miles, when nothing closer exists.
#
# A sanity bound rather than a target: ordering by distance already handles
# density, so a school in a dense area fills all six slots inside a mile and
# never sees this. It decides one thing — what happens where the area is
# sparse — and the answer differs by phase because catchments do. Primary
# catchments are routinely under a mile; beyond two, a primary is not a weaker
# option but not an option, and no section is the honest answer.
PRIMARY_RADIUS_MILES = 2.0
SECONDARY_RADIUS_MILES = 6.0
POST16_RADIUS_MILES = 10.0
EARTH_RADIUS_MILES = 3958.8
_SPECIAL = re.compile(r"\bspecial\b|pupil referral|alternative provision", re.I)
# Values that mean "this school has no religious character".
_NO_FAITH = {"", "none", "does not apply", "not applicable"}
def is_special_provision(school_type: str | None) -> bool:
"""Mirror of isSpecialSchool() in nextjs-app/lib/utils.ts.
Special schools carry a mainstream phase, so phase alone cannot identify
them. The two implementations must agree: a school the frontend treats as
special for benchmarking but this treats as mainstream would be dropped
from its own England comparison and then offered as a peer to a mainstream
school on the next page along.
"""
return bool(_SPECIAL.search(school_type or ""))
def is_selective(admissions_policy: str | None) -> bool:
"""Strictly selective. Unknown counts as non-selective, which is the safe
direction: it can only ever exclude a pairing, never invent one."""
return (admissions_policy or "").strip().lower() == "selective"
def faith_key(denomination: str | None) -> str:
value = (denomination or "").strip().lower()
return "" if value in _NO_FAITH else value
def faith_label(denomination: str | None) -> str:
return denomination.strip() if faith_key(denomination) else "No religious character"
def genders_compatible(a: str | None, b: str | None) -> bool:
single = {"boys", "girls"}
left, right = (a or "").strip().lower(), (b or "").strip().lower()
return not (left in single and right in single and left != right)
def is_secondary_phase(phase: str | None) -> bool:
"""Whether this phase takes the secondary side: secondary group membership,
minus all-through.
Membership is read from PHASE_GROUPS rather than tested with `"secondary" in
phase`, because that substring misses "16 plus" — GIAS phase 6, which
PHASE_GROUPS deliberately files as secondary. The substring version fails
silently rather than loudly: a sixth-form college is simply handed the
primary bucket and offered infant schools as peers.
All-through is the exception. PHASE_GROUPS lists it on both sides because it
belongs on both phases' place pages, but the detail page renders it with the
primary template, and the metric follows the template.
"""
text = (phase or "").strip().lower()
return text != "all-through" and text in PHASE_GROUPS["secondary"]
def radius_miles(phase: str | None) -> float:
"""How far this phase's section will reach when nothing closer exists."""
if (phase or "").strip().lower() == "16 plus":
return POST16_RADIUS_MILES
return SECONDARY_RADIUS_MILES if is_secondary_phase(phase) else PRIMARY_RADIUS_MILES
def _phase_group(is_secondary: bool) -> set[str]:
return PHASE_GROUPS["secondary" if is_secondary else "primary"]
def _haversine_miles(lat1: float, lon1: float, lat2, lon2):
"""Vectorised, matching the postcode search in app.py."""
lat1_r, lon1_r = np.radians(lat1), np.radians(lon1)
lat2_r, lon2_r = np.radians(lat2.astype(float)), np.radians(lon2.astype(float))
dlat, dlon = lat2_r - lat1_r, lon2_r - lon1_r
a = np.sin(dlat / 2) ** 2 + np.cos(lat1_r) * np.cos(lat2_r) * np.sin(dlon / 2) ** 2
return 2 * EARTH_RADIUS_MILES * np.arcsin(np.sqrt(a))
def _native(value):
"""NaN and numpy scalars both reach JSONResponse badly; normalise here so
the caller never has to remember to."""
if value is None:
return None
if isinstance(value, np.generic):
value = value.item()
if isinstance(value, float) and np.isnan(value):
return None
return value
def _mask(series: pd.Series, predicate) -> pd.Series:
"""A boolean mask that survives an empty frame.
`Series.apply` on an empty Series returns an empty *DataFrame*, and using
that as a mask silently drops every column — so the next column lookup
raises KeyError rather than yielding no rows. This is not hypothetical: a
special school with no special school near it empties the frame at the
provision filter, which is the ordinary case for most special schools.
"""
return pd.Series([predicate(value) for value in series], index=series.index, dtype=bool)
def _shared(subject: pd.Series, candidate: pd.Series, is_secondary: bool) -> list[str]:
"""What this candidate genuinely has in common with the subject.
Empty is a real answer, and renders no chips at all. A card claiming a
shared characteristic it does not have would be worse than a bare one —
and since these no longer affect the order, an empty list costs the school
nothing but its place in the row, which distance already decided.
"""
shared: list[str] = []
gender = str(subject.get("gender") or "").strip()
if gender and str(candidate.get("gender") or "").strip().lower() == gender.lower():
shared.append(gender)
if is_secondary:
policy = str(candidate.get("admissions_policy") or "").strip()
subject_policy = str(subject.get("admissions_policy") or "").strip()
if (
policy
and policy.lower() == subject_policy.lower()
and policy.lower() not in {"not applicable", "unknown"}
):
shared.append(policy)
if faith_key(candidate.get("religious_denomination")) == faith_key(
subject.get("religious_denomination")
):
shared.append(faith_label(candidate.get("religious_denomination")))
return shared
def select_nearby(frame: pd.DataFrame, urn: int) -> list[dict]:
"""The nearest eligible schools, closest first — at most MAX_SCHOOLS, and
none at all below MINIMUM.
The phase is read from the subject's own row rather than passed in, so a
caller cannot hand this a phase that disagrees with the data it selects
from.
"""
subject_rows = frame[frame["urn"] == urn]
if subject_rows.empty:
return []
subject = subject_rows.iloc[0]
lat, lon = _native(subject.get("latitude")), _native(subject.get("longitude"))
if lat is None or lon is None:
return []
phase = subject.get("phase")
is_secondary = is_secondary_phase(phase)
reach = radius_miles(phase)
metric_key = "attainment_8_score" if is_secondary else "rwm_expected_pct"
candidates = frame[frame["urn"] != urn].copy()
for column in ("latitude", "longitude"):
candidates = candidates[candidates[column].notna()]
if candidates.empty:
return []
# ── Hard filters ────────────────────────────────────────────────────
allowed_phases = _phase_group(is_secondary)
candidates = candidates[
candidates["phase"].fillna("").str.lower().isin(allowed_phases)
]
candidates = candidates[candidates["status"].fillna("").str.lower().str.startswith("open")]
subject_special = is_special_provision(subject.get("school_type"))
special = _mask(candidates["school_type"], is_special_provision)
candidates = candidates[special if subject_special else ~special]
subject_selective = is_selective(subject.get("admissions_policy"))
selective = _mask(candidates["admissions_policy"], is_selective)
candidates = candidates[selective if subject_selective else ~selective]
subject_gender = subject.get("gender")
candidates = candidates[
_mask(candidates["gender"], lambda g: genders_compatible(subject_gender, g))
]
if candidates.empty:
return []
candidates["distance_miles"] = _haversine_miles(
lat, lon, candidates["latitude"].values, candidates["longitude"].values
).round(1)
# ── Nearest first, and nothing else has a say ───────────────────────
within = candidates[candidates["distance_miles"] <= reach]
if len(within) < MINIMUM:
return []
selected = within.sort_values(["distance_miles", "urn"]).head(MAX_SCHOOLS)
return [
{
"urn": int(row["urn"]),
"school_name": str(row.get("school_name") or ""),
"distance_miles": float(row["distance_miles"]),
"school_type": _native(row.get("school_type")),
"age_range": _native(row.get("age_range")),
"shared": _shared(subject, row, is_secondary),
"metric_value": _native(row.get(metric_key)),
"metric_key": metric_key,
"metric_year": _native(row.get("year")),
}
for _, row in selected.iterrows()
]
+356
View File
@@ -0,0 +1,356 @@
"""The place registry: what places the site publishes, and what is in each.
One module owns this question. The pages, the sitemap and the internal-link
modules all read from here, so the threshold and the collision rules exist in
exactly one place and are testable without a browser or a database.
Two namespaces, never one. 67 viable town names collide with a local
authority name, and the authority is the larger set in only 43 of them —
postal towns cross authority boundaries, so neither can absorb the other.
Keys are "<kind>:<slug>" so the collision cannot reappear in the dict.
"""
from __future__ import annotations
import logging
import re
from dataclasses import dataclass, field
logger = logging.getLogger(__name__)
# Five schools with publishable data. Below this a place has nothing to say
# that a list of schools does not, and publishing it is index bloat.
MIN_SCHOOLS = 5
@dataclass(frozen=True)
class Place:
kind: str # "town" | "locality" | "authority" | "outcode"
slug: str
name: str
urns: tuple[int, ...]
parent_authority: str | None # authority NAME, for the 301 target
# Every authority the place meaningfully sits in, largest first. A quarter
# of outcodes and a third of towns straddle a boundary — SW19 is mostly
# Merton but partly Wandsworth — so naming only one asserts something
# false. parent_authority stays single because a redirect needs one
# target; this is what the page shows.
authorities: tuple[tuple[str, int], ...] = ()
# URNs per phase, so the per-phase threshold can be applied without
# re-querying. A place with 30 primaries and 2 secondaries publishes a
# primary variant and no secondary one.
phase_urns: dict[str, tuple[int, ...]] = field(default_factory=dict)
def publishes_phase(self, phase: str) -> bool:
return len(self.phase_urns.get(phase, ())) >= MIN_SCHOOLS
@property
def key(self) -> str:
return f"{self.kind}:{self.slug}"
def _publishable_urns(df) -> set[int]:
"""URNs with something a page could state, deduplicated across years."""
from backend.app import _PUBLISHABLE_FIELDS
cols = [c for c in _PUBLISHABLE_FIELDS if c in df.columns]
if not cols:
return set()
return set(df.loc[df[cols].notna().any(axis=1), "urn"].astype(int))
# The measure a phase page is built around. A page with no results in this
# column has nothing a list of school names does not already give.
_PHASE_METRIC = {
"primary": "rwm_expected_pct",
"secondary": "attainment_8_score",
}
def _phase_urns(group, publishable: set[int]) -> dict[str, tuple[int, ...]]:
"""URNs per phase, counting only schools with a result for that phase.
Not merely "publishable". A school with an Ofsted grade and no results is
worth a page of its own and belongs in the place list, but it cannot
populate a phase page's results column — and the threshold is there to ask
whether that column will have anything in it.
Counting publishable schools instead let /schools/kent/primary publish
with none of its five rows carrying a result, and left 44 phase pages
majority-blank. It is the same rule as "no page without a local average",
which was never extended per phase.
All-through schools count toward both phases, matching the PHASE_GROUPS
mapping the search filters already use.
"""
from backend.app import PHASE_GROUPS
if "phase" not in group.columns:
return {}
lowered = group["phase"].fillna("").str.lower()
out: dict[str, tuple[int, ...]] = {}
for phase in ("primary", "secondary"):
wanted = PHASE_GROUPS.get(phase, set())
subset = group[lowered.isin(wanted)]
# The page lists every school of the phase; the threshold counts only
# those carrying a result, so a mostly-empty table never publishes.
metric = _PHASE_METRIC[phase]
with_result = (
{int(u) for u in subset.loc[subset[metric].notna(), "urn"]}
if metric in subset.columns else set()
)
if len(with_result & publishable) < MIN_SCHOOLS:
continue
urns = tuple(sorted({int(u) for u in subset["urn"]} & publishable))
if urns:
out[phase] = urns
return out
# A place is described by an authority when it holds at least a tenth of the
# schools, and at least two. GIAS carries occasional postcode errors — EN6
# lists two Shropshire schools among fourteen in Hertfordshire — and a bare
# "any authority present" rule would print those as though they were real.
# There is deliberately no cap on how many are named. An earlier cut stopped
# at three, which silently dropped the fourth in exactly the case where the
# information matters most — a genuinely fragmented place. The share rule is
# the only limit, and it already bounds the list at ten.
_AUTHORITY_MIN_SHARE = 0.10
_AUTHORITY_MIN_SCHOOLS = 2
def _authorities(group) -> tuple[tuple[str, int], ...]:
"""Authorities this place meaningfully sits in, largest first."""
from backend.app import EXCLUDED_FILTER_VALUES
if "local_authority" not in group.columns:
return ()
counts = group["local_authority"].dropna().value_counts()
total = int(counts.sum())
if not total:
return ()
kept = [
(str(name), int(n)) for name, n in counts.items()
if str(name) not in EXCLUDED_FILTER_VALUES
and n >= _AUTHORITY_MIN_SCHOOLS
and n / total >= _AUTHORITY_MIN_SHARE
]
# A place too small or too fragmented for the share rule still names its
# largest authority, or the page would say nothing about where it is.
if not kept:
for name, n in counts.items():
if str(name) not in EXCLUDED_FILTER_VALUES:
return ((str(name), int(n)),)
return ()
return tuple(kept)
def _parent_authority(authorities: tuple[tuple[str, int], ...]) -> str | None:
"""The 301 target: the largest authority a place sits in.
Derived from `authorities` rather than computed separately. The first cut
used `mode()` here while `authorities` used `value_counts()`, and on an
exact tie pandas does not guarantee the two pick the same name — so the
redirect could have pointed somewhere other than the authority the page
named first. One computation, one answer.
Deriving it also inherits the sentinel filter, so a place can no longer
redirect to /schools/authority/does-not-apply.
"""
return authorities[0][0] if authorities else None
def _group(df, column: str, kind: str, publishable: set[int]) -> dict[str, Place]:
"""One Place per distinct SLUG in `column` that clears the threshold.
Grouped by slug, not by raw value, because GIAS spells the same place
several ways and they all resolve to one URL. Five town slugs come from
more than one spelling: "London" (1,819 schools) and "LONDON" (12) both
slugify to `london`; Weston-super-Mare is split 14/19 across two
spellings; Newcastle-under-Lyme across three.
Grouping by raw value meant the later group simply overwrote the earlier
one in this dict — so /schools/london could have shown twelve schools
instead of 1,819, silently and depending on row order.
The display name is the most common spelling, which is the one a reader
expects to see.
"""
from backend.app import _slugify
if column not in df.columns:
return {}
working = df.assign(_slug=df[column].map(
lambda v: _slugify(str(v).strip()) if isinstance(v, str) and v.strip() else None))
working = working[working["_slug"].notna() & (working["_slug"] != "")]
out: dict[str, Place] = {}
for slug, group in working.groupby("_slug"):
slug = str(slug)
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
continue
spellings = group[column].dropna().value_counts()
if spellings.empty:
continue
name = str(spellings.index[0]).strip()
authorities = () if kind == "authority" else _authorities(group)
place = Place(
kind=kind, slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable),
)
out[place.key] = place
return out
# "SW11 2AA" -> "SW11". Two letters max, one or two digits, optional letter.
_OUTCODE_RE = re.compile(r"^([A-Z]{1,2}\d{1,2}[A-Z]?)\s")
def _outcode(postcode) -> str | None:
if not isinstance(postcode, str):
return None
m = _OUTCODE_RE.match(postcode.upper().strip())
return m.group(1) if m else None
def _outcode_places(df, publishable: set[int]) -> dict[str, Place]:
"""One Place per postcode district clearing the threshold.
These carry no phase variants: nobody searches "primary schools in SW11",
so the spec gives them no /primary or /secondary route. `phase_urns` is
left empty rather than computed and then filtered downstream — the
registry is the one place that decides which phases a place publishes,
and the page links whatever it reports.
Computing them here put a link to a route that does not exist on every one
of the 1,720 outcode pages.
"""
if "postcode" not in df.columns:
return {}
working = df.assign(_oc=df["postcode"].map(_outcode))
working = working[working["_oc"].notna()]
out: dict[str, Place] = {}
for oc, group in working.groupby("_oc"):
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
continue
authorities = _authorities(group)
place = Place(kind="outcode", slug=str(oc).lower(), name=str(oc),
urns=urns, parent_authority=_parent_authority(authorities),
authorities=authorities)
out[place.key] = place
return out
def _locality_places(df, publishable: set[int],
town_slugs: set[str]) -> dict[str, Place]:
"""One Place per curated locality clearing the threshold."""
from backend.localities import LOCALITY_OUTCODES
if "postcode" not in df.columns:
return {}
working = df.assign(_oc=df["postcode"].map(_outcode))
out: dict[str, Place] = {}
for slug, (name, outcodes) in LOCALITY_OUTCODES.items():
if slug in town_slugs:
# Skip, do not raise. The guard exists so a locality never
# silently shadows a town — skipping achieves that, and the error
# log makes it loud.
#
# Raising here took down sitemap generation for all 25,000 school
# pages when "richmond" met the GIAS town Richmond in North
# Yorkshire. Worse, GIAS town names change without any code change,
# so a raise means curated data can break the site spontaneously.
# A curation mistake must cost one page, not the sitemap.
logger.error(
"locality %r collides with the published town of the same "
"slug and has been skipped; rename it or remove it", slug)
continue
group = working[working["_oc"].isin(outcodes)]
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
# Not an error — a locality can legitimately be too small. Logged
# because one you meant to publish quietly vanishing is the
# failure worth hearing about.
logger.warning(
"locality %s (%s) has %d publishable schools, below the "
"threshold of %d - not published",
slug, ", ".join(outcodes), len(urns), MIN_SCHOOLS)
continue
authorities = _authorities(group)
place = Place(kind="locality", slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable))
out[place.key] = place
return out
# Ordered authority → town/locality → outcode, widest first, because that is
# the order a breadcrumb reads. The link module re-sorts for its own purposes.
_PLACE_ORDER = {"authority": 0, "town": 1, "locality": 2, "outcode": 3}
def build_place_index(registry: dict[str, Place]) -> dict[int, tuple[Place, ...]]:
"""URN → the published places containing it, built once per registry.
The reverse of the registry, and the thing school pages link out through.
Derived from the registry rather than maintained beside it, so the two
cannot disagree about which places exist: a place below the publish
threshold is absent from the registry, so it is absent from here too, and
a link is never offered for a page that does not exist.
Built as an index rather than scanned per call because /api/schools/{urn}
is the site's highest-traffic endpoint. Scanning meant walking every place
and doing a tuple membership test against each — on the order of 10^5
comparisons per request, repeated for every school page view. One pass at
registry-build time replaces all of it with a dict lookup.
"""
grouped: dict[int, list[Place]] = {}
for place in registry.values():
for urn in place.urns:
grouped.setdefault(int(urn), []).append(place)
return {
urn: tuple(sorted(places,
key=lambda p: (_PLACE_ORDER.get(p.kind, 9), p.slug)))
for urn, places in grouped.items()
}
def places_for_urn(index: dict[int, tuple[Place, ...]], urn: int) -> tuple[Place, ...]:
"""The published places containing this school, widest first.
Empty is a real answer, not a failure: a school whose town and authority
both fall below the publish threshold has nowhere to link, and the page
renders without the module.
"""
return index.get(int(urn), ())
def build_place_registry(df) -> dict[str, Place]:
"""Every place the site publishes, keyed by "<kind>:<slug>"."""
if df.empty or "urn" not in df.columns:
return {}
publishable = _publishable_urns(df)
registry: dict[str, Place] = {}
registry.update(_group(df, "local_authority", "authority", publishable))
towns = _group(df, "town", "town", publishable)
registry.update(towns)
town_slugs = {p.slug for p in towns.values()}
registry.update(_locality_places(df, publishable, town_slugs))
registry.update(_outcode_places(df, publishable))
return registry
+12
View File
@@ -532,6 +532,18 @@ RANKING_COLUMNS = [
"gcse_grade_91_pct",
]
# Maps user-facing phase filter values to the GIAS PhaseOfEducation values they
# include. All-through schools appear in both primary and secondary results,
# which is why this is a set per phase rather than a single string comparison.
#
# Lives here rather than in app.py because nearby_schools.py needs it too, and
# importing app from there would be a cycle.
PHASE_GROUPS: dict[str, set[str]] = {
"primary": {"primary", "middle deemed primary", "all-through"},
"secondary": {"secondary", "middle deemed secondary", "all-through", "16 plus"},
"all-through": {"all-through"},
}
# School listing columns
SCHOOL_COLUMNS = [
"urn",
+269
View File
@@ -0,0 +1,269 @@
"""The destinations serialiser's contract.
Not rendering a figure is not the same as not publishing it. This endpoint is
public and unauthenticated, so whatever the payload carries is published,
whatever the UI draws. The categories sum to the cohort and the pupil groups
sum to each other, so a lone suppressed cell is solvable by subtraction — the
serialiser adds secondary suppression to prevent it.
See docs/superpowers/specs/2026-08-28-destination-measures-design.md.
"""
from backend.data_loader import (
_destinations_block, _format_cohort_year, disclosure_invariant_holds,
)
def _row(group, measure, pupils, status, cohort=180, percentage=None, year=202223):
return {
"pupil_group": group,
"destination_measure": measure,
"pupils": pupils,
"percentage": percentage,
"status": status,
"cohort_pupils": cohort,
"year": year,
}
def test_suppressed_category_serialises_as_suppressed_with_null_pupils():
rows = [
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
_row("all", "sixth_form_college", None, "suppressed"),
]
block = _destinations_block(rows)
cats = {c["category"]: c for c in block["groups"]["all"]["categories"]}
assert cats["sixth_form_college"]["status"] == "suppressed"
assert cats["sixth_form_college"]["pupils"] is None
assert cats["sixth_form_college"]["percentage"] is None
def test_published_category_keeps_its_figures():
block = _destinations_block([
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
])
cat = block["groups"]["all"]["categories"][0]
assert cat["pupils"] == 75
assert cat["percentage"] == 41.7
assert cat["status"] == "published"
def test_only_the_latest_year_is_served():
rows = [
_row("all", "school_sixth_form", 60, "published", year=202122),
_row("all", "school_sixth_form", 75, "published", year=202223),
]
block = _destinations_block(rows)
assert block["cohort_year"] == "2022/23"
assert len(block["groups"]["all"]["categories"]) == 1
assert block["groups"]["all"]["categories"][0]["pupils"] == 75
def test_all_three_pupil_groups_are_carried():
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("disadvantaged", "school_sixth_form", 17, "published", cohort=62),
_row("other", "school_sixth_form", 58, "published", cohort=118),
]
block = _destinations_block(rows)
assert set(block["groups"]) == {"all", "disadvantaged", "other"}
assert block["groups"]["disadvantaged"]["cohort"] == 62
def test_cohort_year_is_reported_so_the_page_can_date_itself():
block = _destinations_block([_row("all", "school_sixth_form", 75, "published")])
assert block["cohort_year"] == "2022/23"
def test_format_cohort_year_handles_the_six_digit_form():
assert _format_cohort_year(202223) == "2022/23"
assert _format_cohort_year(None) is None
def test_empty_rows_yield_none_not_an_empty_shell():
assert _destinations_block([]) is None
# ── Disclosure control ──────────────────────────────────────────────────────
#
# The rendering guards in lib/destinations.ts stop a withheld figure being
# DRAWN. They do nothing about it being COMPUTED: this endpoint is public and
# unauthenticated, so whatever the payload carries is published. These tests
# are the ones that matter.
def _solve_residual(group):
"""What any caller can work out: cohort minus everything published."""
published = [c["pupils"] for c in group["categories"] if c["pupils"] is not None]
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
return group["cohort"] - sum(published), len(hidden)
def test_a_lone_suppressed_category_cannot_be_solved_for():
"""Whitley Bay High School's real 2022/23 disadvantaged group: further
education withheld, everything else published, cohort 41. Before secondary
suppression the payload gave the answer away as 41 - 23 = 18."""
rows = [
_row("disadvantaged", "school_sixth_form", 15, "published", cohort=41),
_row("disadvantaged", "sixth_form_college", 0, "published", cohort=41),
_row("disadvantaged", "further_education", None, "suppressed", cohort=41),
_row("disadvantaged", "apprenticeship", 1, "published", cohort=41),
_row("disadvantaged", "employment", 2, "published", cohort=41),
_row("disadvantaged", "not_sustained", 3, "published", cohort=41),
_row("disadvantaged", "not_captured", 2, "published", cohort=41),
]
group = _destinations_block(rows)["groups"]["disadvantaged"]
residual, hidden = _solve_residual(group)
assert hidden >= 2, "a lone suppressed cell must gain a companion"
assert residual != 18, "the withheld figure is recoverable from the payload"
def test_every_group_hides_none_or_at_least_two_categories():
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("all", "sixth_form_college", None, "suppressed"),
_row("all", "further_education", 61, "published"),
_row("all", "apprenticeship", 8, "published"),
_row("all", "employment", 6, "published"),
_row("all", "not_sustained", 5, "published"),
_row("all", "not_captured", 4, "published"),
]
group = _destinations_block(rows)["groups"]["all"]
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
assert len(hidden) >= 2
def test_a_category_hidden_in_one_group_is_hidden_in_a_second():
"""disadvantaged + other = all for every category, so a category withheld
in exactly one of the three is recoverable from the other two."""
rows = []
for measure, a, d, o in [
("school_sixth_form", 75, None, 58),
("further_education", 61, 27, 34),
("apprenticeship", 8, 4, 4),
("employment", 6, 1, 5),
("not_sustained", 5, 3, 2),
("not_captured", 4, 2, 2),
]:
rows.append(_row("all", measure, a, "published", cohort=159))
rows.append(_row("disadvantaged", measure, d,
"published" if d is not None else "suppressed", cohort=37))
rows.append(_row("other", measure, o, "published", cohort=122))
groups = _destinations_block(rows)["groups"]
measures = {c["category"] for g in groups.values() for c in g["categories"]}
assert len(measures) == 6, "the fixture's six measures must all be checked"
for measure in sorted(measures):
hidden = sum(
1 for g in groups.values() for c in g["categories"]
if c["category"] == measure and c["status"] == "suppressed"
)
# The invariant is "none, or at least two" — not "at least two".
assert hidden != 1, f"{measure} is solvable across the pupil groups"
def test_a_suppressed_cell_never_keeps_its_percentage():
"""percentage / pupils would hand back the cohort, and with it the residual."""
rows = [
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
_row("all", "sixth_form_college", None, "suppressed", percentage=11.7),
_row("all", "further_education", 61, "published", percentage=33.9),
]
group = _destinations_block(rows)["groups"]["all"]
for cell in group["categories"]:
if cell["status"] != "published":
assert cell["pupils"] is None
assert cell["percentage"] is None
def test_aggregates_are_not_served():
"""An aggregate spanning exactly one suppressed component names it, and
nothing renders them today."""
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("all", "agg_sustained_all", 171, "published"),
]
group = _destinations_block(rows)["groups"]["all"]
assert [c["category"] for c in group["categories"]] == ["school_sixth_form"]
assert "aggregates" not in group
def test_a_fully_published_group_is_left_alone():
"""Secondary suppression must not cost anything where nothing is withheld —
this is the all-pupils view on every mainstream secondary."""
rows = [
_row("all", m, p, "published")
for m, p in [("school_sixth_form", 75), ("sixth_form_college", 21),
("further_education", 61), ("apprenticeship", 8),
("employment", 6), ("not_sustained", 5), ("not_captured", 4)]
]
group = _destinations_block(rows)["groups"]["all"]
assert all(c["status"] == "published" for c in group["categories"])
assert len(group["categories"]) == 7
def test_the_invariant_is_asserted_directly_not_re_derived():
"""A group with one suppressed category and nothing else to withhold."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=9),
_row("all", "sixth_form_college", None, "not_applicable", cohort=9),
_row("all", "further_education", None, "not_applicable", cohort=9),
]
block = _destinations_block(rows)
assert block is None or disclosure_invariant_holds(block["groups"])
def test_a_sparse_cohort_with_no_companion_drops_the_group():
"""Special schools and AP routinely have one suppressed category and every
other one not applicable. There is nothing left to withhold, so the group
goes — an earlier version returned here with the violation intact."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=9),
_row("all", "sixth_form_college", None, "not_applicable", cohort=9),
_row("all", "further_education", None, "not_applicable", cohort=9),
_row("all", "apprenticeship", None, "not_applicable", cohort=9),
_row("all", "employment", None, "not_applicable", cohort=9),
_row("all", "not_sustained", None, "not_applicable", cohort=9),
_row("all", "not_captured", None, "not_applicable", cohort=9),
]
block = _destinations_block(rows)
assert block is None or "all" not in block["groups"], (
"a group that cannot be made safe must not be served"
)
def test_zeros_are_not_treated_as_a_usable_companion():
"""Suppressing a zero protects nothing — the residual is unchanged. With
only zeros available the group must be dropped, not falsely 'fixed'."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=5),
_row("all", "sixth_form_college", 0, "published", cohort=5),
_row("all", "further_education", 0, "published", cohort=5),
]
block = _destinations_block(rows)
if block and "all" in block["groups"]:
group = block["groups"]["all"]
published = sum(c["pupils"] for c in group["categories"]
if c["pupils"] is not None)
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
assert len(hidden) != 1, "a zero companion leaves the figure solvable"
assert group["cohort"] - published != 5
def test_masking_always_terminates_in_a_safe_state():
"""Exhaustive over every suppression pattern of a four-category group."""
from itertools import product
MEASURES = ["school_sixth_form", "sixth_form_college",
"further_education", "apprenticeship"]
for statuses in product(["published", "suppressed", "not_applicable"],
repeat=len(MEASURES)):
rows = [
_row("all", m, 3 if st == "published" else None, st, cohort=12)
for m, st in zip(MEASURES, statuses)
]
block = _destinations_block(rows)
if block is None:
continue
assert disclosure_invariant_holds(block["groups"]), (
f"invariant broken for {statuses}"
)
+163
View File
@@ -0,0 +1,163 @@
"""Tests for the feature flag layer (spec 2026-08-23).
None of these need a running Unleash. That is the point: an unset UNLEASH_URL
means every flag is False, which is what local development and CI get.
"""
from datetime import date, timedelta
from backend import flags
def test_every_declared_flag_is_keyed_by_its_own_name():
# One string is the registry key, the Unleash flag name and the JSON key.
# A mismatch here would mean the UI toggles a flag the code never reads.
for key, flag in flags.REGISTRY.items():
assert key == flag.name
def test_flag_names_are_snake_case():
# Matches the API's existing convention (admission_distance,
# rwm_expected_pct) so no case transformation exists to get wrong.
for name in flags.REGISTRY:
assert name == name.lower()
assert "-" not in name and " " not in name
def test_an_unconfigured_client_evaluates_every_flag_false(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
for name in flags.REGISTRY:
assert flags.is_enabled(name) is False
def test_an_undeclared_flag_is_false_rather_than_an_error(monkeypatch):
# A typo'd flag name must not raise in a request path. It is logged as an
# error, because an undeclared flag is always a bug.
monkeypatch.setattr(flags, "_client", None)
assert flags.is_enabled("no_such_flag") is False
def test_an_exploding_client_is_false_rather_than_a_500(monkeypatch):
class Boom:
def is_enabled(self, *a, **kw):
raise RuntimeError("unleash is on fire")
monkeypatch.setattr(flags, "_client", Boom())
name = next(iter(flags.REGISTRY))
assert flags.is_enabled(name) is False
def test_all_flags_reports_every_declared_flag(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
assert set(flags.all_flags()) == set(flags.REGISTRY)
assert all(v is False for v in flags.all_flags().values())
def test_a_flag_older_than_the_limit_fails_this_test():
"""A tripwire, not an assertion about correctness.
Flags are temporary scaffolding and the failure mode of every flag system
is accumulation. This fails on the day a flag turns 90, on whatever PR
happens to be open — which is the point: someone has to decide.
To fix: delete the flag and the branches that read it, or, if it genuinely
still needs to exist, move its `added` date and say why in the commit.
"""
stale = [
f.name for f in flags.REGISTRY.values()
if date.today() - f.added > timedelta(days=flags.MAX_FLAG_AGE_DAYS)
]
assert not stale, (
f"Flags older than {flags.MAX_FLAG_AGE_DAYS} days: {stale}. "
"Remove the flag and the code branches it guards, or move its `added` "
"date deliberately."
)
def _client():
from fastapi.testclient import TestClient
from backend import app as app_module
return TestClient(app_module.app, raise_server_exceptions=False)
def test_the_flags_endpoint_lists_every_declared_flag(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
body = _client().get("/api/flags").json()
assert set(body) == set(flags.REGISTRY)
def test_the_flags_endpoint_answers_false_when_unleash_is_unreachable(monkeypatch):
# The endpoint must still answer. A frontend that cannot read flags renders
# everything dark, which is right; one that gets a 500 renders nothing.
monkeypatch.setattr(flags, "_client", None)
res = _client().get("/api/flags")
assert res.status_code == 200
assert all(v is False for v in res.json().values())
def _school_payload(monkeypatch, *, flag_on: bool):
"""Fetch one school's payload with the distance flag forced on or off.
The DataFrame shape is copied from test_school_details.py rather than
minimised: the endpoint reads a wide set of GIAS columns, and a trimmed
frame fails for reasons that have nothing to do with flags.
"""
import numpy as np
import pandas as pd
from fastapi.testclient import TestClient
from backend import app as app_module
df = pd.DataFrame([{
"urn": 150275,
"school_name": "West London Performing Arts Academy",
"phase": "Secondary",
"school_type": "Special post 16 institution",
"trust_name": None,
"religious_denomination": "Does not apply",
"gender": None,
"age_range": "16-25",
"admissions_policy": None,
"capacity": np.nan,
"gias_total_pupils": np.nan,
"headteacher_name": None,
"website": None,
"ofsted_grade": np.nan,
"local_authority": "Ealing",
"address": "268 Northfield Avenue, London, W5 4UB",
"postcode": "W5 4UB",
"latitude": 51.4986,
"longitude": -0.3148,
"year": np.nan,
"total_pupils": np.nan,
"eligible_pupils": np.nan,
"rwm_expected_pct": np.nan,
}])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
# Two arguments: get_supplementary_data(db, urn). See backend/app.py.
monkeypatch.setattr(
app_module, "get_supplementary_data",
lambda db, urn: {"admission_distance": {"distance_m": 772.49,
"year": 2024}})
monkeypatch.setattr(flags, "is_enabled", lambda name: flag_on)
client = TestClient(app_module.app, raise_server_exceptions=False)
res = client.get("/api/schools/150275")
assert res.status_code == 200, res.text
return res.json()
def test_the_distance_field_is_absent_when_the_flag_is_off(monkeypatch):
"""Absent, not null, and withheld at the source.
/api/schools/ is public and unauthenticated. Leaving a withheld field in
the payload while declining to render it hands the record to anyone who
opens the network tab — the reasoning already recorded in c9a1892.
"""
body = _school_payload(monkeypatch, flag_on=False)
assert "admission_distance" not in body
def test_the_distance_field_is_present_when_the_flag_is_on(monkeypatch):
body = _school_payload(monkeypatch, flag_on=True)
assert body["admission_distance"]["distance_m"] == 772.49
+383
View File
@@ -0,0 +1,383 @@
"""Selection rules for the nearby-schools section.
Hard filters encode claims the section is not allowed to make — that a
selective school is an alternative to a non-selective one, that a special
school is comparable to a mainstream one, or that a Girls school is an option
for a Boys school's reader. They decide who is eligible.
Distance decides the order, and nothing else does. An earlier version ranked by
intake similarity first, which put a Catholic school 2.9 miles away above the
community school 0.3 miles down the road — for a primary, a school that far is
not a weaker option, it is not an option. Similarity is now reported on the
card and never reorders the row.
"""
import numpy as np
import pandas as pd
from backend.nearby_schools import (
is_secondary_phase,
radius_miles,
select_nearby,
)
BASE_LAT, BASE_LON = 51.5000, -0.1000
def _row(urn, name, **overrides):
base = {
"urn": urn,
"school_name": name,
"local_authority": "Testshire",
"school_type": "Community school",
"phase": "Primary",
"age_range": "4-11",
"status": "Open",
"gender": "Mixed",
"religious_denomination": "None",
"admissions_policy": "Not applicable",
"latitude": BASE_LAT,
"longitude": BASE_LON,
"year": 202425,
"rwm_expected_pct": 70.0,
"attainment_8_score": np.nan,
}
base.update(overrides)
return base
def _frame(*rows):
return pd.DataFrame(list(rows))
def _at(miles):
"""A latitude `miles` north of BASE_LAT."""
return BASE_LAT + miles / 69.0
# ---------------------------------------------------------------------------
# Order: distance, and only distance
# ---------------------------------------------------------------------------
def test_returns_nearest_first():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Mid", latitude=_at(1.0)),
_row(100003, "Near", latitude=_at(0.4)),
_row(100004, "Far", latitude=_at(1.8)),
)
result = select_nearby(frame, 100001)
assert [s["urn"] for s in result] == [100003, 100002, 100004]
assert result[0]["distance_miles"] == 0.4
def test_a_faith_match_never_outranks_a_closer_school():
"""The reported defect. A Catholic primary surrounded by Catholic primaries
showed six of them and omitted the community school down the road."""
frame = _frame(
_row(100001, "St Jude's RC Primary", religious_denomination="Roman Catholic"),
_row(100002, "Elm Grove Primary", religious_denomination="None", latitude=_at(0.3)),
_row(100003, "Holy Cross RC", religious_denomination="Roman Catholic", latitude=_at(0.8)),
_row(100004, "Sacred Heart RC", religious_denomination="Roman Catholic", latitude=_at(1.2)),
_row(100005, "St Peter's RC", religious_denomination="Roman Catholic", latitude=_at(1.6)),
)
result = select_nearby(frame, 100001)
assert result[0]["urn"] == 100002, "the nearest school leads, whatever its intake"
assert [s["distance_miles"] for s in result] == sorted(s["distance_miles"] for s in result)
def test_the_nearest_eligible_school_is_always_shown():
"""Whatever else changes, a section titled "nearby" cannot omit the nearest
school while listing one four times further away."""
frame = _frame(
_row(100001, "Subject", gender="Boys", religious_denomination="Roman Catholic"),
_row(100002, "Nearest", gender="Mixed", religious_denomination="None", latitude=_at(0.2)),
*[
_row(100010 + n, f"Match {n}", gender="Boys",
religious_denomination="Roman Catholic", latitude=_at(0.9 + n * 0.1))
for n in range(6)
],
)
assert select_nearby(frame, 100001)[0]["urn"] == 100002
def test_caps_at_six_taking_the_nearest():
frame = _frame(
_row(100001, "Subject"),
*[_row(100010 + n, f"Peer {n}", latitude=_at(0.1 * (n + 1))) for n in range(7)],
)
result = select_nearby(frame, 100001)
assert len(result) == 6
assert 100016 not in {s["urn"] for s in result}, "the seventh-nearest is the one dropped"
def test_fewer_than_two_matches_returns_empty():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Only neighbour", latitude=_at(0.5)),
)
assert select_nearby(frame, 100001) == []
def test_excludes_the_subject_school():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
assert 100001 not in {s["urn"] for s in select_nearby(frame, 100001)}
def test_a_school_is_never_listed_twice():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
result = select_nearby(frame, 100001)
assert len(result) == len({s["urn"] for s in result})
# ---------------------------------------------------------------------------
# Reach: a sanity bound, not a target
# ---------------------------------------------------------------------------
def test_primary_does_not_reach_past_two_miles():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Just inside", latitude=_at(1.9)),
_row(100003, "Just outside", latitude=_at(2.4)),
_row(100004, "Miles away", latitude=_at(4.0)),
)
# One inside the cap is below the minimum, so nothing renders at all —
# a primary with nothing within two miles has no nearby schools.
assert select_nearby(frame, 100001) == []
def test_secondary_reaches_further_than_primary():
frame = _frame(
_row(100001, "Subject", phase="Secondary"),
_row(100002, "A", phase="Secondary", latitude=_at(3.0)),
_row(100003, "B", phase="Secondary", latitude=_at(5.5)),
)
assert {s["urn"] for s in select_nearby(frame, 100001)} == {100002, 100003}
def test_the_cap_follows_the_phase():
assert radius_miles("Primary") == 2.0
assert radius_miles("Middle deemed primary") == 2.0
assert radius_miles("All-through") == 2.0
assert radius_miles("Secondary") == 6.0
assert radius_miles("Middle deemed secondary") == 6.0
# Post-16 is the phase people travel furthest for.
assert radius_miles("16 plus") == 10.0
# ---------------------------------------------------------------------------
# Hard filters: eligibility, never order
# ---------------------------------------------------------------------------
def test_selective_never_meets_non_selective():
frame = _frame(
_row(100001, "Grammar", phase="Secondary", admissions_policy="Selective"),
_row(100002, "Comp A", phase="Secondary", admissions_policy="Non-selective", latitude=_at(0.5)),
_row(100003, "Comp B", phase="Secondary", admissions_policy="Non-selective", latitude=_at(0.6)),
)
assert select_nearby(frame, 100001) == []
assert 100001 not in {s["urn"] for s in select_nearby(frame, 100002)}
def test_special_schools_match_only_each_other():
frame = _frame(
_row(100001, "Special", school_type="Community special school"),
_row(100002, "Mainstream A", latitude=_at(0.5)),
_row(100003, "Mainstream B", latitude=_at(0.6)),
)
assert select_nearby(frame, 100001) == []
assert select_nearby(frame, 100002) == []
def test_boys_never_meets_girls():
frame = _frame(
_row(100001, "Boys School", gender="Boys"),
_row(100002, "Girls School", gender="Girls", latitude=_at(0.5)),
_row(100003, "Mixed School", gender="Mixed", latitude=_at(0.6)),
_row(100004, "Another Mixed", gender="Mixed", latitude=_at(0.7)),
)
urns = {s["urn"] for s in select_nearby(frame, 100001)}
assert 100002 not in urns
assert urns == {100003, 100004}
def test_closed_schools_and_missing_coordinates_are_dropped():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Closed", status="Closed", latitude=_at(0.5)),
_row(100003, "No coords", latitude=np.nan, longitude=np.nan),
_row(100004, "Good A", latitude=_at(0.6)),
_row(100005, "Good B", latitude=_at(0.7)),
)
assert {s["urn"] for s in select_nearby(frame, 100001)} == {100004, 100005}
def test_all_through_is_offered_on_both_phase_sides():
frame = _frame(
_row(100001, "Primary subject", phase="Primary"),
_row(100002, "All through", phase="All-through", latitude=_at(0.5)),
_row(100003, "Primary peer", phase="Primary", latitude=_at(0.6)),
)
assert 100002 in {s["urn"] for s in select_nearby(frame, 100001)}
secondary = _frame(
_row(100010, "Secondary subject", phase="Secondary"),
_row(100002, "All through", phase="All-through", latitude=_at(0.5)),
_row(100011, "Secondary peer", phase="Secondary", latitude=_at(0.6)),
)
assert 100002 in {s["urn"] for s in select_nearby(secondary, 100010)}
def test_sixteen_plus_is_matched_against_secondary_not_primary():
"""GIAS phase 6 is "16 plus", and PHASE_GROUPS puts it in the secondary
group — a sixth-form college's peers are secondaries and other colleges,
never primary schools. A substring test for "secondary" misses it silently:
no crash, just a page offering infant schools to a sixth form."""
frame = _frame(
_row(100001, "Sixth Form College", phase="16 plus", age_range="16-19"),
_row(100002, "Nearby Secondary", phase="Secondary", latitude=_at(0.5),
attainment_8_score=52.0),
_row(100003, "Nearby College", phase="16 plus", latitude=_at(0.6)),
_row(100004, "Nearby Primary", phase="Primary", latitude=_at(0.1)),
)
result = select_nearby(frame, 100001)
urns = {s["urn"] for s in result}
assert 100004 not in urns, "a primary school is not a peer for a sixth form"
assert urns == {100002, 100003}
assert all(s["metric_key"] == "attainment_8_score" for s in result)
def test_is_secondary_phase_agrees_with_the_phase_groups_it_selects_from():
for phase in ("Secondary", "Middle deemed secondary", "16 plus"):
assert is_secondary_phase(phase) is True, phase
for phase in ("Primary", "Middle deemed primary", "Nursery", "", None):
assert is_secondary_phase(phase) is False, phase
# In PHASE_GROUPS an all-through school is on both sides, but it renders
# with the primary template, and the metric follows the phase side.
assert is_secondary_phase("All-through") is False
# ---------------------------------------------------------------------------
# What the card reports
# ---------------------------------------------------------------------------
def test_shared_lists_only_what_is_actually_shared():
frame = _frame(
_row(100001, "Subject", phase="Secondary", gender="Mixed",
religious_denomination="None", admissions_policy="Non-selective"),
_row(100002, "Full match", phase="Secondary", gender="Mixed",
religious_denomination="None", admissions_policy="Non-selective", latitude=_at(0.5)),
_row(100003, "Faith differs", phase="Secondary", gender="Mixed",
religious_denomination="Church of England", admissions_policy="Non-selective", latitude=_at(0.6)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert by_urn[100002]["shared"] == ["Mixed", "Non-selective", "No religious character"]
assert by_urn[100003]["shared"] == ["Mixed", "Non-selective"]
def test_a_shared_faith_is_named():
frame = _frame(
_row(100001, "Subject", religious_denomination="Roman Catholic"),
_row(100002, "Also RC", religious_denomination="Roman Catholic", latitude=_at(0.4)),
_row(100003, "Secular", religious_denomination="None", latitude=_at(0.5)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert "Roman Catholic" in by_urn[100002]["shared"]
assert by_urn[100003]["shared"] == ["Mixed"]
def test_shared_is_empty_when_nothing_is_shared():
frame = _frame(
_row(100001, "Subject", gender="Boys", religious_denomination="Roman Catholic"),
_row(100002, "A", gender="Mixed", religious_denomination="None", latitude=_at(0.4)),
_row(100003, "B", gender="Mixed", religious_denomination="Church of England", latitude=_at(0.5)),
)
assert all(s["shared"] == [] for s in select_nearby(frame, 100001))
def test_no_tier_is_reported_because_there_are_no_tiers():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.4)),
_row(100003, "B", latitude=_at(0.5)),
)
assert all("tier" not in s for s in select_nearby(frame, 100001))
def test_metric_follows_the_phase_side_not_the_neighbour():
frame = _frame(
_row(100001, "Subject", phase="Secondary", attainment_8_score=50.0),
_row(100002, "A", phase="Secondary", attainment_8_score=52.8, latitude=_at(0.5)),
_row(100003, "B", phase="Secondary", attainment_8_score=np.nan, latitude=_at(0.6)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert by_urn[100002]["metric_key"] == "attainment_8_score"
assert by_urn[100002]["metric_value"] == 52.8
assert by_urn[100002]["metric_year"] == 202425
assert by_urn[100003]["metric_value"] is None
def test_values_are_json_safe_native_types():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
for school in select_nearby(frame, 100001):
assert isinstance(school["urn"], int)
assert isinstance(school["distance_miles"], float)
assert not isinstance(school["metric_value"], np.generic)
# ---------------------------------------------------------------------------
# The endpoint
# ---------------------------------------------------------------------------
import pytest
from fastapi.testclient import TestClient
def _endpoint_frame():
return _frame(
_row(100001, "Subject Primary"),
_row(100002, "Neighbour A", latitude=_at(0.5)),
_row(100003, "Neighbour B", latitude=_at(0.6)),
)
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_latest_school_data", _endpoint_frame)
monkeypatch.setattr(app_module, "load_school_data", _endpoint_frame)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
return TestClient(app_module.app, raise_server_exceptions=False)
def test_detail_payload_carries_nearby_schools(client):
resp = client.get("/api/schools/100001")
assert resp.status_code == 200, resp.text
similar = resp.json()["nearby_schools"]
assert [s["school_name"] for s in similar] == ["Neighbour A", "Neighbour B"]
assert similar[0]["metric_key"] == "rwm_expected_pct"
def test_a_failure_in_selection_does_not_break_the_page(client, monkeypatch):
from backend import app as app_module
def _explode(*args, **kwargs):
raise ValueError("selection blew up")
monkeypatch.setattr(app_module, "select_nearby", _explode)
resp = client.get("/api/schools/100001")
assert resp.status_code == 200, resp.text
assert resp.json()["nearby_schools"] == []
+497
View File
@@ -0,0 +1,497 @@
"""Tests for the place registry (spec 2026-08-21).
The registry is built from the in-memory school DataFrame, so these build a
small frame directly rather than touching a database.
"""
import numpy as np
import pandas as pd
import pytest
from backend.places import (MIN_SCHOOLS, build_place_index,
build_place_registry, places_for_urn)
def _df(rows: list[dict]) -> pd.DataFrame:
base = {
"year": 202425, "ofsted_grade": 2.0, "ofsted_date": None,
"rwm_expected_pct": 60.0, "attainment_8_score": np.nan,
"phase": "Primary", "postcode": "AA1 1AA",
}
return pd.DataFrame([{**base, **r} for r in rows])
def _town(n: int, town: str, la: str, start: int = 100000, **kw) -> list[dict]:
"""`start` offsets the URNs so two calls can describe different schools —
the Bedford case needs two authorities' worth of distinct URNs in one
town."""
return [
{"urn": start + i, "school_name": f"{town} School {i}",
"town": town, "local_authority": la, **kw}
for i in range(n)
]
def test_town_clearing_the_threshold_is_published():
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
assert "town:brentwood" in reg
assert reg["town:brentwood"].name == "Brentwood"
assert len(reg["town:brentwood"].urns) == MIN_SCHOOLS
def test_town_below_the_threshold_is_not_published():
reg = build_place_registry(_df(_town(MIN_SCHOOLS - 1, "Crosby", "Sefton")))
assert "town:crosby" not in reg
def test_a_town_below_threshold_still_names_its_authority():
# The route layer needs somewhere to 301 to.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS - 1, "Crosby", "Sefton") + _town(MIN_SCHOOLS, "Bootle", "Sefton")))
assert "authority:sefton" in reg
def test_town_and_authority_of_the_same_name_are_separate_places():
# 67 real collisions. Neither set contains the other: Bedford the town has
# 104 schools, Bedford the authority 86, because postal towns cross
# authority boundaries.
rows = (_town(MIN_SCHOOLS, "Bedford", "Bedford")
+ _town(MIN_SCHOOLS, "Bedford", "Central Bedfordshire", start=200000))
reg = build_place_registry(_df(rows))
town, authority = reg["town:bedford"], reg["authority:bedford"]
assert set(town.urns) != set(authority.urns)
assert len(town.urns) == MIN_SCHOOLS * 2 # both authorities' schools
assert len(authority.urns) == MIN_SCHOOLS # only this authority's
def test_schools_without_publishable_data_do_not_count_toward_the_threshold():
rows = _town(MIN_SCHOOLS, "Ghosttown", "Nowhere")
for r in rows:
r["rwm_expected_pct"] = np.nan
r["ofsted_grade"] = np.nan
reg = build_place_registry(_df(rows))
assert "town:ghosttown" not in reg
def test_blank_town_is_ignored():
rows = _town(MIN_SCHOOLS, "", "Essex")
reg = build_place_registry(_df(rows))
assert not any(k.startswith("town:") for k in reg)
def test_a_school_is_counted_once_even_with_several_years_of_rows():
rows = []
for year in (202324, 202425):
rows += [{**r, "year": year} for r in _town(MIN_SCHOOLS, "Beccles", "Suffolk")]
reg = build_place_registry(_df(rows))
assert len(reg["town:beccles"].urns) == MIN_SCHOOLS
def test_locality_groups_schools_by_outcode(monkeypatch):
# The GIAS town field collapses 1,819 London schools into "London", so a
# locality is defined by its postcode districts instead.
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"battersea": ("Battersea", ("SW11",))})
rows = _town(MIN_SCHOOLS, "London", "Wandsworth")
for r in rows:
r["postcode"] = "SW11 2AA"
reg = build_place_registry(_df(rows))
assert reg["locality:battersea"].name == "Battersea"
assert len(reg["locality:battersea"].urns) == MIN_SCHOOLS
def test_locality_below_the_threshold_is_not_published(monkeypatch):
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"nowhere": ("Nowhere", ("ZZ99",))})
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "London", "Wandsworth")))
assert "locality:nowhere" not in reg
def test_a_locality_may_not_shadow_a_viable_town(monkeypatch, caplog):
"""A colliding locality is skipped loudly, and the town survives.
This used to raise, which took down sitemap generation for all 25,000
school pages the first time a curated slug met a real GIAS town. Curated
data must not be able to break the site — and GIAS town names change with
no code change at all, so the raise could fire spontaneously.
"""
import logging
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"brentwood": ("Brentwood", ("CM13",))})
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
with caplog.at_level(logging.ERROR):
reg = build_place_registry(_df(rows))
assert "locality:brentwood" not in reg # skipped
assert "town:brentwood" in reg # the town is untouched
assert "brentwood" in caplog.text # and it was loud about it
def test_a_locality_collision_does_not_break_the_rest_of_the_registry(monkeypatch):
# The whole point of skipping rather than raising.
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"brentwood": ("Brentwood", ("CM13",))})
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
reg = build_place_registry(_df(rows))
assert "authority:essex" in reg
assert "outcode:cm13" in reg
def test_outcode_places_are_built_from_postcodes():
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:cm13"].name == "CM13"
assert len(reg["outcode:cm13"].urns) == MIN_SCHOOLS
def test_malformed_postcodes_do_not_create_places():
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "not a postcode"
reg = build_place_registry(_df(rows))
assert not any(k.startswith("outcode:") for k in reg)
def test_every_curated_locality_is_structurally_valid():
# Guards the hand-maintained file: real slug, real name, real outcodes.
import re
from backend.localities import LOCALITY_OUTCODES
assert LOCALITY_OUTCODES, "the curated locality list must not be empty"
for slug, (name, outcodes) in LOCALITY_OUTCODES.items():
assert re.fullmatch(r"[a-z0-9-]+", slug), slug
assert name.strip() == name and name, slug
assert outcodes, f"{slug} has no outcodes"
for oc in outcodes:
assert re.fullmatch(r"[A-Z]{1,2}\d{1,2}[A-Z]?", oc), (slug, oc)
def test_the_pipeline_seed_mirrors_the_canonical_module():
"""Two copies with no drift guard is worse than one copy.
backend/localities.py is canonical because the backend image does not
contain pipeline/. The seed exists so the warehouse can join on the same
definitions, and this is what stops the two diverging — the same
arrangement assert_gias_code_names_match_seed.sql gives gias_codes.
"""
import csv
from pathlib import Path
from backend.localities import LOCALITY_OUTCODES
seed_path = (Path(__file__).resolve().parents[2]
/ "pipeline/transform/seeds/locality_outcodes.csv")
assert seed_path.exists(), f"missing seed mirror at {seed_path}"
seed = {
row["locality_slug"]: (row["locality_name"],
tuple(row["outcodes"].split("|")))
for row in csv.DictReader(seed_path.open())
}
assert seed == LOCALITY_OUTCODES
def test_no_curated_locality_names_a_london_borough():
"""Boroughs are authorities and already have a page.
A locality defined by two or three outcodes inside a borough would be a
partial, near-duplicate subset of that authority page — the exact
thin-content failure the two-namespace design exists to avoid. Hackney,
Islington, Greenwich and Ealing were all in the first draft.
Hardcoded rather than read from the corpus because this must fail in CI,
where there is no database.
"""
from backend.localities import LOCALITY_OUTCODES
boroughs = {
"barking-and-dagenham", "barnet", "bexley", "brent", "bromley",
"camden", "croydon", "ealing", "enfield", "greenwich", "hackney",
"hammersmith-and-fulham", "haringey", "harrow", "havering",
"hillingdon", "hounslow", "islington", "kensington-and-chelsea",
"kingston-upon-thames", "lambeth", "lewisham", "merton", "newham",
"redbridge", "richmond-upon-thames", "southwark", "sutton",
"tower-hamlets", "waltham-forest", "wandsworth", "westminster",
}
named = boroughs & set(LOCALITY_OUTCODES)
assert not named, (
f"these are boroughs, not districts: {sorted(named)} - they already "
"have an authority page covering every school"
)
def test_a_place_names_every_authority_it_straddles():
"""SW19 is mostly Merton but partly Wandsworth.
A quarter of viable outcodes and a third of viable towns cross an
authority boundary, so naming only the largest asserts something false.
"""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
names = [n for n, _ in reg["outcode:sw19"].authorities]
assert names == ["Merton", "Wandsworth"] # largest first
assert dict(reg["outcode:sw19"].authorities)["Wandsworth"] == 7
def test_the_redirect_target_stays_a_single_authority():
# parent_authority and authorities do different jobs: a 301 needs one
# target, the page needs the truth.
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:sw19"].parent_authority == "Merton"
def test_a_stray_authority_below_the_share_threshold_is_not_named():
# GIAS carries postcode errors — EN6 lists two Shropshire schools among
# fourteen in Hertfordshire. Printing those as though real would be worse
# than omitting them.
rows = (_town(30, "Barnet", "Hertfordshire", start=300000)
+ _town(1, "Barnet", "Shropshire", start=400000))
for r in rows:
r["postcode"] = "EN6 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:en6"].authorities] == ["Hertfordshire"]
def test_a_sentinel_authority_is_never_named():
rows = (_town(20, "London", "Merton", start=300000)
+ _town(6, "London", "Does not apply", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:sw19"].authorities] == ["Merton"]
def test_a_place_always_names_at_least_one_authority():
# Even when every authority is below the share threshold, the page has to
# say where the place is.
rows = []
for i, la in enumerate(["A", "B", "C", "D", "E", "F", "G"]):
rows += _town(1, "Fragmented", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
place = reg.get("town:fragmented")
assert place is not None
assert len(place.authorities) == 1
def test_every_qualifying_authority_is_named_with_no_cap():
"""An earlier cut stopped at three, dropping the fourth silently.
That truncation bit exactly where the information matters most — a
genuinely fragmented place — and nothing recorded it.
"""
rows = []
for i, la in enumerate(["Hackney", "Lambeth", "Westminster", "Lewisham"]):
rows += _town(3, "Fourway", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
assert len(reg["town:fourway"].authorities) == 4
def test_the_redirect_target_is_the_authority_named_first():
"""They were computed separately — mode() against value_counts() — and on
an exact tie pandas does not guarantee the two agree."""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
place = build_place_registry(_df(rows))["outcode:sw19"]
assert place.parent_authority == place.authorities[0][0]
def test_a_place_never_redirects_to_a_sentinel_authority():
# Deriving the parent from `authorities` inherits its sentinel filter.
rows = (_town(6, "Someplace", "Does not apply", start=300000)
+ _town(5, "Someplace", "Essex", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:someplace"].parent_authority == "Essex"
def test_spellings_of_one_place_are_merged_not_overwritten():
"""GIAS spells the same place several ways, and they share a URL.
"London" (1,819 schools) and "LONDON" (12) both slugify to `london`.
Grouping by raw value let the later group overwrite the earlier one, so
the page could have shown twelve schools instead of 1,819 — silently, and
depending on row order.
"""
rows = (_town(6, "Weston-super-Mare", "North Somerset", start=300000)
+ _town(5, "Weston-Super-Mare", "North Somerset", start=400000))
reg = build_place_registry(_df(rows))
assert len(reg["town:weston-super-mare"].urns) == 11
def test_the_merged_place_takes_its_most_common_spelling():
rows = (_town(9, "Newcastle-under-Lyme", "Staffordshire", start=300000)
+ _town(5, "NEWCASTLE-UNDER-LYME", "Staffordshire", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:newcastle-under-lyme"].name == "Newcastle-under-Lyme"
def test_a_phase_page_needs_results_not_merely_publishable_schools():
"""/schools/kent/primary published with none of its five rows scored.
The threshold counted schools that were publishable — a result OR an
Ofsted grade — while the page exists for its results column. Forty-four
phase pages were majority-blank; one had no results at all.
"""
rows = _town(MIN_SCHOOLS, "Kent", "Kent")
for r in rows:
r["rwm_expected_pct"] = np.nan # Ofsted only, no results
reg = build_place_registry(_df(rows))
assert "town:kent" in reg # the place still publishes
assert not reg["town:kent"].publishes_phase("primary")
def test_a_phase_page_publishes_once_enough_schools_carry_a_result():
rows = _town(MIN_SCHOOLS, "Beccles", "Suffolk")
reg = build_place_registry(_df(rows))
assert reg["town:beccles"].publishes_phase("primary")
def test_a_publishing_phase_page_still_lists_its_unscored_schools():
"""The threshold gates whether the page exists; it does not filter rows.
A parent looking up a school by name has to find it whether or not it
published results.
"""
scored = _town(MIN_SCHOOLS, "Beccles", "Suffolk", start=300000)
unscored = _town(2, "Beccles", "Suffolk", start=400000)
for r in unscored:
r["rwm_expected_pct"] = np.nan
reg = build_place_registry(_df(scored + unscored))
place = reg["town:beccles"]
assert place.publishes_phase("primary")
assert len(place.phase_urns["primary"]) == MIN_SCHOOLS + 2
def test_the_secondary_threshold_counts_its_own_metric():
# A town full of scored primaries must not thereby publish a secondary page.
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
reg = build_place_registry(_df(rows))
assert not reg["town:brentwood"].publishes_phase("secondary")
def test_an_outcode_publishes_no_phase_variants():
"""There is no /schools/near/[outcode]/[phase] route, by design.
Nobody searches "primary schools in SW11", so the spec gives outcodes no
phase variants. The registry computed them anyway, and the place page —
which links whatever phases the registry reports — put two 404s on every
outcode page in the site.
This is the single rule now: a kind with no phase route reports no phases,
so neither the page nor the sitemap can offer one.
"""
rows = [{"urn": 500000 + i, "school_name": f"SW11 School {i}",
"town": "London", "local_authority": "Wandsworth",
"postcode": "SW11 1AA"} for i in range(MIN_SCHOOLS + 3)]
reg = build_place_registry(_df(rows))
place = reg["outcode:sw11"]
assert place.phase_urns == {}
assert not place.publishes_phase("primary")
assert not place.publishes_phase("secondary")
def test_an_authority_still_publishes_phase_variants():
"""Authorities keep theirs — "primary schools in Kent" is a real query,
and /schools/authority/[la]/[phase] is the route that serves it."""
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Maidstone", "Kent")))
assert reg["authority:kent"].publishes_phase("primary")
# ── The reverse index: which published places contain a school ──────────────
#
# School pages link out to the location layer through this. It is the whole
# point of the index: before it, ~27k school pages linked to nothing on the
# site and stranded whatever authority they held.
def test_a_school_resolves_to_every_published_place_containing_it():
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
places = places_for_urn(build_place_index(reg), 100000)
kinds = {p.kind for p in places}
assert "town" in kinds
assert "authority" in kinds
def test_a_school_in_an_unpublished_town_still_resolves_to_its_authority():
# A town below the threshold has no page, so there is no link to offer —
# but the authority above it clears the threshold on the same schools and
# is where that reader should be sent.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS - 1, "Tinytown", "Essex")
+ _town(MIN_SCHOOLS, "Brentwood", "Essex", start=200000)
))
places = places_for_urn(build_place_index(reg), 100000)
# The town is below the threshold, so it has no page and must not be
# offered as a link. The authority above it does, and is the right target.
assert all(p.slug != "tinytown" for p in places)
assert "authority" in {p.kind for p in places}
def test_an_unknown_urn_resolves_to_nothing_rather_than_raising():
# A school page renders for any URN the API knows; the link module is not
# entitled to take the page down when it has nothing to say.
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
assert places_for_urn(build_place_index(reg), 999999) == ()
def test_the_index_is_consistent_with_the_registry_it_was_built_from():
# The invariant that matters: a link module must never offer a place whose
# page does not exist, and never omit one that does.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS, "Brentwood", "Essex")
+ _town(MIN_SCHOOLS, "Bedford", "Bedford", start=300000)
))
index = build_place_index(reg)
for key, place in reg.items():
for urn in place.urns:
assert place in places_for_urn(index, urn), (
f"{urn} is in {key} but the index does not say so")
def test_the_index_holds_no_school_the_registry_does_not():
# The reverse direction of the invariant above. An index entry for a URN
# no published place contains would put a link on a page for a place that
# does not list that school.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS, "Brentwood", "Essex")
+ _town(MIN_SCHOOLS - 1, "Tinytown", "Essex", start=400000)
))
index = build_place_index(reg)
for urn, places in index.items():
for place in places:
assert urn in place.urns
assert place.key in reg
def test_the_index_preserves_the_widest_first_order():
# The breadcrumb reads authority then town, and takes this order as given.
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
kinds = [p.kind for p in places_for_urn(build_place_index(reg), 100000)]
assert kinds.index("authority") < kinds.index("town")
+185
View File
@@ -0,0 +1,185 @@
"""Tests for the places API (spec 2026-08-21)."""
import numpy as np
import pandas as pd
import pytest
from fastapi.testclient import TestClient
def _schools_df() -> pd.DataFrame:
base = {
"local_authority": "Essex", "school_type": "Academy",
"phase": "Primary", "year": 202425, "ofsted_grade": 2.0,
"ofsted_date": None, "attainment_8_score": np.nan,
"town": "Brentwood", "postcode": "CM13 1AA", "status": "Open",
"address": "1 Test Street", "latitude": 51.6, "longitude": 0.3,
}
return pd.DataFrame([
{**base, "urn": 100000 + i, "school_name": f"Brentwood School {i}",
"rwm_expected_pct": 50.0 + i}
for i in range(6)
])
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _schools_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_registry_lists_each_published_place(client):
body = client.get("/api/places").json()
slugs = {(p["kind"], p["slug"]) for p in body["places"]}
assert ("town", "brentwood") in slugs
assert ("authority", "essex") in slugs
assert ("outcode", "cm13") in slugs
def test_registry_carries_a_count_per_place(client):
body = client.get("/api/places").json()
town = next(p for p in body["places"] if p["slug"] == "brentwood")
assert town["count"] == 6
def test_place_detail_returns_its_schools_alphabetically(client):
"""A place page is read by someone looking for a school they can name.
Scanning for it is what the order should serve, so the list is A-Z.
/api/rankings is where the league-table ordering lives.
"""
body = client.get("/api/places/town/brentwood").json()
assert body["place"]["name"] == "Brentwood"
names = [s["school_name"] for s in body["schools"]]
assert names == sorted(names, key=str.lower)
def test_place_ordering_ignores_case(client):
body = client.get("/api/places/town/brentwood").json()
names = [s["school_name"] for s in body["schools"]]
# A capitalised name must not sort ahead of every lowercase one.
assert names == sorted(names, key=str.lower)
def test_the_rankings_endpoint_still_ranks_by_metric(client):
# Alphabetical is a place-page decision, not a site-wide one.
body = client.get("/api/rankings?metric=rwm_expected_pct&phase=primary").json()
scores = [r["rwm_expected_pct"] for r in body.get("rankings", [])
if r.get("rwm_expected_pct") is not None]
assert scores == sorted(scores, reverse=True)
def test_place_detail_carries_the_local_average(client):
body = client.get("/api/places/town/brentwood").json()
# 50..55 inclusive
assert body["averages"]["rwm_expected_pct"] == pytest.approx(52.5)
def test_phase_filter_narrows_the_school_list(client):
body = client.get("/api/places/town/brentwood?phase=secondary").json()
assert body["schools"] == []
def test_unknown_place_404s(client):
assert client.get("/api/places/town/atlantis").status_code == 404
def test_unknown_kind_404s(client):
assert client.get("/api/places/planet/mars").status_code == 404
def _straddling_df() -> pd.DataFrame:
"""Eight schools in CM13: six in Essex, which has a page, and two in an
authority too small to have one.
Two, not one: the registry ignores an authority holding a single school in
a place, because GIAS carries occasional postcode errors."""
df = _schools_df()
extra = df.iloc[:2].copy()
extra["urn"] = [200000, 200001]
extra["school_name"] = ["Scilly School 0", "Scilly School 1"]
extra["local_authority"] = "Isles Of Scilly"
return pd.concat([df, extra], ignore_index=True)
@pytest.fixture()
def straddling_client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _straddling_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _straddling_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_an_outcode_reports_no_phases_because_it_has_no_phase_route(client):
body = client.get("/api/places/outcode/cm13").json()
assert body["place"]["phases"] == []
def test_an_authority_reports_the_phases_it_publishes(client):
body = client.get("/api/places/authority/essex").json()
assert body["place"]["phases"] == ["primary"]
def test_an_authority_without_a_page_is_named_but_carries_no_slug(straddling_client):
"""Two English authorities — City of London and the Isles of Scilly — hold
fewer than the five schools a page needs, so they have no page.
Naming them is still right: the page says where the place is. Linking them
would not be. A null slug is what tells the page to print the name plainly
rather than invent a URL that 404s.
"""
body = straddling_client.get("/api/places/outcode/cm13").json()
by_name = {a["name"]: a for a in body["place"]["authorities"]}
assert by_name["Essex"]["slug"] == "essex"
assert by_name["Isles Of Scilly"]["slug"] is None
def _attributed_df() -> pd.DataFrame:
"""The same town, with the four attributes the place table now shows."""
df = _schools_df()
df["age_range"] = "4-11"
df["religious_denomination"] = "Church of England"
df["nursery_provision"] = True
df["parliamentary_constituency"] = "Brentwood and Ongar"
return df
@pytest.fixture()
def attributed_client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _attributed_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _attributed_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_place_detail_carries_the_attributes_the_table_shows(attributed_client):
"""age_range and religious_denomination ride in on SCHOOL_COLUMNS.
nursery_provision and parliamentary_constituency do not, and the place
table needs all four — a column the response cannot fill is a column of
dashes on ~3,900 pages.
"""
body = attributed_client.get("/api/places/town/brentwood").json()
school = body["schools"][0]
assert school["age_range"] == "4-11"
assert school["religious_denomination"] == "Church of England"
assert school["nursery_provision"] is True
assert school["parliamentary_constituency"] == "Brentwood and Ongar"
def test_place_detail_survives_a_mart_without_the_optional_columns(client):
"""The base fixture has neither column, as an unrebuilt mart does not.
data_loader degrades those to NULL rather than failing the load, so the
endpoint must not assume they are present.
"""
res = client.get("/api/places/town/brentwood")
assert res.status_code == 200
assert "nursery_provision" not in res.json()["schools"][0]
+68
View File
@@ -0,0 +1,68 @@
"""Publication must preserve the current dataset until every replacement is ready."""
import asyncio
import pandas as pd
import pytest
from fastapi.testclient import TestClient
from backend import app as api, data_loader
from backend.tests.test_sixth_form_flag import _schools_df
@pytest.fixture
def client(monkeypatch):
old = _schools_df()
monkeypatch.setattr(data_loader, '_df_cache', old)
monkeypatch.setattr(data_loader, '_df_latest_cache', old)
monkeypatch.setattr(api, '_place_registry', {'old': 'registry'})
monkeypatch.setattr(api, '_place_index', {'old': 'index'})
monkeypatch.setattr(api, '_place_index_source', api._place_registry)
monkeypatch.setattr(api, '_sitemaps', {'old.xml': 'old sitemap'})
monkeypatch.setattr(api, '_publication_lock', asyncio.Lock())
monkeypatch.setattr(api.limiter, 'enabled', False)
api.app.dependency_overrides[api.verify_admin_api_key] = lambda: True
yield TestClient(api.app, raise_server_exceptions=False)
api.app.dependency_overrides.clear()
def state():
return (data_loader._df_cache, data_loader._df_latest_cache, api._place_registry,
api._place_index, api._place_index_source, api._sitemaps)
@pytest.mark.parametrize('failure', ['empty', 'database', 'sitemap', 'duplicate'])
def test_failed_reload_preserves_every_published_object(client, monkeypatch, failure):
before = state()
df = _schools_df()
if failure == 'empty':
df = pd.DataFrame()
if failure == 'duplicate':
df = pd.concat([df, df.iloc[:1]], ignore_index=True)
def load():
if failure == 'database':
raise RuntimeError('database unavailable')
return df
monkeypatch.setattr(api, 'load_school_data_as_dataframe', load)
if failure == 'sitemap':
monkeypatch.setattr(api, 'build_sitemaps', lambda *args: (_ for _ in ()).throw(RuntimeError('bad XML')))
response = client.post('/api/admin/reload')
assert response.status_code == 503
assert all(a is b for a, b in zip(before, state()))
def test_success_publishes_school_data_places_and_sitemaps(client, monkeypatch):
df = _schools_df()
df.loc[0, 'school_name'] = 'Replacement School'
monkeypatch.setattr(api, 'load_school_data_as_dataframe', lambda: df)
response = client.post('/api/admin/reload')
assert response.status_code == 200
assert data_loader.load_school_data() is df
assert data_loader.load_latest_school_data().iloc[0].school_name == 'Replacement School'
assert api._place_index_source is api._place_registry
assert 'old.xml' not in api._sitemaps
assert 'replacement-school' in api._sitemaps['schools-1.xml']
def test_failed_sitemap_regeneration_keeps_existing_publication(client, monkeypatch):
before = state()
monkeypatch.setattr(api, 'build_sitemaps', lambda *args: (_ for _ in ()).throw(RuntimeError('bad XML')))
assert client.post('/api/admin/regenerate-sitemap').status_code == 503
assert all(a is b for a, b in zip(before, state()))
+147
View File
@@ -0,0 +1,147 @@
"""The rate-limit bucket must be the caller, not the proxy in front of them.
`get_remote_address` reads request.client.host. In staging and production the
backend has no published ports and its only caller is the Next proxy, so that
host is the Next container — one bucket for every browser user on the site.
Measured before this fix: 70 concurrent requests, 60 served and 10 refused.
"""
from starlette.datastructures import Headers
from backend.app import client_key
class _Req:
"""Enough of a Request for the key function: headers and a client host."""
def __init__(self, headers: dict, host: str = "10.0.0.9"):
self.headers = Headers(headers)
self.client = type("C", (), {"host": host})()
self.scope = {"type": "http", "client": (host, 0),
"headers": [(k.lower().encode(), v.encode())
for k, v in headers.items()]}
def test_cloudflare_header_wins():
# Cloudflare sets CF-Connecting-IP and overwrites any client-supplied
# value, so it is trustworthy in a way a parsed XFF chain is not.
assert client_key(_Req({"cf-connecting-ip": "203.0.113.7"})) == "203.0.113.7"
def test_forwarded_for_is_the_fallback_and_takes_the_first_entry():
# Left-most is the original client; everything after it is proxies.
assert client_key(
_Req({"x-forwarded-for": "203.0.113.7, 10.0.0.2"})) == "203.0.113.7"
def test_remote_address_is_the_last_resort():
assert client_key(_Req({}, host="10.0.0.9")) == "10.0.0.9"
def test_cloudflare_header_beats_forwarded_for():
key = client_key(_Req({"cf-connecting-ip": "203.0.113.7",
"x-forwarded-for": "198.51.100.1"}))
assert key == "203.0.113.7"
def test_two_callers_get_two_buckets():
# The whole point: one user exhausting their limit must not refuse another.
a = client_key(_Req({"cf-connecting-ip": "203.0.113.7"}))
b = client_key(_Req({"cf-connecting-ip": "203.0.113.8"}))
assert a != b
def test_whitespace_is_stripped():
# "a, b" split on comma leaves a leading space on every entry but the
# first; an unstripped key silently creates a second bucket per client.
assert client_key(_Req({"x-forwarded-for": " 203.0.113.7 ,10.0.0.2"})) \
== "203.0.113.7"
# ---------------------------------------------------------------------------
# The ceiling that header rotation cannot raise.
# ---------------------------------------------------------------------------
import pytest
from fastapi.testclient import TestClient
@pytest.fixture()
def api(monkeypatch):
from backend import app as app_module
from backend.config import settings
monkeypatch.setattr(settings, "global_rate_limit_per_minute", 5)
monkeypatch.setattr(app_module, "_global_window", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def _ceiling_req(path: str, host: str):
"""Enough of a Request for exempt_from_ceiling: a path and a peer host."""
return type("R", (), {
"url": type("U", (), {"path": path})(),
"client": type("C", (), {"host": host})(),
})()
def _get(client, path="/api/flags", cf=None):
headers = {"cf-connecting-ip": cf} if cf else {}
return client.get(path, headers=headers)
def test_rotating_the_cloudflare_header_cannot_buy_unlimited_requests(api):
"""The attack the per-client keying opened up.
client_key trusts CF-Connecting-IP, and nothing in this process can tell an
edge-set header from an attacker-set one — that distinction can only be
made at Cloudflare, with Authenticated Origin Pulls or an origin firewall.
A caller reaching the origin directly can therefore mint a fresh
rate-limit bucket per request and evade per-client limits entirely.
Per-client fairness is still the right default; this is the backstop that
bounds what evading it can achieve. Without it, correct keying would be a
net regression against abuse compared with the shared bucket it replaced.
"""
codes = [_get(api, cf=f"203.0.113.{i}").status_code for i in range(8)]
assert codes.count(200) == 5
assert codes.count(429) == 3
def test_the_ceiling_says_which_limit_was_hit(api):
# Distinguishable from slowapi's per-client 429, or an operator reading
# logs cannot tell "one noisy client" from "the origin is saturated".
for i in range(5):
_get(api, cf=f"203.0.113.{i}")
refused = _get(api, cf="203.0.113.99")
assert refused.status_code == 429
assert "capacity" in refused.json()["detail"].lower()
assert refused.headers.get("retry-after")
def test_traffic_below_the_ceiling_is_untouched(api):
codes = [_get(api, cf=f"203.0.113.{i}").status_code for i in range(5)]
assert codes == [200] * 5
def test_the_container_healthcheck_is_exempt(api):
"""The healthcheck runs `curl http://localhost:80/api/data-info` inside the
container. If the ceiling could starve it, saturation would fail the
healthcheck, restart the container, and turn a load spike into an outage
loop — the ceiling has to protect the origin, not kill it.
"""
from backend.app import exempt_from_ceiling
assert exempt_from_ceiling(_ceiling_req("/api/data-info", "127.0.0.1"))
assert exempt_from_ceiling(_ceiling_req("/api/data-info", "::1"))
# Everyone else is counted.
assert not exempt_from_ceiling(_ceiling_req("/api/data-info", "10.0.0.9"))
def test_the_ceiling_ignores_non_api_paths():
# Sitemaps and robots.txt are served by this app too, and a crawler
# fetching them must not be refused because the API is busy.
from backend.app import exempt_from_ceiling
assert exempt_from_ceiling(_ceiling_req("/sitemap.xml", "10.0.0.9"))
assert exempt_from_ceiling(_ceiling_req("/robots.txt", "10.0.0.9"))
+176
View File
@@ -56,6 +56,11 @@ def client(monkeypatch):
monkeypatch.setattr(
app_module, "get_supplementary_data", lambda db, urn: {}
)
# The place registry is a module-level cache, so without this the endpoint
# answers from whatever registry an earlier test happened to leave behind
# — and a `places == []` assertion is satisfied by a stale registry just
# as well as by this fixture's own data, which makes it prove nothing.
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
@@ -69,3 +74,174 @@ def test_nan_gias_fields_serialize_as_null(client):
assert info["capacity"] is None
assert info["total_pupils"] is None
assert info["school_name"] == "West London Performing Arts Academy"
# ── Links out to the location layer ─────────────────────────────────────────
#
# School pages carried no link into the site at all: the only anchor on the
# template pointed at the school's own website, so ~27k pages received
# whatever authority the site had and sent it off-site. `places` is what the
# link module and the breadcrumb are built from.
def test_places_is_present_even_when_the_school_belongs_to_none(client):
# This fixture's single school cannot clear any publish threshold, so the
# honest answer is an empty list. The key must still be there: a missing
# key and "no places" are different things to the page rendering it.
body = client.get("/api/schools/150275").json()
assert body["places"] == []
def test_places_names_only_pages_that_exist(monkeypatch):
from backend import app as app_module
from backend.places import MIN_SCHOOLS
def _df():
return pd.DataFrame([
{
"urn": 100000 + i,
"school_name": f"Brentwood School {i}",
"town": "Brentwood",
"local_authority": "Essex",
"postcode": "CM15 8AA",
"phase": "Primary",
"year": 202425,
"rwm_expected_pct": 60.0,
"attainment_8_score": np.nan,
"ofsted_grade": 2.0,
"ofsted_date": None,
}
for i in range(MIN_SCHOOLS)
])
monkeypatch.setattr(app_module, "load_school_data", _df)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
client = TestClient(app_module.app, raise_server_exceptions=False)
places = client.get("/api/schools/100000").json()["places"]
assert places, "a school in a published town must offer links"
by_kind = {p["kind"]: p for p in places}
assert by_kind["town"]["url"] == "/schools/brentwood"
assert by_kind["authority"]["url"] == "/schools/authority/essex"
# Every entry carries what the link text needs, and a count, so the anchor
# can say what it leads to rather than "click here".
for place in places:
assert place["name"]
assert place["count"] >= 1
assert place["url"].startswith("/schools/")
def _brentwood_df(phase: str = "Primary", n: int = None):
from backend.places import MIN_SCHOOLS
n = n if n is not None else MIN_SCHOOLS
return lambda: pd.DataFrame([
{
"urn": 100000 + i,
"school_name": f"Brentwood School {i}",
"town": "Brentwood", "local_authority": "Essex",
"postcode": "CM15 8AA", "phase": phase, "year": 202425,
"rwm_expected_pct": 60.0, "attainment_8_score": 50.0,
"ofsted_grade": 2.0, "ofsted_date": None,
}
for i in range(n)
])
def _places_for(monkeypatch, df_factory, urn: int):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", df_factory)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
client = TestClient(app_module.app, raise_server_exceptions=False)
return client.get(f"/api/schools/{urn}").json()["places"]
def test_a_place_offers_the_phase_page_this_school_appears_on(monkeypatch):
# "primary schools in brentwood" is the query the phase pages exist for,
# and ~950 of them were once reachable by nothing at all.
places = _places_for(monkeypatch, _brentwood_df("Primary"), 100000)
town = next(p for p in places if p["kind"] == "town")
assert town["phases"], "a primary school in a published primary town has a link"
assert town["phases"][0]["url"] == "/schools/brentwood/primary"
assert town["phases"][0]["count"] >= 1
def test_an_all_through_school_offers_both_phase_pages(monkeypatch):
# It genuinely appears on both, so there is no tie to break.
places = _places_for(monkeypatch, _brentwood_df("All-through"), 100000)
town = next(p for p in places if p["kind"] == "town")
assert {p["phase"] for p in town["phases"]} == {"primary", "secondary"}
def test_outcodes_never_offer_a_phase_page(monkeypatch):
# The registry gives outcodes no phase route — nobody searches "primary
# schools in SW11" — and computing them anyway once put a link to a
# nonexistent route on all 1,720 outcode pages.
places = _places_for(monkeypatch, _brentwood_df("Primary"), 100000)
outcode = next((p for p in places if p["kind"] == "outcode"), None)
if outcode is not None:
assert outcode["phases"] == []
def test_a_school_absent_from_the_phase_page_is_not_linked_to_it(monkeypatch):
# The check is URN membership in the registry's own phase list, not a
# re-derivation of the phase mapping. A secondary school must not be sent
# to a primary phase page that does not list it.
from backend.places import MIN_SCHOOLS
def df():
rows = [
{"urn": 100000 + i, "school_name": f"P{i}", "town": "Brentwood",
"local_authority": "Essex", "postcode": "CM15 8AA",
"phase": "Primary", "year": 202425, "rwm_expected_pct": 60.0,
"attainment_8_score": np.nan, "ofsted_grade": 2.0,
"ofsted_date": None}
for i in range(MIN_SCHOOLS)
]
rows.append({
"urn": 900000, "school_name": "Lone Secondary", "town": "Brentwood",
"local_authority": "Essex", "postcode": "CM15 8AA",
"phase": "Secondary", "year": 202425, "rwm_expected_pct": np.nan,
"attainment_8_score": 50.0, "ofsted_grade": 2.0, "ofsted_date": None,
})
return pd.DataFrame(rows)
places = _places_for(monkeypatch, df, 900000)
town = next(p for p in places if p["kind"] == "town")
# The town publishes a primary page, but this secondary school is not on
# it, and there are too few secondaries for a secondary page.
assert town["phases"] == []
def test_the_place_index_rebuilds_when_the_registry_is_replaced(monkeypatch):
"""The reverse index is cached; a stale one would put another dataset's
places on a school page. Invalidation is an identity check against the
registry rather than a second flag, so this asserts the check works."""
from backend import app as app_module
monkeypatch.setattr(app_module, "_place_registry", None)
monkeypatch.setattr(app_module, "_place_index", None)
monkeypatch.setattr(app_module, "_place_index_source", None)
monkeypatch.setattr(app_module, "load_school_data", _brentwood_df("Primary"))
first = app_module.get_place_index()
assert 100000 in first
# Same registry object, so the index is reused rather than rebuilt.
assert app_module.get_place_index() is first
# Drop the registry the way every test that touches place data does. The
# index must follow it, not survive it.
app_module._place_registry = None
monkeypatch.setattr(app_module, "load_school_data",
_brentwood_df("Primary", n=0))
rebuilt = app_module.get_place_index()
assert rebuilt is not first
assert 100000 not in rebuilt, "the index outlived the registry it came from"
+114
View File
@@ -0,0 +1,114 @@
from types import SimpleNamespace
import pytest
from fastapi.testclient import TestClient
from backend import app as api, data_loader
from backend.tests.test_sixth_form_flag import _schools_df
def client_for(monkeypatch, search):
client = SimpleNamespace(collections={'schools': SimpleNamespace(documents=SimpleNamespace(search=search))})
monkeypatch.setattr(data_loader, '_get_typesense_client', lambda: client)
def test_search_returns_matches_beyond_first_page(monkeypatch):
pages = []
def search(params):
pages.append(params['page'])
urns = range(100000, 100250) if params['page'] == 1 else [100999]
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in urns]}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('academy')
assert len(result) == 251
assert result[-1] == 100999
assert pages == [1, 2]
def test_search_caps_broad_queries_at_a_bounded_number_of_pages(monkeypatch):
requests = []
def search(params):
requests.append(params)
return {
'found': 10_000,
'hits': [
{'document': {'urn': 100000 + params['page'] * 1000 + i}}
for i in range(params['per_page'])
],
}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('school')
assert len(result) == data_loader.SEARCH_MAX_CANDIDATES
assert len(requests) == data_loader.SEARCH_MAX_CANDIDATES // data_loader.SEARCH_PAGE_SIZE
assert all(request['per_page'] == data_loader.SEARCH_PAGE_SIZE for request in requests)
assert requests[-1]['page'] == len(requests)
def test_search_uses_a_smaller_final_page_when_the_cap_is_not_a_page_multiple(monkeypatch):
monkeypatch.setattr(data_loader, 'SEARCH_MAX_CANDIDATES', 251)
requests = []
def search(params):
requests.append(params)
return {
'found': 10_000,
'hits': [{'document': {'urn': 100000 + len(requests) * 1000 + i}}
for i in range(params['per_page'])],
}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('school')
assert len(result) == 251
assert [request['per_page'] for request in requests] == [250, 1]
def test_later_page_failure_does_not_return_partial_results(monkeypatch):
def search(params):
if params['page'] == 2:
raise RuntimeError('timeout')
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in range(100000, 100250)]}
client_for(monkeypatch, search)
assert data_loader.search_schools_typesense('academy') is None
def test_zero_matches_are_distinct_from_unavailable(monkeypatch):
client_for(monkeypatch, lambda _: {'found': 0, 'hits': []})
assert data_loader.search_schools_typesense('academy') == []
monkeypatch.setattr(data_loader, '_get_typesense_client', lambda: None)
assert data_loader.search_schools_typesense('academy') is None
@pytest.mark.parametrize('matches, expected', [([], []), (None, [100001])])
def test_fallback_only_on_dependency_failure(monkeypatch, matches, expected):
monkeypatch.setattr(api.limiter, 'enabled', False)
monkeypatch.setattr(api, 'load_latest_school_data', _schools_df)
monkeypatch.setattr(api, 'search_schools_typesense', lambda _: matches)
response = TestClient(api.app).get('/api/schools?search=Alpha')
assert response.status_code == 200
assert [s['urn'] for s in response.json()['schools']] == expected
def test_filtered_api_keeps_match_from_second_search_page(monkeypatch):
monkeypatch.setattr(api.limiter, 'enabled', False)
df = _schools_df()
monkeypatch.setattr(api, 'load_latest_school_data', lambda: df)
def search(params):
urns = range(200000, 200250) if params['page'] == 1 else [100001]
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in urns]}
client_for(monkeypatch, search)
response = TestClient(api.app).get('/api/schools?search=Alpha&local_authority=Testshire')
assert response.status_code == 200
assert response.json()['total'] == 1
assert response.json()['schools'][0]['urn'] == 100001
def test_unavailable_dataset_is_not_a_missing_school_or_empty_search(monkeypatch):
import pandas as pd
monkeypatch.setattr(api.limiter, 'enabled', False)
monkeypatch.setattr(api, 'load_school_data', lambda: pd.DataFrame())
monkeypatch.setattr(api, 'load_latest_school_data', lambda: pd.DataFrame())
client = TestClient(api.app)
assert client.get('/api/schools/100001').status_code == 503
assert client.get('/api/schools?search=school').status_code == 503
+306
View File
@@ -0,0 +1,306 @@
"""Tests for sitemap generation (spec 2026-08-20, workstream W1).
The sitemap is built from the in-memory school DataFrame, so these inject a
small frame via monkeypatch rather than touching a database.
"""
import numpy as np
import pandas as pd
import pytest
def _schools_df() -> pd.DataFrame:
"""Two schools: one with results, one with neither results nor Ofsted."""
base = {
"local_authority": "Testshire",
"school_type": "Academy",
"phase": "Primary",
"year": 202425,
"ofsted_date": None,
}
return pd.DataFrame(
[
{**base, "urn": 100001, "school_name": "Alpha Primary",
"rwm_expected_pct": 62.0, "attainment_8_score": np.nan,
"ofsted_grade": 2.0},
{**base, "urn": 100002, "school_name": "Ghost Primary",
"rwm_expected_pct": np.nan, "attainment_8_score": np.nan,
"ofsted_grade": np.nan},
]
)
@pytest.fixture()
def sitemap(monkeypatch) -> str:
"""The sitemap index."""
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
return app_module.build_sitemap()
@pytest.fixture()
def schools_child(monkeypatch) -> str:
"""The first school child sitemap, where school URLs actually live."""
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
return app_module.build_sitemaps()["schools-1.xml"]
@pytest.fixture()
def static_child(monkeypatch) -> str:
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
return app_module.build_sitemaps()["static.xml"]
def test_every_loc_uses_the_www_host(sitemaps):
# The apex 301s to www. A <loc> that redirects burns a crawl per URL.
# Checked across every file, index included, not just one.
#
# A child can legitimately be empty — this fixture holds two schools and no
# town clearing the threshold — so the presence check applies only to files
# that carry URLs. The absence check applies to all of them.
for name, xml in sitemaps.items():
assert "https://schoolcompare.co.uk" not in xml, name
if "<loc>" in xml:
assert "https://www.schoolcompare.co.uk" in xml, name
def test_school_with_results_is_listed(schools_child):
assert "/school/100001-alpha-primary" in schools_child
def test_school_with_no_results_and_no_ofsted_is_omitted(schools_child):
# Nothing for a search result to say about it. Submitting it spends crawl
# budget and drags the corpus-wide quality signal down.
#
# Asserted against the child, not the index: the index carries no school
# URLs at all, so it would pass this trivially and prove nothing.
assert "/school/100002" not in schools_child
def test_no_invented_priority_or_changefreq(sitemaps):
# Google ignores both. They were noise dressed as signal.
for name, xml in sitemaps.items():
assert "<priority>" not in xml, name
assert "<changefreq>" not in xml, name
def test_ofsted_date_becomes_lastmod(monkeypatch):
from backend import app as app_module
import datetime
def _df():
base = _schools_df()
base.loc[base["urn"] == 100001, "ofsted_date"] = datetime.date(2024, 3, 14)
return base
monkeypatch.setattr(app_module, "load_school_data", _df)
xml = app_module.build_sitemaps()["schools-1.xml"]
assert "<lastmod>2024-03-14</lastmod>" in xml
def test_no_lastmod_invented_when_date_unknown(monkeypatch):
# An always-now lastmod is a claim Google learns to distrust. Absent
# honestly means unknown.
from backend import app as app_module
def _df():
df = _schools_df()
df["ofsted_date"] = None
return df
monkeypatch.setattr(app_module, "load_school_data", _df)
# The child only. The index legitimately carries a lastmod, because there
# it means "when this sitemap file changed", which we do know.
xml = app_module.build_sitemaps()["schools-1.xml"]
assert "<lastmod>" not in xml
def test_static_routes_are_listed(static_child):
for path in ("/", "/rankings", "/compare", "/admissions"):
assert f"<loc>https://www.schoolcompare.co.uk{path}</loc>" in static_child
@pytest.fixture()
def sitemaps(monkeypatch) -> dict:
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
return app_module.build_sitemaps()
def test_index_lists_each_child(sitemaps):
index = sitemaps["sitemap.xml"]
assert "<sitemapindex" in index
assert "https://www.schoolcompare.co.uk/sitemaps/static.xml" in index
assert "https://www.schoolcompare.co.uk/sitemaps/schools-1.xml" in index
def test_index_carries_no_url_elements(sitemaps):
# A sitemap index holds <sitemap> entries only; mixing in <url> is invalid.
assert "<url>" not in sitemaps["sitemap.xml"]
def test_index_does_not_list_itself(sitemaps):
assert "<loc>https://www.schoolcompare.co.uk/sitemap.xml</loc>" not in sitemaps["sitemap.xml"]
def test_static_child_holds_the_static_routes(sitemaps):
static = sitemaps["static.xml"]
for path in ("/", "/rankings", "/compare", "/admissions"):
assert f"<loc>https://www.schoolcompare.co.uk{path}</loc>" in static
def test_school_child_holds_the_schools(sitemaps):
assert "/school/100001-alpha-primary" in sitemaps["schools-1.xml"]
def test_children_are_chunked_under_the_limit(monkeypatch):
# Sitemaps cap at 50,000 URLs per file. Chunk at 10,000 so a child stays
# small enough to eyeball in Search Console.
from backend import app as app_module
import pandas as _pd
rows = [
{"urn": 200000 + i, "school_name": f"School {i}", "year": 202425,
"rwm_expected_pct": 60.0, "attainment_8_score": None,
"ofsted_grade": 2.0, "ofsted_date": None}
for i in range(10_001)
]
monkeypatch.setattr(app_module, "load_school_data", lambda: _pd.DataFrame(rows))
maps = app_module.build_sitemaps()
assert maps["schools-1.xml"].count("<url>") == 10_000
assert maps["schools-2.xml"].count("<url>") == 1
def test_build_sitemap_still_returns_the_index(sitemap):
# lifespan and the admin endpoint call build_sitemap(); keep it working.
assert "<sitemapindex" in sitemap
def test_school_with_results_in_an_earlier_year_is_still_listed(monkeypatch):
"""Regression: The Mallard Academy (150367).
Real KS2 results 2015-16 to 2018-19, then null rows from 2022-23 onward
because the school stopped reporting. The first cut tested the latest
year's row alone and dropped it, along with ~220 others, even though its
detail page shows all four years of results.
"""
from backend import app as app_module
import pandas as _pd
base = {"local_authority": "Testshire", "school_type": "Academy",
"phase": "Primary", "ofsted_date": None, "ofsted_grade": np.nan,
"attainment_8_score": np.nan, "urn": 150367,
"school_name": "Mallard Academy"}
df = _pd.DataFrame([
{**base, "year": 201819, "rwm_expected_pct": 67.0},
{**base, "year": 202324, "rwm_expected_pct": np.nan},
{**base, "year": 202425, "rwm_expected_pct": np.nan},
])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
xml = app_module.build_sitemaps()["schools-1.xml"]
assert "/school/150367-mallard-academy" in xml
def test_school_with_no_results_in_any_year_is_still_omitted(monkeypatch):
"""The fix must not turn into "list everything"."""
from backend import app as app_module
import pandas as _pd
base = {"local_authority": "Testshire", "school_type": "Academy",
"phase": "Primary", "ofsted_date": None, "ofsted_grade": np.nan,
"attainment_8_score": np.nan, "rwm_expected_pct": np.nan,
"urn": 100002, "school_name": "Ghost Primary"}
df = _pd.DataFrame([{**base, "year": y} for y in (202324, 202425)])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
assert "/school/100002" not in app_module.build_sitemaps()["schools-1.xml"]
def _places_df() -> pd.DataFrame:
base = {
"local_authority": "Essex", "school_type": "Academy",
"phase": "Primary", "year": 202425, "ofsted_grade": 2.0,
"ofsted_date": None, "attainment_8_score": np.nan,
"town": "Brentwood", "postcode": "CM13 1AA",
}
return pd.DataFrame([
{**base, "urn": 100000 + i, "school_name": f"Brentwood School {i}",
"rwm_expected_pct": 60.0}
for i in range(6)
])
@pytest.fixture()
def place_sitemaps(monkeypatch) -> dict:
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _places_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return app_module.build_sitemaps()
def test_place_children_are_listed_in_the_index(place_sitemaps):
index = place_sitemaps["sitemap.xml"]
assert "/sitemaps/places-1.xml" in index
assert "/sitemaps/outcodes-1.xml" in index
def test_town_and_authority_urls_use_their_own_namespaces(place_sitemaps):
xml = place_sitemaps["places-1.xml"]
assert "<loc>https://www.schoolcompare.co.uk/schools/brentwood</loc>" in xml
assert "<loc>https://www.schoolcompare.co.uk/schools/authority/essex</loc>" in xml
def test_outcode_urls_live_in_their_own_child(place_sitemaps):
assert "/schools/near/cm13" in place_sitemaps["outcodes-1.xml"]
assert "/schools/near/cm13" not in place_sitemaps["places-1.xml"]
def test_place_urls_carry_no_priority_or_changefreq(place_sitemaps):
for name in ("places-1.xml", "outcodes-1.xml"):
assert "<priority>" not in place_sitemaps[name]
assert "<changefreq>" not in place_sitemaps[name]
def test_phase_variants_are_submitted_where_the_phase_clears_the_threshold(place_sitemaps):
# "primary schools in beccles" is the query shape the baseline showed, so
# each variant is its own page and has to be submitted. Emitting only the
# bare place URL left ~950 of them reachable by nothing.
xml = place_sitemaps["places-1.xml"]
assert "<loc>https://www.schoolcompare.co.uk/schools/brentwood/primary</loc>" in xml
def test_a_phase_below_its_own_threshold_is_not_submitted(place_sitemaps):
# The fixture is six primaries and no secondaries.
xml = place_sitemaps["places-1.xml"]
assert "/schools/brentwood/secondary" not in xml
def test_outcodes_get_no_phase_variants(place_sitemaps):
# Nobody searches "primary schools in CM13"; the routes do not exist.
xml = place_sitemaps["outcodes-1.xml"]
assert "/primary" not in xml and "/secondary" not in xml
def test_authority_phase_variants_are_submitted_in_their_own_namespace(place_sitemaps):
"""302 of these were already in the sitemap, and every one 404'd.
The spec gives authorities a phase route; the plan built the bare
authority route and dropped it. Nothing noticed because the sitemap was
written from the registry, which was right, while the routes were written
by hand. This test fails if the URL ever leaves the sitemap; the e2e
journey fails if the route ever leaves the app.
"""
xml = place_sitemaps["places-1.xml"]
assert ("<loc>https://www.schoolcompare.co.uk"
"/schools/authority/essex/primary</loc>") in xml
# And never in the town namespace, which is a different set of schools.
assert "/schools/essex/primary" not in xml
+157
View File
@@ -0,0 +1,157 @@
"""Tests for school autosuggest (spec 2026-08-26)."""
from backend import data_loader
class _FakeDocs:
def __init__(self, hits, explode=False):
self._hits = hits
self._explode = explode
self.last_params = None
def search(self, params):
self.last_params = params
if self._explode:
raise RuntimeError("typesense is down")
return {"hits": [{"document": d} for d in self._hits]}
class _FakeClient:
def __init__(self, hits, explode=False):
self.docs = _FakeDocs(hits, explode)
self.collections = {"schools": type("C", (), {"documents": self.docs})()}
_HIT = {
"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school",
}
def _use(monkeypatch, client):
monkeypatch.setattr(data_loader, "_get_typesense_client", lambda: client)
def test_returns_the_fields_a_suggestion_needs(monkeypatch):
# Local authority is not decoration: there are many schools called
# "St Mary's", and a list without it cannot be chosen between.
_use(monkeypatch, _FakeClient([_HIT]))
out = data_loader.suggest_schools_typesense("breck")
assert out == [{
"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school",
}]
def test_a_missing_optional_field_becomes_an_empty_string(monkeypatch):
# phase and school_type are optional in the Typesense schema. A missing
# key must not KeyError in the keystroke path.
_use(monkeypatch, _FakeClient([{"urn": 1, "school_name": "X",
"local_authority": "Y", "postcode": "Z"}]))
out = data_loader.suggest_schools_typesense("x")
assert out[0]["phase"] == "" and out[0]["school_type"] == ""
def test_typesense_unavailable_gives_no_suggestions_rather_than_raising(monkeypatch):
_use(monkeypatch, None)
assert data_loader.suggest_schools_typesense("anything") == []
def test_a_typesense_error_gives_no_suggestions_rather_than_raising(monkeypatch):
_use(monkeypatch, _FakeClient([], explode=True))
assert data_loader.suggest_schools_typesense("anything") == []
def test_the_limit_is_passed_through_and_clamped(monkeypatch):
client = _FakeClient([])
_use(monkeypatch, client)
data_loader.suggest_schools_typesense("x", limit=500)
assert client.docs.last_params["per_page"] == 20
def _client(monkeypatch, rows, *, blow_up_dataframe=False):
from fastapi.testclient import TestClient
from backend import app as app_module
monkeypatch.setattr(app_module, "suggest_schools_typesense",
lambda q, limit=8: rows)
if blow_up_dataframe:
def _boom():
raise AssertionError("the suggest path must not load the DataFrame")
monkeypatch.setattr(app_module, "load_school_data", _boom)
monkeypatch.setattr(app_module, "load_latest_school_data", _boom)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_the_endpoint_returns_suggestions(monkeypatch):
body = _client(monkeypatch, [_HIT]).get("/api/suggest?q=breck").json()
assert body["suggestions"][0]["school_name"] == "Brecknock Primary School"
def test_the_endpoint_never_touches_the_dataframe(monkeypatch):
"""The whole reason this is not a mode of /api/schools.
That endpoint filters and sorts 25,000 rows of pandas per query, holding
the GIL. Per keystroke, that is the cost this endpoint exists to avoid.
"""
res = _client(monkeypatch, [_HIT], blow_up_dataframe=True).get("/api/suggest?q=breck")
assert res.status_code == 200
assert res.json()["suggestions"]
def test_a_one_character_query_returns_nothing_and_does_not_error(monkeypatch):
# The keystroke path never errors on ordinary input.
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=b")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_a_blank_query_returns_nothing_and_does_not_error(monkeypatch):
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_typesense_down_is_an_empty_list_not_a_500(monkeypatch):
res = _client(monkeypatch, []).get("/api/suggest?q=breck")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_the_response_is_cacheable(monkeypatch):
# Prefix queries repeat enormously across users, and school names change
# once a year. Without this the endpoint pays full price every keystroke.
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=breck")
assert "s-maxage" in res.headers.get("cache-control", "")
assert res.headers.get("etag")
def test_a_malformed_urn_does_not_raise(monkeypatch):
"""The docstring promises "never raises"; the parsing loop sat outside the
try, so int(None) or int("abc") would have turned a keystroke into a 500.
Typesense declares urn as int32, so this should be unreachable — but the
contract is what the caller relies on, and a search index is a separate
system that can be reindexed by something other than this code.
"""
_use(monkeypatch, _FakeClient([{"urn": None, "school_name": "X",
"local_authority": "Y", "postcode": "Z"}]))
assert data_loader.suggest_schools_typesense("x") == []
def test_a_malformed_row_does_not_discard_the_good_ones(monkeypatch):
# One bad document must not blank the whole dropdown.
_use(monkeypatch, _FakeClient([
{"urn": "not-a-number", "school_name": "Bad", "local_authority": "Y",
"postcode": "Z"},
_HIT,
]))
out = data_loader.suggest_schools_typesense("x")
assert [r["urn"] for r in out] == [100010]
def test_a_hit_with_no_document_does_not_raise(monkeypatch):
_use(monkeypatch, _FakeClient([{}]))
assert data_loader.suggest_schools_typesense("x") == []
+76 -4
View File
@@ -8,8 +8,31 @@ from backend import data_loader
from backend.data_loader import get_supplementary_data_batch
def _sort_key(criterion):
"""(column name, descending) for a SQLAlchemy order_by argument.
A bare column (Model.year) arrives as an InstrumentedAttribute carrying
.key; Model.year.desc() wraps it in a UnaryExpression whose column sits on
.element.
"""
name = getattr(criterion, "key", None)
if name is not None:
return name, False
element = getattr(criterion, "element", None)
name = getattr(element, "key", None)
return name, "DESC" in str(criterion).upper()
class _FakeQuery:
"""Records that a query ran and serves canned rows filtered by an in-list."""
"""Records that a query ran and serves canned rows filtered by an in-list.
order_by is honoured rather than ignored. The batch loader picks a row per
URN by position — first for "latest Ofsted", last for "latest cut-off
distance" — which is only correct because the database returned them
sorted. A double that drops the ORDER BY makes those picks depend on
fixture insertion order instead, so the test would pass with the sort
reversed or removed and prove nothing about the query.
"""
def __init__(self, recorder, model_name, rows):
self._rec = recorder
@@ -19,7 +42,20 @@ class _FakeQuery:
def filter(self, *args, **kwargs):
return self
def order_by(self, *args, **kwargs):
def order_by(self, *criteria):
for crit in reversed(criteria): # reversed = stable multi-key sort
name, descending = _sort_key(crit)
if not name:
continue
values = [getattr(r, name, None) for r in self._rows]
# Only sort on plainly comparable values. Some fixtures stand dates
# up as namespace objects, which raise on <; leaving those in their
# given order matches what the real query would produce for them.
if not all(isinstance(v, (int, float, str)) for v in values):
continue
self._rows = sorted(
self._rows, key=lambda r: getattr(r, name), reverse=descending
)
return self
def all(self):
@@ -69,6 +105,13 @@ def _adm_row(urn, year):
)
def _dist_row(urn, year, distance_m, route_count=1):
return types.SimpleNamespace(
urn=urn, year=year, distance_m=distance_m, route_count=route_count,
la_name="Camden", distance_unit_raw="miles", source_file="camden/guide.pdf",
)
def test_one_query_per_table_and_latest_row_per_urn():
rows = {
# URN 1 has two Ofsted rows; the batch must keep the most recent (2023).
@@ -78,19 +121,34 @@ def test_one_query_per_table_and_latest_row_per_urn():
_ofsted_row(2, "2021-06-01", 1),
],
"FactAdmissions": [_adm_row(1, 202526), _adm_row(1, 202627), _adm_row(2, 202627)],
# URN 1 has three years of cut-offs; only the most recent is served.
# Deliberately not in year order — the ordering is the query's job.
"FactAdmissionDistance": [
_dist_row(1, 2026, 529.47),
_dist_row(1, 2024, 772.49),
_dist_row(1, 2025, 1421.05),
_dist_row(2, 2023, 2029.38, route_count=4),
],
"FactPupilCharacteristics": [],
"FactDeprivation": [],
"FactFinance": [],
"FactKs4Destinations": [],
"FactKs5Destinations": [],
}
session = _FakeSession(rows)
out = get_supplementary_data_batch(session, [1, 2])
# Exactly one query per table — five total, regardless of two URNs.
# Exactly one query per table — eight total, regardless of two URNs.
assert sorted(session.queries) == [
"FactAdmissions", "FactDeprivation", "FactFinance",
"FactAdmissionDistance", "FactAdmissions", "FactDeprivation",
"FactFinance", "FactKs4Destinations", "FactKs5Destinations",
"FactOfstedInspection", "FactPupilCharacteristics",
]
# A school with no destination rows gets null, not an empty shell — the
# frontend renders the section from the block's presence.
assert out[1]["destinations"] is None
# Latest Ofsted kept per URN
assert out[1]["ofsted"]["overall_effectiveness"] == 2
assert out[2]["ofsted"]["overall_effectiveness"] == 1
@@ -100,6 +158,18 @@ def test_one_query_per_table_and_latest_row_per_urn():
assert out[1]["admissions"]["year"] == 202627
assert out[2]["admissions_history"] == [{**out[2]["admissions_history"][0]}]
# Cut-off distance: the latest year only. Earlier years stay in the mart
# but are held back as a paid feature, and this API is public — serving
# them here would hand them to anyone reading the response. The fixture
# rows are deliberately out of order, so "latest" only comes out right if
# the query's ORDER BY is doing the work.
assert out[1]["admission_distance"]["year"] == 2026
assert "admission_distance_history" not in out[1]
assert out[1]["admission_distance"]["distance_m"] == 529.47
# route_count travels with the figure — the page needs it to say the
# distance is the furthest of several bands rather than the only one.
assert out[2]["admission_distance"]["route_count"] == 4
# Empty tables degrade to the null block, not a crash
assert out[1]["census"] is None and out[1]["deprivation"] is None
@@ -109,3 +179,5 @@ def test_single_wrapper_matches_batch(monkeypatch):
single = data_loader.get_supplementary_data(session, 5)
assert single["ofsted"]["overall_effectiveness"] == 2
assert single["admissions_history"] == []
assert single["admission_distance"] is None
assert "admission_distance_history" not in single
-26
View File
@@ -1,26 +0,0 @@
"""
Schema versioning for database migrations.
HOW TO USE:
- Bump SCHEMA_VERSION when making changes to database models
- This triggers an automatic full data reimport on next app startup
WHEN TO BUMP:
- Adding/removing columns in models.py
- Changing column types or constraints
- Modifying CSV column mappings in schemas.py
- Any change that requires fresh data import
"""
# Current schema version - increment when models change
SCHEMA_VERSION = 6
# Changelog for documentation
SCHEMA_CHANGELOG = {
1: "Initial schema with School and SchoolResult tables",
2: "Added pupil absence fields (reading, maths, gps, writing, science)",
3: "Added supplementary data tables: ofsted, parent_view, census, admissions, sen_detail, phonics, deprivation, finance; GIAS columns on schools",
4: "Added Ofsted Report Card columns to ofsted_inspections (new framework from Nov 2025)",
5: "Apply ALTER TABLE additions for RC columns missed by create_all on existing tables",
6: "Removed the Ofsted Parent View feature: dropped fact_parent_view table and model",
}
+29 -126
View File
@@ -1,134 +1,37 @@
# SchoolCompare.co.uk - Project Context
# SchoolCompare project context
## Overview
## Maintained documentation
SchoolCompare is a web application for comparing UK primary school (KS2) performance data. It allows users to:
- Search and browse schools by name, location (postcode), or local authority
- Compare multiple schools side-by-side with charts and tables
- View school rankings by various KS2 metrics
- See historical performance trends across years
Read [README.md](README.md), [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) and
[docs/DEVELOPMENT.md](docs/DEVELOPMENT.md) for the current implementation.
[docs/LEGACY_CODE.md](docs/LEGACY_CODE.md) records obsolete paths and deliberate
compatibility code. Historical design documents are not current setup instructions.
## Architecture
## Architecture constraints
### Backend (Python/FastAPI)
- **Framework**: FastAPI with uvicorn
- **Database**: PostgreSQL with SQLAlchemy ORM
- **Data Source**: UK Government "Compare School Performance" CSV downloads
Key files:
- `backend/app.py` - Main FastAPI application, API routes
- `backend/config.py` - Configuration via pydantic-settings (env vars, .env file)
- `backend/database.py` - SQLAlchemy engine, session management
- `backend/models.py` - Database models (School, SchoolResult)
- `backend/data_loader.py` - Data queries, geocoding, legacy DataFrame compatibility
- `backend/schemas.py` - Column mappings, metric definitions, LA code mappings
### Frontend (Vanilla JS)
- Single-page application with hash-based routing
- Chart.js for data visualization
- No build step required
Key files:
- `frontend/index.html` - Main HTML structure
- `frontend/app.js` - All application logic, API calls, rendering
- `frontend/styles.css` - Styling (CSS variables, responsive design)
### Database Schema
```
schools school_results
├── id (PK) ├── id (PK)
├── urn (unique, indexed) ├── school_id (FK → schools.id)
├── school_name ├── year (indexed)
├── local_authority ├── rwm_expected_pct
├── school_type ├── reading_expected_pct
├── postcode ├── ... (all KS2 metrics)
├── latitude, longitude └── unique(school_id, year)
└── results → SchoolResult[]
```
## Configuration
Environment variables (or `.env` file):
- `DATABASE_URL` - PostgreSQL connection string (default: `postgresql://schoolcompare:schoolcompare@localhost:5432/schoolcompare`)
- `HOST`, `PORT` - Server binding (default: `0.0.0.0:80`)
- `ALLOWED_ORIGINS` - CORS origins
## Running Locally
1. Start PostgreSQL:
```bash
docker compose up -d db
```
2. Run migration to import CSV data:
```bash
python scripts/migrate_csv_to_db.py --drop
# Add --geocode to geocode postcodes (slower, adds lat/long)
```
3. Start the app:
```bash
uvicorn backend.app:app --host 0.0.0.0 --port 8000
```
## Docker Deployment
```bash
docker compose up -d
```
This starts:
- `db` - PostgreSQL 16 with persistent volume
- `app` - FastAPI application on port 80
## Data
- Source: UK Government Compare School Performance downloads
- Location: `data/` directory with year folders (e.g., `2023-2024/england_ks2final.csv`)
- The `scripts/download_data.py` can fetch data from the government website
## Key Features
- **Location Search**: Enter postcode to find nearby schools (uses postcodes.io API)
- **Multi-school Comparison**: Select multiple schools, view metrics across years
- **Rankings**: Top schools by any KS2 metric, filterable by local authority
- **Variability Analysis**: Shows standard deviation of scores across years
## API Endpoints
- `GET /api/schools` - List/search schools (supports pagination, location search)
- `GET /api/schools/{urn}` - School details with all yearly data
- `GET /api/compare?urns=123,456` - Compare multiple schools
- `GET /api/rankings` - School rankings by metric
- `GET /api/filters` - Available filter options (LAs, types, years)
- `GET /api/metrics` - Metric definitions (single source of truth)
- `GET /api/data-info` - Database stats
- Next.js serves the public UI. FastAPI serves school data from dbt-built `marts.*`.
The backend does not create school tables or import CSVs at startup.
- School coverage spans England and multiple phases, not only primary schools in
Wandsworth and Merton.
- `/api/*` belongs to the FastAPI proxy. Payload uses `/cms-api` and `/admin`.
- Payload runs inside Next.js, with its own `payload` schema and persistent media.
Keep CMS migrations independent of school-data transformations.
- Public and Payload route groups have separate root layouts. Do not introduce
`app/layout.tsx`. Keep site-wide metadata files at the `app/` root.
- Builds must succeed with `DATABASE_URL` unset. Do not call `getCachedPayload()`
at module scope or add DB-backed `generateStaticParams`.
- After changing CMS fields/editors, run `npm run generate:importmap` and commit
the generated import map. See `nextjs-app/docs/PUBLISHING.md`.
- The backend and pipeline GIAS dictionary copies are generated together; preserve
their parity. Tests enforce it.
## SDLC
Full details in `docs/DEPLOY.md`. The short version:
Follow [docs/DEPLOY.md](docs/DEPLOY.md).
- **Never push to `main` directly.** Work on a feature branch and open a PR;
branch protection requires the PR checks (typecheck, tests, builds, AI review)
to pass before merge.
- Merging to `main` deploys automatically **to staging only**: images are
built once, deployed to the staging Portainer stack, and verified by the
Playwright journeys in `e2e/`. Production is a second, manual approval:
the "Promote to Production (manual)" workflow in Gitea Actions, run after
testing the feature on staging. It refuses commits whose staging E2E gate
isn't green. Never trigger it yourself — promotion is the human's call.
- If you change user-facing behaviour, update or extend the `e2e/` journey
tests in the same PR — they gate whether staging is fit for human testing
and whether a commit is promotable.
## Recent Changes
- Added staging environment + automated staging→prod pipeline (Gitea Actions)
- Migrated from CSV file storage to PostgreSQL database
- Added location-based search using postcode geocoding
- Added local authority filter to rankings
- Improved frontend with featured schools, loading states, API caching
# Important
- Do not attempt to start a local server to test the application, it does not work
- Never push directly to `main`. Use a feature branch and a PR with passing checks.
- Merges deploy staging only. Production promotion is a separate human decision;
do not trigger the promotion workflow yourself.
- Update E2E journeys in the same PR when changing user-facing behaviour.
- Do not attempt to start a local server to test the application; use unit checks
and the configured integration environment.
+48 -2
View File
@@ -16,7 +16,16 @@
# ADMIN_API_KEY — Backend admin API key
# TYPESENSE_API_KEY — Typesense admin API key
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: development
# PAYLOAD_SECRET — Payload CMS encryption secret. REQUIRED: long and
# random, and DIFFERENT from production's. Sharing
# it would let a staging session authenticate
# against production.
# AIRFLOW_ADMIN_USER — Airflow admin username (default: admin)
# AIRFLOW_ADMIN_PASSWORD — Airflow admin password. REQUIRED: the api-server
# refuses to start without it, rather than falling
# back to a generated one that changes on restart.
# STAGING_DB_IP — macvlan IP for staging Postgres (default 10.0.1.190)
# STAGING_FRONTEND_IP — macvlan IP for staging frontend (default 10.0.1.151)
@@ -55,6 +64,12 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- unleash_cache:/app/.unleash
depends_on:
sc_database:
condition: service_healthy
@@ -78,9 +93,20 @@ services:
- FASTAPI_URL=http://backend:80/api
- TYPESENSE_URL=http://typesense:8108
- TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme}
# Payload CMS runs inside this container, in the `payload` schema of the
# staging database. Staging has its own stack, its own Postgres and its
# own admin account — never production's.
- DATABASE_URL=postgresql://${DB_USERNAME}:${DB_PASSWORD}@sc_database:5432/${DB_DATABASE_NAME}
- PAYLOAD_SECRET=${PAYLOAD_SECRET:?set PAYLOAD_SECRET in the staging Portainer stack environment}
volumes:
# Portainer prefixes volume names with the stack name, so this is
# automatically isolated from production's media.
- payload_media:/app/media
depends_on:
backend:
condition: service_healthy
sc_database:
condition: service_healthy
networks:
backend: {}
macvlan:
@@ -116,7 +142,23 @@ services:
airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:staging
container_name: sc_staging_airflow_api
command: airflow api-server --port 8080
# The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports:
- "8081:8080"
environment:
@@ -128,6 +170,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-staging-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "${AIRFLOW_ADMIN_USER:-admin}:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:?set AIRFLOW_ADMIN_PASSWORD in the Portainer stack environment}
AIRFLOW__LOGGING__BASE_LOG_FOLDER: /opt/airflow/logs
PG_HOST: sc_database
PG_PORT: "5432"
@@ -212,3 +256,5 @@ volumes:
postgres_data:
typesense_data:
airflow_logs:
unleash_cache:
payload_media:
+73
View File
@@ -0,0 +1,73 @@
# Portainer Stack Definition for School Compare — UNLEASH (feature flags)
#
# Deploy as a *separate* Portainer stack ("schoolcompare-unleash"), alongside
# the production and staging stacks. It deliberately belongs to neither: a
# staging redeploy must not be able to disturb production's flag state, and a
# production redeploy must not disturb staging's.
#
# One instance serves both environments. Open-source Unleash ships with
# `development` and `production` environments and environment-scoped client
# tokens, so the same flag holds independent state in each — which is what
# lets a feature be on in staging, where the E2E journeys exercise it, while
# production stays dark.
#
# Portainer environment variables (set in Portainer UI -> Stack -> Environment):
# UNLEASH_DB_PASSWORD — PostgreSQL password for the Unleash database
# UNLEASH_ADMIN_PASSWORD — initial admin password for the Unleash UI
# UNLEASH_IP — macvlan IP for the Unleash server (default 10.0.1.152)
services:
# ── PostgreSQL (Unleash's own; nothing else uses it) ──────────────────
unleash_db:
container_name: sc_unleash_postgres
image: postgres:16-alpine
environment:
POSTGRES_USER: unleash
POSTGRES_PASSWORD: ${UNLEASH_DB_PASSWORD}
POSTGRES_DB: unleash
volumes:
- unleash_postgres_data:/var/lib/postgresql/data
networks:
- unleash
healthcheck:
test: ["CMD-SHELL", "pg_isready -U unleash"]
interval: 10s
timeout: 5s
retries: 5
start_period: 10s
restart: unless-stopped
# ── Unleash server (UI + client API on 4242) ──────────────────────────
unleash:
container_name: sc_unleash
image: unleashorg/unleash-server:6
environment:
DATABASE_URL: postgres://unleash:${UNLEASH_DB_PASSWORD}@unleash_db:5432/unleash
DATABASE_SSL: "false"
INIT_ADMIN_API_TOKENS: ""
UNLEASH_DEFAULT_ADMIN_PASSWORD: ${UNLEASH_ADMIN_PASSWORD}
depends_on:
unleash_db:
condition: service_healthy
networks:
unleash: {}
macvlan:
ipv4_address: ${UNLEASH_IP:-10.0.1.152}
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:4242/health || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 30s
restart: unless-stopped
networks:
unleash:
driver: bridge
macvlan:
external:
name: macvlan
volumes:
unleash_postgres_data:
+48 -2
View File
@@ -7,7 +7,15 @@
# ADMIN_API_KEY — Backend admin API key
# TYPESENSE_API_KEY — Typesense admin API key
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: production
# PAYLOAD_SECRET — Payload CMS encryption secret. REQUIRED: long and
# random. Changing it invalidates every admin
# session. Staging MUST use a different value.
# AIRFLOW_ADMIN_USER — Airflow admin username (default: admin)
# AIRFLOW_ADMIN_PASSWORD — Airflow admin password. REQUIRED: the api-server
# refuses to start without it, rather than falling
# back to a generated one that changes on restart.
services:
@@ -44,6 +52,12 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- unleash_cache:/app/.unleash
depends_on:
sc_database:
condition: service_healthy
@@ -67,9 +81,21 @@ services:
- FASTAPI_URL=http://backend:80/api
- TYPESENSE_URL=http://typesense:8108
- TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme}
# Payload CMS runs inside this container. It reaches Postgres over the
# `backend` network and keeps its tables in the `payload` schema, so no
# pipeline operation on `public` can touch blog content.
- DATABASE_URL=postgresql://${DB_USERNAME}:${DB_PASSWORD}@sc_database:5432/${DB_DATABASE_NAME}
# Same :? form as AIRFLOW_ADMIN_PASSWORD: refuse to start rather than
# boot with an empty secret and silently accept forged sessions.
- PAYLOAD_SECRET=${PAYLOAD_SECRET:?set PAYLOAD_SECRET in the Portainer stack environment}
volumes:
# Blog images. Not reproducible from the pipeline — must be backed up.
- payload_media:/app/media
depends_on:
backend:
condition: service_healthy
sc_database:
condition: service_healthy
networks:
backend: {}
macvlan:
@@ -105,7 +131,23 @@ services:
airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:prod
container_name: schoolcompare_airflow_api
command: airflow api-server --port 8080
# The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports:
- "8080:8080"
environment:
@@ -117,6 +159,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "${AIRFLOW_ADMIN_USER:-admin}:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:?set AIRFLOW_ADMIN_PASSWORD in the Portainer stack environment}
AIRFLOW__LOGGING__BASE_LOG_FOLDER: /opt/airflow/logs
PG_HOST: sc_database
PG_PORT: "5432"
@@ -201,3 +245,5 @@ volumes:
postgres_data:
typesense_data:
airflow_logs:
unleash_cache:
payload_media:
+23 -1
View File
@@ -36,6 +36,10 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- ./data:/app/data:ro
depends_on:
@@ -101,7 +105,23 @@ services:
airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:latest
container_name: schoolcompare_airflow_api
command: airflow api-server --port 8080
# The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports:
- "8080:8080"
environment: &airflow-env
@@ -113,6 +133,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "admin:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:-admin}
PG_HOST: db
PG_PORT: "5432"
PG_USER: schoolcompare
+125
View File
@@ -0,0 +1,125 @@
# Architecture
This describes the implementation as reviewed on 2026-09-14. It distinguishes
current behaviour from improvements still to be implemented.
## Request flow
```text
Browser → Next.js public routes
├─ /api/* proxy → FastAPI → cached DataFrames / PostgreSQL marts
│ ├─ Typesense (search and suggestions)
│ └─ postcodes.io (postcode lookup)
└─ /admin, /cms-api, /blog → Payload → payload schema + media volume
Next.js server rendering → FastAPI directly through FASTAPI_URL
```
`nextjs-app/lib/api.ts` contains typed fetch wrappers and revalidation defaults.
The proxy is `nextjs-app/app/(frontend)/api/[...path]/route.ts`. Payload uses
`/cms-api` so its routes do not collide with the FastAPI proxy. The proxy denies
`/api/flags`; server-side rendering reads flags directly from FastAPI.
## Data ownership
| Layer | Owner and role |
|---|---|
| Source data | GIAS, DfE EES, Ofsted, finance, deprivation and council admission-distance sources |
| `raw` | Singer taps and the PostgreSQL target configured in `pipeline/meltano.yml` |
| Staging/intermediate/marts | dbt models in `pipeline/transform`; marts are materialized tables |
| `marts.dim_school`, `marts.dim_location` | School identity and location, filtered to supported England establishments |
| `marts.fact_*` | Performance and supplementary datasets; coverage and years vary |
| Typesense `schools` alias | Search documents built by `pipeline/scripts/sync_typesense.py` |
| `payload` | CMS collections and migrations in `nextjs-app/`; independent of dbt |
| Media volume | Uploaded blog media; requires backup and cannot be regenerated from school datasets |
`backend/models.py` maps existing marts for reading. It does not create the school
schema. There is no startup schema-version migration or CSV reimport. Payload's
`nextjs-app/migrations/` is active and must not be confused with the removed
legacy backend migration code.
Coordinates normally come from GIAS British National Grid coordinates transformed
by PostGIS in `dim_location.sql`. `pipeline/scripts/geocode_postcodes.py` is a
manual fallback utility, not a task wired into the current school-data DAG.
Backend postcode searches also use postcodes.io; that lookup does not populate
school coordinates in the database.
## Backend boundaries
- `app.py`: routes, middleware, search filtering, sitemap/place publication and response assembly.
- `data_loader.py`: SQL loading, process-local DataFrame caches, Typesense calls,
postcode lookups, supplementary queries and benchmark calculation.
- `database.py`: synchronous SQLAlchemy engine and sessions.
- `schemas.py`: metric definitions, column mappings and display metadata; despite
its name this is not a collection of Pydantic API response models.
- `places.py` and `localities.py`: place registry and curated locality information.
- `flags.py`: Unleash-backed feature flags, disabled when no server is configured.
- `gias_codes.py` / `ofsted_codes.py`: source-code translation and display rules.
Search starts from a cached latest-row-per-school snapshot. Detail pages read
history from the full DataFrame and supplementary data from marts. Comparisons
batch supplementary queries across selected URNs. Async routes still contain
synchronous dependency calls; a fully asynchronous database layer is not present.
## Frontend boundaries
`app/(frontend)` owns the public root layout and pages. `app/(payload)` owns the
CMS root layout. Do not add a shared `app/layout.tsx`: these groups deliberately
have separate root layouts. Root metadata files remain in `app/`.
Server pages fetch initial data and pass it to client views. Client state uses
React hooks, URL search parameters and the comparison context/localStorage.
There is no SWR dependency. Leaflet maps are loaded through dynamic wrappers;
Chart.js renders performance and comparison charts.
`components/school/` contains detail sections, with section decisions and data
preparation in `lib/schoolSections.ts`. The nearby-schools section is selected in
`backend/nearby_schools.py` — hard filters decide eligibility (phase, provision,
selectivity, gender) and distance alone decides the order, capped per phase —
and served on `/api/schools/{urn}`. Its rules are presentation logic,
deliberately kept out of `marts.*` so they can be tuned by deploy rather than by
pipeline run. `lib/types.ts` contains manually maintained
API types. `payload-types.ts` and the Payload import map are generated artifacts.
## Publication and caching today
1. Airflow DAGs extract and validate source data, then run selected dbt builds.
2. Relevant DAGs rebuild Typesense and swap the `schools` alias.
3. They call `POST /api/admin/reload` with `X-API-Key`. It builds and validates
replacement DataFrames, places, reverse membership and sitemaps off the request
loop, then publishes them together. Failure returns 503 and preserves live data.
4. A separate weekly sitemap DAG can regenerate the derived publication from the
current DataFrame without clearing the live registry first.
GIAS is scheduled daily, Ofsted monthly, and annual datasets are manually
triggered. The DAG definitions are authoritative for selectors and dependencies.
Caches exist in several independent layers: backend DataFrames and registries,
backend HTTP Cache-Control/ETags, Next.js fetch/page revalidation, and browser or
shared HTTP caches where configured. Place fetches request a one-week revalidation
interval. HTTP ETags are computed after route execution, not before database work.
Typesense publication validates every import response and the final document
count before switching aliases. A session-scoped PostgreSQL advisory lock
serialises index reads/publication across DAGs. The previous collection remains
available for rollback; old unaliased collections are pruned after success.
Failed drafts are retained until a later successful cleanup, because an uncertain
alias-update response must never cause deletion of a potentially live index.
The backend snapshot swap is process-local and assumes the current single-worker
deployment. It is not an atomic transaction spanning PostgreSQL marts, Typesense
and Next.js caches. Next.js caches are not explicitly purged by the pipeline.
School search retrieves a relevance-ordered candidate prefix (currently capped at
1,000 URNs) before applying API filters. This keeps scoped searches useful while
putting a hard ceiling on Typesense round trips; only a dependency failure invokes
substring fallback, not a valid empty match set.
## Deployment references
See [DEPLOY.md](DEPLOY.md). PR checks include frontend typechecking/tests, backend
unit tests, image builds and AI review. Staging journeys run after merging.
Staging runs are serialised across builds, deployment and E2E. Build-stamped
frontend/backend identities are checked before and after journeys. Only then are
the captured image digests marked verified. Promotion resolves and validates the
complete verified image set before retagging production. See the runbook for
first-rollout requirements and remaining integration checks.
+159 -4
View File
@@ -19,8 +19,9 @@ PR checks (.gitea/workflows/pr-checks.yml)
▼
Stage pipeline (.gitea/workflows/deploy.yml) — automatic
1. build & push images → tags sha-<sha>, staging
2. staging Portainer webhook → wait for staging health
2. staging Portainer webhook → verify frontend/backend SHA + build ID
3. Playwright E2E journeys against staging ← gate before human testing
4. verify identity again; tag tested digests verified-<full-sha>
▼
Manual testing on staging (stx.schoolcompare.co.uk)
│ Actions → "Promote to Production (manual)" ← approval #2
@@ -28,14 +29,15 @@ Manual testing on staging (stx.schoolcompare.co.uk)
Promote pipeline (.gitea/workflows/promote.yml) — manual dispatch
1. resolve target sha (input, or latest main if empty)
2. REFUSE unless that commit's "E2E Journeys against Staging" status is green
3. retag sha-<sha> → :prod (same bytes — build once, promote the image)
3. resolve verified-<full-sha> digests, validate labels, retag digests → :prod
previous :prod saved as :prod-previous
4. prod Portainer webhook → wait for prod health
4. prod Portainer webhook → verify expected SHA + build ID
```
Key principle: **build once, promote the exact image**. Production pins `:prod`,
which only moves when a human runs the promote workflow — and the workflow
only accepts commits that passed the staging E2E gate. Nothing tags `:latest`
only accepts commits that passed the staging E2E gate and have a complete verified
image set. Nothing tags `:latest`
anymore.
## Branch & PR workflow
@@ -98,6 +100,12 @@ fail the E2E gate. That's the point: staging absorbs the risk.
pr-checks status checks (frontend, backend, builds, ai-review) to pass.
5. **Bootstrap staging data via Airflow** (no prod dump — staging populates
itself from source, exercising the pipeline image end-to-end):
- Set `AIRFLOW_ADMIN_PASSWORD` in the stack environment first. The
api-server refuses to start without it. Airflow's simple auth manager
otherwise generates a password on first start and writes it to a file, so
the login changes every time the container restarts; the stack writes that
file itself from this variable instead. `AIRFLOW_ADMIN_USER` defaults to
`admin`.
- Open the staging Airflow UI (`http://<host>:8081`) and trigger, in order:
`school_data_daily`, `school_data_monthly_ofsted`, then the manual-schedule
`school_data_annual_ees` and `school_data_annual_idaci`.
@@ -150,3 +158,150 @@ token Gitea Actions provides automatically (`secrets.GITEA_TOKEN` — no setup
needed), and fails the check only when a finding is rated
**severe** (would break prod, leak data, or corrupt data). Minor findings are
informational and never block a merge.
## Rate limiting, and the Cloudflare gap
Two independent limits protect the API:
- **Per client**, via slowapi, keyed on `CF-Connecting-IP` (falling back to
`X-Forwarded-For`, then the peer address). 60/minute by default;
`/api/suggest` gets 120/minute because typing is bursty.
- **Globally**, via `GlobalRateLimitMiddleware`: a fixed 60-second window over
all `/api/` traffic, `GLOBAL_RATE_LIMIT_PER_MINUTE` (default 3000),
independent of any client identity. Requests from `127.0.0.1` are exempt so
the container healthcheck cannot be starved into a restart loop.
### Open: the origin must only accept Cloudflare
`CF-Connecting-IP` is only meaningful for requests that actually reached the
origin through Cloudflare, and **the application cannot verify that they did**.
Anything able to reach the origin directly can set that header freely and, by
rotating it, mint a fresh rate-limit bucket per request — defeating per-client
limits on every endpoint.
The global ceiling bounds the damage to total origin capacity. It does not fix
the underlying gap, and nothing in the code can. Closing it needs one of:
- **Authenticated Origin Pulls** — Cloudflare presents a client certificate the
origin requires, so non-Cloudflare traffic is refused at TLS.
- **An origin firewall** restricted to Cloudflare's published IP ranges.
Until one is in place, treat per-client limits as protection against accidents
and ordinary load, not against a determined caller.
## Feature flags (Unleash)
Flag state lives in a self-hosted Unleash instance, deployed as its own
Portainer stack from `docker-compose.portainer.unleash.yml`. It is separate
from the application stacks on purpose — redeploying staging must not be able
to disturb production's flags.
The flags themselves are declared in `backend/flags.py`. Unleash holds the
state; the registry holds the list. A flag in the UI that is not in the
registry is orphaned and nothing reads it.
### First-time setup
1. Deploy the stack in Portainer. Set `UNLEASH_DB_PASSWORD`,
`UNLEASH_ADMIN_PASSWORD` and (optionally) `UNLEASH_IP`.
2. Log in to the UI at `http://<UNLEASH_IP>:4242` as `admin`.
3. Create one **client** API token per environment:
- `schoolcompare-staging`, environment **development**
- `schoolcompare-prod`, environment **production**
Client tokens, not admin tokens — the backend only reads.
4. Put each token in the matching Portainer stack's `UNLEASH_API_TOKEN`
variable, and set `UNLEASH_URL` to `http://<UNLEASH_IP>:4242/api`.
5. Redeploy the application stacks.
### Adding a flag to Unleash
**Unleash does not create flags by itself.** The SDK reads definitions from the
server and never registers anything, and metrics for a flag the server has
never heard of are discarded. So a flag declared in `backend/flags.py` will be
evaluated on every request, stay `False` forever, and never appear in the UI
until someone creates it there by hand.
For each flag in the registry, create one in Unleash with:
- **Name** — character for character what `backend/flags.py` declares.
snake_case, no hyphens or spaces. A typo produces a flag that looks correct
in the UI and is read by nothing.
- **Type** — Release. No strategies, constraints or variants: these are plain
on/off switches, by design.
### Turning a feature on
Toggle the flag in the environment matching the stack you mean: **development**
for staging, **production** for prod. The token in each stack is scoped to one
environment, so toggling the other one has no visible effect.
The SDK refreshes every 15 seconds, so the API reflects the change almost at
once; the pages follow on their own schedule, below.
A flip reaches school pages within about five minutes and place pages within
the hour. Next's ISR does the propagating — it revalidates a route at the
*lowest* `revalidate` among that route's fetches, which is 300s for
`/school/[slug]` and 3600s for the place pages. There is no webhook, and
adding one would only be worth it if flips ever needed to be instant.
### When Unleash is unreachable
Every flag evaluates to `False` and the site serves as though nothing were
switched on. That is deliberate — an unfinished feature staying hidden is the
safe direction — but it means a *released* feature disappears if a backend
container cold-starts with an empty cache while Unleash is down. The SDK's
disk cache is on a named volume so restarts keep last-known state, and flags
are removed from the code within 90 days (enforced by a test), which bounds
how long any feature is exposed to this.
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
attempted. That is the correct behaviour for local development and CI, and it
means the test suites need no flag server.
## Release identity and the P1 reliability gate
Every staging run creates a random build ID before building its three images.
Each image carries the commit and build ID as labels. Frontend/backend images
also contain a build-time JSON file; environment overrides cannot rewrite it.
`/release.json` returns both identities with `Cache-Control: no-store`. It fails
with 503 when either identity cannot be read. FastAPI's internal endpoint is
`/api/release`.
The entire staging workflow shares one concurrency group, with cancellation
disabled. This needs Gitea 1.26 or newer, where workflow concurrency is supported
([release notes](https://blog.gitea.com/release-of-1.26.0/)); the configured server
reported 1.27.3 during this change. Do not run the workflow on an older server
that ignores the concurrency key. Manual deployments outside this workflow must
also avoid changing staging during journeys.
The gate checks both identities before and after Playwright. It then validates
labels on the captured build output digests and tags them `verified-<full-sha>`.
The manual promotion script resolves all three verified tags to immutable digests
and confirms one matching commit/build ID before moving any `:prod` tag. It polls
production for that same identity using a locally saved release manifest.
A registry error can still interrupt the three tag writes; the Portainer webhook
only runs after successful promotion, and rerunning promotion resolves the full
verified set again. There is no cross-registry atomic tag transaction.
**First rollout:** old green commits without verified tags/build identities are
not promotable through this gate. Build and test a commit containing the new
workflow first. The release route must be reachable through the configured
`STAGING_BASE_URL`/`PROD_BASE_URL`; it deliberately avoids the public staging
`/api` proxy limitation. No new deployment secret is required.
`scripts/ci/release.py` implements identity polling and digest verification.
The poller identifies itself as `SchoolCompare-Release-Check/1.0`: the public
staging proxy has returned HTTP 403 to Python's default urllib user agent even
while the release endpoint was healthy. It logs changes in HTTP/connection
failures or observed release identities, and includes the last observation in
the timeout error. If verification fails, use that observation to distinguish
proxy rejection (403), an unavailable release endpoint (503), and containers
still reporting an older SHA/build ID. Check the configured base URL from the
CI runner; a successful request from another machine does not establish runner
connectivity. Do not bypass identity verification to unblock a deployment.
Its mocked tests run in PR checks alongside backend and index-publication tests.
The new Playwright journeys also check deployed identity and stale pagination.
Local unit checks do not validate registry credentials, Portainer behaviour,
proxy routing or a deployed image; those require the staging run. Production
promotion remains a separate human action.
+94
View File
@@ -0,0 +1,94 @@
# Development and validation
## Prerequisites and environment boundaries
Use a feature branch. The deployed stack is the integration environment; do not
assume a local server can run from a fresh checkout. This cleanup did not start
local servers or provision databases. Unit tests use fixtures and mocks.
The current versions are not yet aligned:
| Component | Container | PR checks |
|---|---|---|
| Backend | Python 3.11 | Python 3.12 |
| Frontend | Node 24 | Node 22 |
| Pipeline | Python 3.13 | Pipeline image build |
Use the component's container version when reproducing deployment behaviour.
The backend dependency pins predate Python 3.14; do not assume the system Python
can install or run them. Version alignment is a separate maintenance task.
## Frontend checks
```sh
cd nextjs-app
npm ci
npm run typecheck
npm test -- --runInBand
```
`npm run build` is the production build check. There is no `lint` script.
Tests live in `__tests__/` and use Jest/React Testing Library. These checks do not
prove that live PostgreSQL queries, Typesense or a deployed proxy work.
The frontend `.env.example` documents runtime variables. Browser traffic normally
uses `/api`; `FASTAPI_URL` is an absolute server-side URL ending in `/api`.
Payload additionally needs `DATABASE_URL` and `PAYLOAD_SECRET` when used at runtime.
Never commit credentials or real `.env` files.
## Backend checks
From the repository root, using an available Python 3.11 or 3.12 interpreter:
```sh
python3.11 -m venv /tmp/schoolcompare-backend-venv
/tmp/schoolcompare-backend-venv/bin/python -m pip install -r requirements.txt pytest 'httpx<0.28' pyyaml
/tmp/schoolcompare-backend-venv/bin/python -m pytest backend/tests pipeline/tests scripts/ci/tests -q
```
Substitute `python3.12` if matching PR CI. The test dependencies above match the
current workflow; they are not yet captured in a dedicated development lockfile.
Backend configuration is defined in `backend/config.py`; `.env.example` documents
commonly used values. `ALLOWED_ORIGINS` uses a JSON array, not a comma-separated string.
## Data and pipeline work
The app needs populated `marts.*` tables. A new Postgres instance alone is not a
working school-data environment. Use the existing managed pipeline or an approved
snapshot; the removed CSV importer cannot build the current schema.
The pipeline container includes Meltano, dbt/Postgres, Airflow and the custom taps.
Airflow commands/selectors live in `pipeline/dags/`. Schema tests live in
`pipeline/transform/tests/` and model YAML files. Run the relevant `dbt build`
selector in an isolated data environment for model changes; it writes tables and
is not a read-only smoke test. Prefer `python -m dbt.cli.main` as the DAGs do.
GIAS dictionaries are generated together by
`pipeline/scripts/generate_gias_codes.py`. The backend and pipeline copies are
intentional; `backend/tests/test_gias_codes.py` checks that they stay identical.
For Payload collection/editor changes, run `npm run generate:importmap` in
`nextjs-app/` and include the generated map. Preserve CMS migrations and the
separate `payload` schema. See [publishing](../nextjs-app/docs/PUBLISHING.md).
## End-to-end checks
Against an existing, authorised test environment:
```sh
cd e2e
npm ci
npx playwright install chromium
BASE_URL=https://your-test-environment.example npx playwright test
```
The suite does not start a web server. CI installs Chromium with system dependencies
and runs against staging. Use the configured staging target: `docs/DEPLOY.md`
records the public staging proxy limitation. User-visible behaviour changes should
update the corresponding journeys.
## Before requesting review
Run checks relevant to the change, inspect `git diff --check`, and report checks
that could not run. Do not publish or promote as part of local validation.
[DEPLOY.md](DEPLOY.md) documents the PR and human promotion gates.
+89
View File
@@ -0,0 +1,89 @@
# Legacy and unused-code inventory
Reviewed 2026-09-14. This inventory records source evidence, not production usage
telemetry. A command with no repository caller may still be run manually or from
an external scheduler. Historical specs and prototypes are not runtime imports.
## Method and scope
Searched backend imports, tests, CLI scripts, Airflow DAGs, Meltano configuration,
Gitea workflows, Dockerfiles and documentation. For frontend candidates, inspected
TypeScript imports, re-exports, literal dynamic imports and `require` calls,
resolving relative and `@/` paths while excluding tests, dependencies and build
output. Checked candidates again with text searches including tests.
Next.js route files, generated Payload import-map entries and plugin discovery
are entry points even without ordinary imports. This is why a zero-import count
alone is not sufficient grounds for deletion. Computed imports and external
operators are outside this static audit.
## Removed in this cleanup
These names are recorded for Git-history lookup; they are no longer file links.
| Removed path or symbol | Evidence and replacement |
|---|---|
| `backend/migration.py` | Imported `School` and `SchoolResult`, which no longer exist in `backend/models.py`. Only the legacy CSV CLI imported it. Current tables are built by dbt. |
| `backend/version.py` | Only the legacy importer consumed `SCHEMA_VERSION`. FastAPI lifespan does not perform version-triggered imports. This is unrelated to active Payload migrations. |
| `scripts/migrate_csv_to_db.py` | Imported removed `init_db`/`set_db_schema_version` helpers and the obsolete models indirectly. No runtime, DAG or workflow calls it. Use the managed pipeline for current marts. |
| `scripts/geocode_schools.py` | Imported the removed `School` ORM model. No pipeline/workflow calls it. Coordinates now come from GIAS/PostGIS; a separate mart-aware manual utility remains under `pipeline/scripts/`. |
| `backend.data_loader.haversine_distance` | No callers. Search uses its inline vectorised NumPy calculation. |
| `nextjs-app/lib/api.ts: fetcher` | No callers; SWR is not installed. Application fetches use the named API wrappers. |
| `nextjs-app/lib/api.ts: kmToMiles` | No callers. `calculateDistance` remains because `CutoffMapPanel` uses it. |
The removed command files could not import successfully against the current
backend. This cleanup does not run replacements, migrate data or modify databases.
Their previous implementations remain recoverable from Git history.
## Unused candidates retained for a separate cleanup
| Candidate | Evidence | Recommended next step |
|---|---|---|
| `nextjs-app/components/LoadingSkeleton.tsx` and its CSS | No application or test imports found. | Remove together after confirming no planned use. |
| `nextjs-app/components/Pagination.tsx` and its CSS | No application or test imports found; HomeView implements load-more behaviour. | Remove as a pair if numbered pagination will not return. |
| `nextjs-app/components/SchoolCard.tsx` and its CSS | Imported by its own tests, not application code. HomeView uses SchoolRow/SecondarySchoolRow. | Decide whether to retire the card design; if removed, remove its dedicated tests as well. Passing tests do not establish runtime use. |
| `backend/database.py: get_db`, `get_db_session` | No remaining callers after removing the importer. Current code creates SessionLocal directly. | Either adopt these helpers during session-lifecycle cleanup or remove them; do not rewrite active sessions in a documentation change. |
| `backend/schemas.py: COLUMN_MAPPINGS`, `NULL_VALUES`, `LA_CODE_TO_NAME` | No remaining Python consumers found after importer removal. Other constants in this module are active. | Remove individual constants after checking external data utilities; retain the module. |
| `backend/config.py: data_dir`, `max_page_size`, `rate_limit_burst` | No active consumers found. `default_page_size` appears only in a branch that expects None, although the route supplies a concrete default. | Reconcile settings with route validation in a focused API change. |
## Legacy/manual paths requiring operational verification
| Path | Status and reason to retain for now |
|---|---|
| FastAPI `/`, `/compare`, `/rankings`, `/favicon.svg`, `/robots.txt`, and conditional `/static` | Old frontend-serving routes reference a `frontend/` directory absent from the checkout and backend image. Next.js owns these public surfaces. Removal changes externally callable routes, so first check proxy/operator usage and define replacement responses. |
| `scripts/fetch_real_data.py`, `scripts/download_data.py` | Historical standalone CSV utilities. The fetch script targets Wandsworth/Merton; neither is wired into the managed pipeline. Marked historical, retained pending confirmation of manual use. |
| `pipeline/scripts/geocode_postcodes.py` | Mart-aware postcode fallback, not called by the current DAGs. Do not confuse it with the removed legacy ORM geocoder. Verify the target schema before manual use. |
| `docker-compose.yml` | Uses unpublished `:latest` release tags and lacks frontend Payload DB/secret/media configuration. Retained as an old development topology, not recommended onboarding. |
| `nextjs-app/docker-compose.yml` | Standalone legacy recipe with old backend port assumptions and no CMS persistence setup. Retained until its consumers are checked. |
| `MIGRATION_SUMMARY.md`, `docs/superpowers/`, `mockups/` | Historical designs and prototypes. Retain as history; do not follow as current deployment instructions. |
| `scripts/sql/drop_fact_parent_view.sql` | One-off maintenance SQL. Not an application entry point; repository call-site searches cannot establish whether it is still needed operationally. |
## Active code that can look obsolete
- `backend/data_loader.py` older-mart query fallbacks are covered by backend tests
and support databases at different migration stages. Remove only after verifying
the deployed schemas in every supported environment.
- `backend/gias_codes.py` and `pipeline/scripts/gias_codes.py` are intentionally
generated copies for separate runtime images. Their parity is tested.
- `nextjs-app/migrations/`, `payload-types.ts` and the Payload import map are active
CMS artifacts, not remnants of the removed school importer.
- `get_available_years`, `get_available_local_authorities` and `get_schools_count`
in `data_loader.py` are called through `get_data_info`, which serves the backend
data-info endpoint. They are not dead functions.
- `get_supplementary_data` is an intentional single-school wrapper around the
batch implementation.
- `pipeline/transform` models named `legacy` can be active data sources: annual
DAG selectors explicitly include legacy KS2/KS4 lineage. Names alone do not
establish obsolescence.
## Suggested next passes
1. Decide the fate of the three unused UI components and remove paired assets/tests.
2. Consolidate backend session usage and remove abandoned settings/constants.
3. Verify external consumers, then retire static-serving API routes and old compose recipes.
4. Audit manual data utilities with pipeline operators before deleting them.
5. Revisit compatibility fallbacks only after documenting supported schema versions.
Validation for this cleanup should include frontend typechecking/tests, Python
syntax checks, reference searches and documentation link checks. Live database,
external scheduler and deployed route usage require separate integration evidence.
+85
View File
@@ -0,0 +1,85 @@
# Remote branch cleanup, 2026-08-20
# Restore any branch with: git push origin <sha>:refs/heads/<name>
## Deleted: fully merged into main (content is in main)
75677f4252b759ef895e7d5f7c19f8f1745bdb59 add-contact-form-footer
fa1abff642683dfd26ba88a295a0a6710d77147a chore/byline-removal-and-audit-figure
95081d38bdf87764ef5d298676c25fae4cd3b792 chore/remove-parent-view
6877abedebfc1c7d95f1f6ebe68945c62be528ab chore/staged-prod-promotion
090d5f7bec824e083d3252e2c6e636686016304d ci/frontend-checks-speedup
8c3a5cc4e9f551f0190d85357ce7741ad87f3a4c design/cohort-identity
955659580067ca8b81bd77e31e4e2554103f094d feat/allow-analytics-iframe-embed
6828f6cd4417284ea3eb6f088fa20945b8b40ed3 feat/compare-chips-two-per-row
d5cd0abfee226885119665da5d2aa8288b59217f feat/compare-data-foundation
6dd9b04b50bee146682da87efad8fc8b526251c5 feat/compare-frontend-rebuild
96d5fcf5b07b6f175b48e9b20fcb765a320a907f feat/detail-header-details-reveal
f1388ff5bd0af1409823a1e047b7ba84246a0f70 feat/gias-sixth-form-flag
eddf74745f86c9c6d9eb07d867246ff9cc90dc20 feat/hero-artwork-v2
3015c37bac6dc28db58c80fdb9235942c25b83d7 feat/hero-byline
8e763e39d17a4964cf558e51c03f044186371f6d feat/info-popover-tooltip
88c653215d520ab6e902c9de55bf27d86eeb90c3 feat/last-distance-offered
a72323874f7aebdb2e64b6d64a5febd61152d09d feat/last-distance-offered-full
c9a1892bfb0370e0672e5849cba294ddfabe3c65 feat/latest-cutoff-only
4e8df006d75d8be2a1d8529ddf855c445854cba1 feat/near-me-by-search
45ab479062c6a1639facad636fc0cc0cf0fd9155 feat/proposed-to-close-schools
3bf2e8f262cbe058fda6de6f8ea3e050224a51b9 feat/school-detail-visualisations
1f80571b1ff217dc92b660a936c02b5f49d07f0b feat/umami-heatmap-recorder
609bb923d96aa5730131463ea35d5efdd404bf96 feature/ingest-independent-schools
94151c58ea38a9256d66d15161293486a500c7a2 fix/admissions-section-height
79246edc22961c2beb3520437e9d064d3b809d10 fix/annual-dag-ks4-national-selector
59ac9c10b97e0ef1143f57fea06b324e72ac3d4a fix/chart-marker-contrast
9f8dba227c95706ca3527bd48d381e7622cc0a5e fix/compare-chart-refetch-resilience
e74d3882ce78a141fa1a57daa3102d7a58852dc3 fix/compare-expert-fixes
80176cac4db4820e76ea2a156c7a2974bee2f203 fix/compare-final-review-mustfix
f579630fab6c456e26a6984a3e8eebdbe3184202 fix/compare-mockup-drift
dc85254ad2ddf134b4434065d4762cef374d2220 fix/compare-null-year-blanks-chart
43a2c4a6bc539b621f31655aec05ef319a25f343 fix/compare-refresh-and-fetch
d677b5453365b72c81d6df2de62b1fa0d05d684d fix/daily-dag-cache-invalidation
f6bb037c471553e8195b5a8b147467ce0d07a688 fix/detail-all-through
17bd4d5a5eb0b12ca79b97db587f14f7d671e85b fix/detail-chart-truthfulness
4e6be0ce65647410b4ff74f8b763207c28920c26 fix/detail-inclusion-admissions
fdda52ff0af3fae03a4b059a655973cdbb269f91 fix/detail-ofsted-correctness
e36125b24aba254a8d15c5c33b4d2a296e691995 fix/detail-provenance-anchoring
b31e71ac884df6507f567adc046c9fc52d9d310a fix/detail-report-card-render-date
32f8a02862be6a1d49f4c3928b17fcf15d4c94cc fix/detail-trend-chart-taller
e65688d600a86818fe21ae4c61ba27e5b6ec8d7c fix/e2e-brand-assertions
3adea73ee04cdedfab54b0351878b297f72756ad fix/e2e-compare-chips-phase
06e4898c30feaedc471f97aba28ddb0d61379f4d fix/e2e-compare-samephase
9abd020967670a855e80fe5a908c8048a3aa9f14 fix/e2e-distance-locator
acec8135e1ec7c9c3c255e5b23733a0ef862b590 fix/e2e-rankings-year-pick
5944d88f0b1517ef1ef56af1b62270d2cf28e717 fix/expert-signoff-mustfixes
2433101fa08be5df6f170d41790512ffe823d33e fix/font-cascade-and-map-palette
74ca76d150deec6725259d9637ea86d7bb90c683 fix/gias-legacy-fallback
bdaa05cd542f563ef74c45307cd8f7fc465193c9 fix/hero-fallback-and-sharp
d52d384cf23d282b44e9251176f8f3402d808600 fix/hero-map-ios-fullscreen
4043270a77bbe4edb18207fa5f1d94d4747fe12f fix/hero-mobile-and-wording
22e9eb2d48b0d6623e88fd67cb6ca8e5e4074583 fix/homepage-education-accuracy
8d50afef1e8a2b621b7344eadf475b0d609ad7a2 fix/leaflet-specificity-and-font-assertion
b2b2cad5acf534ae7a667d3fb2be15efff37c4a7 fix/list-map-report-card-signal
dc21e80a5e9eebd13aaab84735642eb77cef35e6 fix/mobile-cell-name-size
e5f7f4c959f024c073472122d858333a0f24866c fix/mobile-compare-polish
a00cbe916182d1e04661c750a1ce2e7ad27a68ae fix/mobile-sort-select-overflow
2fd997bfe640c419a6713e85df008463de7f56e8 fix/modal-keyboard-viewport
3e7705756776a0c44d27966dfc972023f1f38b69 fix/ofsted-link-text
ce422e64363e2b03c186ef316832b6de15502d67 fix/promote-status-token
4522cbf64560db1e1cd519e119aa466b42cd16a1 fix/proposed-to-close-copy
15da060e4af37fbae919e0edf25f108266de2585 fix/rankings-admissions-accuracy
6c872ce726f210433354ca38dc5314bc6a467534 fix/rankings-year-validation
1c1df7796194af3d47f8e5ac0a0fbe6f323700e7 fix/report-card-chip-alignment
b2dc4d0779ced02709429afaa5985acf4abda794 fix/results-map-ios-fullscreen
95a5783da1fc994df76cb97238b55596dce4cd8f fix/runtime-api-proxy
8a9ba30cc24e29653a72c03c0e817684b7db7c07 fix/sats-per-level-national
536832a524fc4f9ed858f047e00b9a4d9d429c46 fix/school-detail-nan-500
fef83b3bf244a9bf3cb4afa75dbc433d8725014f fix/schoolbar-sticky-offset
b0c5b6bb57c879477da12c37b15cff950c29ebd3 fix/secondary-anchors-button-affordance
261403bcd2b01aa4f26ee212e26302fc0f769bf9 fix/special-note-full-width
ea5249a2ea6faf6bfa5a1522644387ba59770ec8 fix/standardise-distance-units
e4565e9f158721d4df6b918f2b065de82851f8d9 fix/trends-chart-height
3aad5101a842539105022f9059e85c224fc5973a fix/welsh-establishment-leak
315f1feede70bdf3101d2fdd305b3d06e037fdac perf/batch-supplementary
d2dc78aeb599e16b7ef5019be2b360b08df463bc perf/compare-loading
e098ad4bd1130152705788d4773837b7d1e7112e perf/server-client-split
## Deleted: superseded by PR #110 (content preserved on feat/seo-crawl-hygiene-main)
786ec80dd4de4a3cb674a89e35b3b5e461639289 feat/seo-crawl-hygiene
a5ac0bcd1b37bc10dcbce88f8601d01bf7b3eaaf feat/england-only-corpus
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -19,7 +19,7 @@
- Screenshot `j1-postcode-results-mobile.png` confirms cards carry only the LA name ("Solihull"), no "X miles away".
**Desktop (1440×900) — Attempt A repeat + keyboard + zoom.**
- Desktop hero additionally shows a trust badge ("● UPDATED WITH 2026/2027 ADMISSIONS RESULTS") and a value-prop subheading ("24,000+ primary and secondary schools with Key Stage 2 SATs, GCSE results, Ofsted grades, progress scores and admissions data — side by side, in one place"). **Both are absent on the mobile hero** (mobile jumps H1 → search box).
- Desktop hero additionally shows a trust badge ("● UPDATED WITH 2026/2027 ADMISSIONS RESULTS") and a value-prop subheading ("27,000+ primary and secondary schools with Key Stage 2 SATs, GCSE results, Ofsted grades, progress scores and admissions data — side by side, in one place"). **Both are absent on the mobile hero** (mobile jumps H1 → search box).
- Name search identical, correct, instant (`j1-search-results-desktop.png`).
- Keyboard: tab order is logical — Skip link → logo → Search/Compare/Rankings/Admissions nav → search input → Search button → Schools near me → content. Skip link, logo, nav links and Search button all get a clear **2px solid orange (#E07256) focus outline**. The search input uses an orange border + very faint ring (`box-shadow rgba(224,114,86,0.12) 0 0 0 3px`, `outline:none`) — visible but weaker than the other controls. Search → results → school link is fully keyboard-operable (standard links/buttons).
- 200% reflow (720×450): **no horizontal scroll** (`scrollWidth == clientWidth == 720`); deadline cards reflow from 1×4 to 2×2, no overlap or clipping. Pass.
@@ -57,7 +57,7 @@
- Severity guess: P2
- **F6. Mobile hero omits the value proposition shown on desktop**
- Evidence: Desktop hero has the "UPDATED WITH 2026/2027…" badge + subheading "24,000+ primary and secondary schools with KS2 SATs, GCSE results, Ofsted grades… side by side, in one place" (`j1-home-desktop-fold.png`). The mobile hero (`j1-home-mobile-fold.png`) drops both — below the poetic-but-vague H1 there is only a search box.
- Evidence: Desktop hero has the "UPDATED WITH 2026/2027…" badge + subheading "27,000+ primary and secondary schools with KS2 SATs, GCSE results, Ofsted grades… side by side, in one place" (`j1-home-desktop-fold.png`). The mobile hero (`j1-home-mobile-fold.png`) drops both — below the poetic-but-vague H1 there is only a search box.
- Criterion violated: Nielsen #1 (Visibility of system status) / recognition-over-recall; mobile content-parity best practice (the primary 63%-of-entries viewport should not lose the core "what is this and why trust it" copy).
- Argument: A first-time parent landing on mobile sees "Every school in England, compared." + a bare box, with no statement of coverage, data sources, or freshness. Weak/absent value proposition above the fold is a classic driver of immediate exits — directly relevant to the 46% home-exit rate.
- Severity guess: P2 (candidate P1 given mobile is the primary, highest-traffic viewport)
@@ -0,0 +1,309 @@
# SEO Programme — Design
Date: 2026-08-20
Status: awaiting review
## Problem
schoolcompare ranks second for "school compare" — an exact match for the
brand and the domain. It ranks poorly for "compare schools", "school
comparison" and "schools near me". The first is a naming artefact and
transfers to nothing; the rest are the queries that actually carry parent
demand.
The cause is structural, not editorial. The site publishes five route
families:
/ /compare /rankings /admissions /school/[slug]
Location intent has no landing page at all. Every competitor outranking us
on those queries wins with programmatic location pages:
| Competitor | URL pattern |
|---------------|--------------------------------------------|
| School Guide | `/best-schools-in/manchester` |
| Locrating | `/the-best-primary-schools-in-Manchester_…`|
| FindMySchool | `/best-primary-schools/manchester` |
| Snobe | `/best-primary-schools/manchester` |
| School Atlas | `/guides/best-primary-schools-manchester` |
"Schools near me" is a local-intent query. Google resolves it against the
user's coordinates and serves pages that are *about a place*. A national
homepage cannot win it. No title or description change fixes this; only
pages Google can localise will.
## Baseline (measured 2026-08-20, production API)
| Measure | Value |
|--------------------------------------------|---------|
| Unique schools | 27,229 |
| URLs in sitemap.xml | 27,232 |
| Schools with 2024/25 performance data | 21,266 |
| Schools with **no** current performance data| ~5,963 (22%) |
| Welsh establishments (all metrics null) | 1,569 |
| Overseas / offshore establishments | 467 |
| Static URLs in sitemap | 3 |
| Routes setting a canonical | 1 of 5 |
Three findings from that table drive the plan.
**We submit ~6,000 thin pages to Google.** `build_sitemap()`
(`backend/app.py:73`) enumerates every URN regardless of whether the school
has any data. Welsh establishments return `school_type: "Welsh
establishment"` with every performance metric, Ofsted grade and phase field
null. Overseas and offshore establishments ("BFPO Overseas Establishments",
"Gibraltar Overseas Establishments", "Jersey Offshore Establishments") are
in the local-authority list too. At 22% of the submitted corpus this is a
site-wide quality signal problem and a crawl-budget waste, not a rounding
error.
W1 item 4 removes 2,036 of those — every non-England establishment — taking
the corpus to 25,193. The 3,927 that remain are English schools with no
current data: mostly newly opened, special, nursery or alternative provision.
Those are a template problem, not a corpus problem, and item 5 handles them
separately.
**The homepage is its own competitor.** `app/page.tsx` accepts eleven search
params (`search`, `local_authority`, `school_type`, `phase`, `page`,
`postcode`, `radius`, `sort`, `gender`, `admissions_policy`,
`has_sixth_form`) and sets no canonical. Every filter combination is a
crawlable near-duplicate of the single page we are asking to rank for
"compare schools".
**School pages are near-orphans.** Reachable from the sitemap and from site
search, but almost nothing links to them contextually, so they accrue no
internal authority.
Also noted: the sitemap emits invented `priority` values and no `lastmod`.
Google ignores `priority` and `changefreq` entirely; `lastmod` is the field
it does read, and we omit it.
## Keyword clusters
Ranked by judgement of UK parent search behaviour and by the competitive
SERP evidence above. Google Search Console is connected, so cluster
priorities are to be re-derived from measured impressions before build
starts (see Workstream 0).
**C1 — Head "compare" terms.** compare schools · school comparison · school
comparison tool · compare school performance · compare primary schools ·
compare secondary schools · compare two schools
**C2 — League tables and rankings.** primary school league tables ·
secondary school league tables · school league tables 2026 · SATs results by
school · GCSE results by school · KS2 league tables · Progress 8 rankings ·
best primary schools in [town] · top 10 primary schools in [LA]
**C3 — Local / near me.** schools near me · primary schools near me ·
secondary schools near me · best schools near me · good schools near me ·
schools in [town] · primary schools in [LA] · schools near [postcode] ·
[postcode] school catchment
**C4 — Individual school long tail.** [school] ofsted · [school] SATs
results · [school] catchment area · [school] reviews · [school] URN
**C5 — Admissions.** primary school admissions 2027 · national offer day
2027 · school application deadline · school admissions appeal ·
oversubscription criteria · distance criteria school admissions · didn't get
first choice school · school admissions [LA]
**C6 — Metric explainers.** what is a good SATs score · what is Progress 8 ·
what is Attainment 8 · expected standard KS2 meaning · scaled score
explained · Ofsted grades explained · Ofsted report cards · pupil premium
explained
**C7 — Head to head.** [school A] vs [school B] · academy vs community
school · grammar school vs comprehensive · faith school vs community school
## Workstreams
### W0 — Measure before touching anything
Export a Google Search Console baseline: impressions, clicks, average
position and CTR by query and by page, for the trailing 16 months. Bucket
queries into C1–C7. This sets the counterfactual — without it, no later
claim about lift is defensible, because school-search traffic is strongly
seasonal (results day in December, offer day in March/April).
Re-rank C1–C7 against measured impressions and adjust the sequence below if
the data disagrees with the judgement calls.
### W1 — Crawl hygiene and index sanity
Cheap, and it unblocks everything after it. Adding 5,000 pages on top of a
corpus that is 22% thin would compound the existing problem.
1. Canonical on every route. `/`, `/rankings`, `/compare` and `/admissions`
currently set none.
2. The homepage canonicalises to `/` regardless of search params.
3. `/compare?urns=…` gets `noindex, follow` — it is an unbounded parameter
space with no standalone value.
4. **England only — DONE.** Wales, the Crown Dependencies, Gibraltar and the
service/overseas schools are removed from the corpus at the mart boundary,
not hidden at the view layer. `dim_school` and `dim_location` both exclude
`TypeOfEstablishment` in {25, 26, 30, 37} — Offshore schools, Service
children's education, Welsh establishment, British schools overseas —
listed once as `vars.non_england_school_type_codes` in `dbt_project.yml`.
That removes 2,036 establishments and 29 local authorities, and because
`build_sitemap()` reads the same marts, it drops them from the sitemap in
the same stroke. `assert_england_only_schools` fails the pipeline if a GIAS
refresh reintroduces them or if the two models drift apart.
5. Prune the remaining thin pages: exclude any school with no performance data
**and** no Ofsted record. Distinct from item 4 — these are English schools
with nothing yet to show, so the fix may be a better template rather than
removal.
6. Rebuild the sitemap as a sitemap **index**: one child per page family,
real `lastmod` from the data-load timestamp, `priority` and `changefreq`
dropped.
### W2 — The location layer
The dominant lever. `dim_location` already carries `town`, `county`,
`local_authority_name`, `parliamentary_constituency`, `latitude`,
`longitude` and `postcode`, so no new ingestion is required.
Routes:
/schools/[la] e.g. /schools/manchester
/schools/[la]/primary
/schools/[la]/secondary
/best-primary-schools/[town]
/best-secondary-schools/[town]
/schools/near/[outcode] e.g. /schools/near/m20
/schools/near-me geolocating hub
**Thin-page threshold: generate a town or outcode page only where at least
five schools have current performance data.** Below that, 301 to the parent
LA page. This is the single most important constraint in the workstream —
it is what separates a location layer from index bloat.
Each page must earn its place with content a parent would actually use, not
a template shell:
- H1 matching the query intent ("Best primary schools in Manchester")
- Counts framed usefully: "137 primary schools, 9 rated Outstanding"
- A ranked table of the top 20 on the headline metric
- Local average against the England average
- Ofsted grade distribution
- Map
- Links to neighbouring towns and to the parent LA
- An FAQ block (feeds `FAQPage` in W4)
- Links to every school page in scope — this is what de-orphans W1's corpus
Sizing estimate: ~150 usable LAs × 3 ≈ 450; towns clearing the threshold
≈ 1,200 × 2 ≈ 2,400; outcodes ≈ 2,300. Roughly **5,000 new pages**,
comfortably inside a sitemap index and well under the per-file 50,000 limit.
### W3 — Make rankings indexable
`/rankings` is driven entirely by query params, so Google indexes
approximately one page where there should be hundreds.
/rankings/[phase]/[metric]
/rankings/[phase]/[metric]/[la]
The interactive filter UI stays; its state moves into real paths. Param
forms canonicalise to the clean path. This is the direct play for C2.
### W4 — Structured data and internal linking
- Replace the bare `EducationalOrganization` on school pages with `School`,
and populate it properly.
- `BreadcrumbList` site-wide.
- `ItemList` on every rankings and location page.
- `FAQPage` on admissions and on location pages.
- New school-page modules: "Other schools in [town]", "Nearby schools",
"Compare with similar schools". Each links out to W2 and W3 pages, which
is what circulates authority instead of stranding it.
Explicitly **not** doing `Dataset` or `AggregateRating` — no review corpus
exists, and fabricating one would be both useless and a policy violation.
### W5 — Admissions expansion
One static page currently carries an entire cluster.
/admissions/[la] per-authority deadlines and offer day
/admissions/appeals
/admissions/national-offer-day
`school admissions [LA]` is high-intent and highly seasonal; per-authority
pages are the natural unit.
### W6 — Explainer content
/guides/progress-8
/guides/attainment-8
/guides/sats-scaled-scores
/guides/ofsted-grades
/guides/expected-standard
Each links into the corresponding W3 rankings page. Cheap to build, and it
is what gives the metric vocabulary enough topical weight to support C1–C4.
### W7 — Head-to-head pages
/compare/[school-a]-vs-[school-b]
C7 is uncontested and native to the product. It is also the easiest way to
destroy everything W1 fixes: 27,229 schools generate 370 million pairs.
**Curated pairs only** — same town, both with current data, both with real
search demand — capped in the low thousands. Gated behind evidence that W2
is indexing cleanly.
### W8 — Metadata rewrite for C1
Current homepage title is `schoolcompare | Compare every school in England`,
which spends the most valuable position on the brand. Rewrite the homepage,
rankings and compare titles and descriptions around C1 phrasing. Small
change, and the cheapest item in the programme.
## Sequencing
W0 → W1 → W2 → W3 → W4 → W5 → W6 → (W7 if W2 indexes cleanly)
W1 before W2 is not negotiable: adding pages to a corpus that is 22% thin
compounds the problem rather than diluting it.
This spec is a programme, not a single implementation plan. Each workstream
gets its own plan and its own PR; W2 will likely need several. Only W0 and W1
are ready to plan against today — the rest should be re-read after W0's
Search Console baseline lands, because that data may reorder them.
## Testing
Per CLAUDE.md, user-facing behaviour changes extend the `e2e/` journeys in
the same PR. Each workstream adds:
- W1: canonical present and correct on every route; `/compare?urns=` carries
`noindex`; sitemap excludes a known dataless URN.
- W2: a known LA, town and outcode page renders with the expected school
count; a below-threshold town redirects to its LA.
- W3: a clean rankings path renders; the param form canonicalises to it.
- W4: JSON-LD parses and validates against the declared types.
## Risks
**Index bloat.** The failure mode of every programmatic SEO programme. The
five-school threshold, the W1 prune and the W7 gate are the three controls.
**Helpful-content exposure.** Google's stance on templated location pages
has hardened. The mitigation is that each page carries genuinely local
computed data — real counts, real distributions, real local-vs-national
comparison — rather than a name substituted into boilerplate.
**Build cost.** School pages already use ISR with a 7-day revalidate and
`PRERENDER_SCHOOLS` gating full prerender. 5,000 more routes need the same
treatment; full static generation of 32,000 pages is likely impractical in
CI.
**Seasonality.** Results day and offer day dominate the traffic curve.
Judging the programme on a mid-summer window would misread it in either
direction. W0's baseline must be year-on-year, not month-on-month.
## Open questions
1. Catchment areas are Locrating's moat and a strong C3 driver
(`[postcode] school catchment`). `fact_admissions` carries admission
distances. Is deriving approximate catchment a later workstream, or out
of scope?
@@ -0,0 +1,278 @@
# W2: The Location Layer — Design
Date: 2026-08-21
Status: awaiting review
Supersedes: workstream W2 in `2026-08-20-seo-programme-design.md`
Scope note: this covers four page families in one spec. Splitting them — towns
and authorities first, outcodes and localities after — was proposed and
declined in favour of building the layer in one pass. The decomposition
argument was that the curated locality seed needs human review and would hold
up 783 pages of measured demand behind it; that risk is accepted here, and the
implementation plan should sequence the seed early enough that review time
does not become the critical path.
## Problem
Location intent is the largest unserved demand the site has. In the 16-month
Search Console baseline it draws **874 impressions, one click, average
position 49.5**. The site does not compete.
Unlike named-school queries — which the same baseline showed to be
navigational and unwinnable, since a parent typing "audley junior school"
wants that school's own website — location queries have no incumbent owner.
Nobody owns "primary schools in Brentwood" the way a school owns its name.
The cause is structural: the site has no page about a place. Every competitor
ranking above it does.
## What the demand actually looks like
Every location query in the baseline is **town or district level**. Not one is
an administrative area:
| Query | Impressions | Position |
|-------|-------------|----------|
| colleges in solihull | 112 | 51.2 |
| schools in ramsey | 64 | 42.5 |
| schools in crosby | 57 | 47.7 |
| primary schools in beccles | 44 | 40.9 |
| private schools in battersea | 41 | 71.9 |
| secondary schools in brentwood | 37 | 56.1 |
| secondary schools in canary wharf | 30 | 35.9 |
Three patterns follow directly, and they drive the whole design.
**Towns, not authorities.** The superseded W2 put `/schools/[la]` first and
towns second. The data inverts that. Brentwood appears four times in different
phrasings; Beccles twice. Both are towns, not authorities.
**Phase is part of the query**, not a filter applied afterwards: "primary
schools in beccles", "secondary schools in brentwood", "colleges in solihull".
**London is searched by district** — Battersea, Canary Wharf — and the GIAS
`town` field cannot serve it at all.
## Measured sizing
Counted against the live corpus of 25,185 schools, not estimated.
| Family | Viable (≥5 schools) | Below threshold |
|--------|--------------------|-----------------|
| Towns | **783** | 907 → redirect to authority |
| Outcodes | **1,760** | 305 |
| Local authorities | 154 | — |
| London localities | ~100–150 (curated) | — |
With phase variants — 783 town pages plus roughly 700 primary and 250
secondary variants, 154 authorities across three variants, 1,760 outcodes and
the curated localities — the total lands near **4,000 pages**. Phase variants
need their own threshold: there are 17,426 primaries but only 4,456 secondaries nationally,
so most towns will support a primary page and not a secondary one.
## Two design problems this spec exists to solve
### 1. Town and authority names collide, and neither contains the other
67 viable towns share a name with a local authority. The obvious fix — let the
authority absorb the town, since it sounds like a superset — **does not work**:
| Place | Schools in the town | Schools in the authority |
|-------|--------------------|-----------------------|
| Bedford | 104 | 86 |
| Birmingham | 520 | 518 |
| Derby | 157 | 119 |
| Doncaster | 152 | 145 |
The authority is the larger set in only 43 of the 67. Postal towns cross
authority boundaries, so these are overlapping sets that happen to share a
name. Publishing both into one namespace produces near-duplicate pages, which
is the specific failure that sinks programmatic SEO.
**Resolution: two namespaces.**
```
/schools/[place] towns and London localities
/schools/[place]/primary
/schools/[place]/secondary
/schools/authority/[la] local authorities
/schools/authority/[la]/primary
/schools/authority/[la]/secondary
/schools/near/[outcode]
```
Outcodes carry no phase variants: nobody searches "primary schools in SW11",
so the variants would be pages without demand.
Every collision disappears by construction. `/schools/[place]` keeps the clean
URL for the pattern that carries the demand; authorities get a namespace whose
purpose is genuinely different — admissions are authority-run, and the
authority page is the one that can speak to catchment policy and LA averages.
A place page and an authority page of the same name must each say plainly
which set of schools they cover, or they read as duplicates to a reader even
when they differ in fact.
### 2. London has no locality field
`town` collapses **1,819 London schools into the single value "London"**. A
page listing all of them is useless, and borough pages do not help because
people search "Battersea", not "Wandsworth".
No single field solves it:
| Search term | `parliamentary_constituency` | postcodes.io `admin_ward` |
|-------------|------------------------------|---------------------------|
| Battersea | **Battersea** ✓ | Northcote / Wandsworth Town ✗ |
| Canary Wharf | Poplar and Limehouse ✗ | **Canary Wharf** ✓ |
| Vauxhall | Vauxhall and Camberwell Green ✗ | **Vauxhall** ✓ |
And neither covers Clapham, Shoreditch or Peckham, which are postal and
colloquial rather than administrative.
**Resolution: a curated seed mapping locality to outcodes.**
```
pipeline/transform/seeds/locality_outcodes.csv
locality_slug,locality_name,outcodes,region
battersea,Battersea,"SW11|SW8",London
canary-wharf,Canary Wharf,"E14",London
clapham,Clapham,"SW4|SW9",London
```
This needs **no new ingestion** — the corpus already has postcodes. It puts
the fuzzy, contested part of the problem in a reviewable file rather than in
derived logic, which suits it: locality boundaries are a judgement, not a
fact. The repo already uses dbt seeds for curated reference data
(`la_code_names.csv`, `gias_code_names.csv`), so this follows an established
pattern.
The seed generalises past London. Any colloquial place — Jesmond, Chorlton,
Clifton — can be defined by its outcodes without a schema change.
**Constraint:** a locality slug may not collide with a viable town slug. The
place registry enforces this and fails the build rather than silently
shadowing a town.
## Architecture
### The place registry
One module owns the question "what places do we publish, and what is in each".
Everything else reads from it: the pages, the sitemap, the internal links.
```
backend/places.py
Place = { kind: "town"|"locality"|"authority"|"outcode",
slug, name, urn_list, parent_authority | None }
build_place_registry(df) -> dict[str, Place]
place_schools(slug, phase=None) -> list[School]
```
Built once at startup from the same DataFrame the sitemap uses, and rebuilt by
the existing `/api/admin/regenerate-sitemap` path after a pipeline run.
Registry construction is where the threshold, the collision rules and the
seed's uniqueness constraint are enforced — in one place, testable without a
browser or a database.
### API
```
GET /api/places the registry: slug, kind, name, count
GET /api/places/{slug}?phase= aggregate + ranked schools for one place
```
`/api/places` is what the sitemap and the internal-link modules enumerate.
### Routes
Next App Router, ISR with the same 7-day revalidate the school pages use.
`generateStaticParams` gated behind an env flag, matching
`PRERENDER_SCHOOLS`, because 3,900 more routes cannot be statically built in
CI on every deploy.
## What each page must contain
A place page that is a name substituted into a template is the thing Google's
helpful-content stance exists to demote. Each page carries computed local
facts that exist nowhere else on the site:
- **H1** matching the query: "Primary schools in Brentwood"
- **Counts framed usefully**: "29 schools, 4 rated Outstanding"
- **A ranked table** of the top 20 on the phase's headline metric —
`rwm_expected_pct` for primary, `attainment_8_score` for secondary, and for
an unphased place page the metric matching whichever phase it holds more of
- **The local average against the England average** — the one number a parent
cannot get from a list
- **Ofsted grade distribution** for the place
- **A map**
- **Links to neighbouring places** and to the parent authority
- **An FAQ block**, feeding `FAQPage` structured data
- **A link to every school page in scope** — this is what finally de-orphans
the 23,000 school pages the original spec identified as near-orphans
## Thin-page controls
Three, and they are the difference between a location layer and index bloat:
1. **Five schools with current data minimum.** Below it, 301 to the parent
authority. This drops 907 towns and 305 outcodes.
2. **Per-phase thresholds.** A town with 30 primaries and 2 secondaries
publishes a primary page and no secondary page.
3. **No page without a local average.** If a place has too few schools with
results to compute one, it has nothing to say that a list does not, and it
falls back to the authority.
## Sitemap
Two new children in the existing index: `/sitemaps/places-{n}.xml` and
`/sitemaps/outcodes-{n}.xml`. Per-family children are why the index was built
in W1 — Search Console reports coverage per submitted sitemap, so indexation
of the location layer is measurable separately from the school pages.
## Testing
Per `CLAUDE.md`, user-facing behaviour extends `e2e/tests/journeys.spec.ts` in
the same PR.
**Unit (registry, no DB):** threshold enforcement; a sub-threshold town
resolves to its authority; a locality slug colliding with a town fails the
build; Bedford's town and authority pages hold different URN sets; per-phase
thresholds.
**Backend:** `/api/places` shape; `/api/places/{slug}` aggregate correctness
against a fixture; unknown slug 404s.
**e2e:** a known town, authority, locality and outcode page each render with
the expected count; a below-threshold town 301s; every place page declares a
canonical and appears in the sitemap; `/schools/bedford` and
`/schools/authority/bedford` both resolve and state which set they cover.
## Risks
**Index bloat** is the failure mode of every programmatic SEO programme. The
three controls above are the answer, and the per-family sitemap is how we find
out early if they were not enough.
**Helpful-content exposure.** Templated location pages are exactly what
Google's stance targets. The mitigation is that every page carries real
computed local data — counts, distributions, local-versus-national comparison
— rather than a name dropped into boilerplate. If indexation of the places
sitemap stalls below roughly half, that is the signal to stop and rethink
rather than to add more pages.
**Build cost.** ~4,000 additional ISR routes on top of 23,000 school pages.
The env-flag gate on `generateStaticParams` keeps CI viable.
**Curation drift.** The locality seed is hand-maintained and will go stale as
places change. It is small and reviewable, and a dbt test asserts every seed
outcode matches at least one school so a typo fails the pipeline rather than
publishing an empty page.
## Out of scope
Catchment-area estimation. It is a strong driver for this cluster and
`fact_admissions` carries the distances, but it is a modelling problem with
real accuracy risk and deserves its own design.
@@ -0,0 +1,294 @@
# Feature Flags — Design
**Date:** 2026-08-23
**Status:** approved for planning
**First consumer:** the last-distance-offered feature (`admission_distance`)
## Goal
Let work merge to `main` and deploy to production without becoming visible,
so that releasing a feature stops being the same event as deploying it.
The site has no way to do this today. A feature is either on `main` and live,
or it is on a branch. That forces long-lived branches for anything not ready,
and it makes every promotion to production an all-or-nothing decision about
everything queued behind it.
This is a **ship-dark** capability, not a kill switch. Flags are expected to
flip on the order of once a month, by a person, deliberately. Nothing here is
designed for flipping something off in seconds under pressure, and nothing
here does percentage rollouts, user targeting or A/B tests — the site has no
user identity to target.
## Decision: Unleash
Flag state is held in a self-hosted [Unleash](https://www.getunleash.io/)
instance (Apache-2.0), not in the repository.
A lighter option was considered and rejected by the project owner: a typed
registry in each runtime with environment-variable overrides set in the
Portainer stack files, which would have needed no new container and kept flag
state in git. The argument for Unleash is that it provides a UI and an audit
log without a deploy, and that flags are expected to become an ongoing
operational tool rather than an occasional one.
Two consequences follow from choosing a service, and this design exists mostly
to handle them:
1. **Flag state lives outside the repository.** `main` is no longer the whole
truth about what is switched on. The registry in §2 exists to bound that.
2. **A flag can change without a deploy**, so nothing else clears the caches
that a deploy would have cleared. §4 establishes how long a flip takes to
become visible, and why that is short enough to need no extra mechanism.
Also considered: Flagsmith (heavier — Django, Postgres and Redis), GrowthBook
(requires MongoDB), and Flipt v2 (the closest conceptual fit, git-native, but
now under the Fair Core Licence — source-available, not OSI open source).
## 1. Topology
A third Portainer stack, `docker-compose.portainer.unleash.yml`, holding
`unleashorg/unleash-server` and its own PostgreSQL 16. It is on the macvlan so
both application stacks can reach it, and it belongs to neither of them — a
staging redeploy must not be able to disturb production's flag state, and vice
versa.
One instance serves both environments. Open-source Unleash ships with
`development` and `production` environments and environment-scoped client
tokens, so the same flag holds independent state in each: staging's FastAPI
carries a `development` token, production's carries a `production` one.
That property is what makes ship-dark testable. A feature can be **on in
staging and off in production** for as long as it takes, which means the `e2e/`
journeys exercise it against staging while production stays unchanged.
## 2. The registry
Unleash supplies flag *state* and the toggle UI. It does not supply the list of
flags. `backend/flags.py` declares every flag the code knows about:
```python
@dataclass(frozen=True)
class Flag:
name: str # identical in the registry, in Unleash, and in JSON
description: str # one line: what turning this on reveals
added: date # for the staleness test in §8
```
**Every flag defaults to `False`.** There is no per-flag default field, because
a flag that defaults on is not a ship-dark flag — it is a kill switch, and this
design does not offer one. A single unconditional default also means the
fallback path has no branching to get wrong.
Three reasons the registry is not optional:
- The Unleash SDK evaluates an unknown flag to `False`. Without a registry that
is an *undeclared* false — indistinguishable from a typo in a flag name.
- `/api/flags` needs a key set to return when Unleash is unreachable. It cannot
enumerate flags it has never heard of.
- A flag present in the Unleash UI but absent from the registry is orphaned,
and should be visibly so rather than quietly authoritative.
**Naming.** One string, used unchanged as the registry key, the Unleash flag
name, and the JSON key in `/api/flags`. It is snake_case, matching the API's
existing convention (`admission_distance`, `rwm_expected_pct`) and the mirrored
types in `nextjs-app/lib/types.ts`. No case transformation anywhere, so there
is no mapping layer to get wrong.
## 3. Read paths
### Backend
`backend/flags.py` wraps `UnleashClient` behind `is_enabled(name: str) -> bool`.
Fail-closed is the default rather than something added: the Python SDK
evaluates every flag to `False` until it has synchronised with the server. An
unfinished feature therefore stays hidden when Unleash is unreachable, which is
the correct direction for ship-dark.
The SDK's fcache directory is mounted on a named volume so a container restart
during an Unleash outage keeps last-known state rather than reverting a
released feature to dark. The registry default remains `False`, so the worst
case is a feature disappearing, never one appearing.
### Frontend
`nextjs-app/lib/flags.ts` exposes `getFlags(): Promise<Flags>`, a single
server-side fetch of `/api/flags` returning a typed record. Server components
only — no flag value reaches the browser bundle, and `package.json` gains no
Unleash dependency. The Unleash client library stays entirely inside the
service that already owns every other piece of data the frontend renders.
The cost, named plainly: a purely front-end flag must still be declared in a
Python file. It is a flat data edit rather than programming, and the return is
one list, so nobody has to ask which service knows about a given flag.
### `/api/flags` must not be publicly reachable
`nextjs-app/app/api/[...path]/route.ts` proxies **everything** under `/api/` to
FastAPI. Left alone, `https://www.schoolcompare.co.uk/api/flags` would return
`{"admission_distance": false, ...}` — publishing the name and state of every
unreleased feature, which defeats the purpose of shipping dark.
The proxy therefore gains a denylist, and `flags` is on it: a request for a
denied path returns 404 rather than being forwarded. Next's own `getFlags()` is
unaffected because it calls `FASTAPI_URL` directly across the Docker network
and never transits the public proxy.
This is a general hole rather than a flags-specific one — the proxy will
forward any future internal endpoint too — so the denylist is written as a
named constant with a comment saying what belongs on it.
## 4. Propagation
**Time-based revalidation is sufficient. There is no webhook.**
An earlier draft of this section specified two Unleash webhooks and a
`revalidateTag('flags')` purge, on the premise that pages cache for seven days.
That premise was wrong, and checking it removed the most complex part of the
design.
Next uses the **lowest** `revalidate` among a route's fetches to set the
revalidation frequency of the whole route — the segment-level
`export const revalidate` does not override a lower value inside it. Measured
against this codebase:
| Page family | Segment | Lowest fetch | Effective |
|---|---|---|---|
| `/school/[slug]` | 604800 | `fetchSchoolDetails` at 300 | **5 minutes** |
| `/schools/*` | 604800 | `fetchNationalAverages` at 3600 | **1 hour** |
The Unleash SDK polls every 15 seconds, so a flip reaches school pages within
about five minutes and place pages within the hour, unaided. Flags flip
monthly, by hand, deliberately. That is fast enough.
What this removes: two webhook integrations, a `/api/revalidate-flags` route, a
shared-secret-in-a-query-string scheme, an idempotency requirement against
duplicate and out-of-order delivery, and a rule that every fetch in
`nextjs-app/lib/` carry a cache tag. None of it has to be built, maintained, or
kept correct as new fetches are added.
**If instant flips are ever wanted**, the webhook is the way to add them, and it
is purely additive — nothing in this design has to change first.
### Two constraints this leaves behind
**Never flag content on a `force-static` page.** `app/admissions/page.tsx`
declares `export const dynamic = 'force-static'`, so it is baked at build time
and never revalidates. A flag gating anything on such a page would not take
effect until the next deploy, silently. If a flag ever needs to reach one, that
page must first move to ISR.
**A route-family flag still needs the sitemap rebuilt.** The sitemap is held in
memory and rebuilt only at startup or via `POST /api/admin/regenerate-sitemap`.
No flag in scope touches the sitemap (§6), so this is deferred with the route
case rather than solved now — but a route flag must not ship without it, or the
sitemap will advertise URLs that `notFound()`.
## 5. What "off" means, per surface
| Surface | Off |
|---|---|
| Route | `notFound()`, **and** absent from the sitemap, **and** absent from nav |
| UI element | Not rendered; surrounding page byte-identical to today |
| API field | Key **absent**, not `null` |
| API endpoint | 404, not 403 |
The three parts of the route rule move together or not at all. Submitting URLs
to Google that return 404 is the bug fixed in PR #124, and a flag is a new way
to reintroduce it.
An API field is withheld **at the source**, never rendered-but-hidden. The
precedent is already set in this codebase by commit `c9a1892`: `/api/schools/`
is public and unauthenticated, so leaving a withheld field in the payload hands
the record to anyone who opens the network tab.
## 6. First consumer: `admission_distance`
The last-distance-offered feature is merged to `main` and live on staging.
Production has never received it: `/api/schools/100010` on production carries
no `admission_distance` key, and no Distance section renders.
It needs **exactly one gate** — `backend/app.py:809`, where the field is
attached to the school payload:
```python
"admission_distance": (
supplementary.get("admission_distance")
if flags.is_enabled("admission_distance") else None
),
```
The frontend follows with no change. `DistanceSection` already returns `null`
when `admission_distance?.distance_m == null`, and `PrimarySchoolSections`
already conditions the admissions block on `(admissions || admissionDistance)`.
The off-state is the commonest state on the site — only 57 local authorities
publish cut-off distances at all — so it is well covered by construction.
The flag does not touch the sitemap: school pages exist either way.
Intended lifecycle: default off, so production receives the code dark on the
next promotion; on in the `development` environment so staging keeps testing
it; flipped on in `production` when the owner chooses.
**This flag exercises two of the three surfaces** in §5 — API field and UI
element. No route case ships with it. The route rule is specified but unproven
until a route-shaped flag exists, and should be treated as such.
## 7. Testing
**Backend unit.** The registry is well-formed; an unknown flag evaluates
`False`; `/api/flags` returns every declared flag with its default when the
SDK is unreachable; `admission_distance` is absent from the school payload when
the flag is off and present when on.
**Frontend unit.** `getFlags()` returns declared defaults when `/api/flags`
fails, rather than throwing and taking the page with it.
**E2E.** Journeys read `/api/flags` and gate flag-dependent assertions on it,
matching the `test.skip` shape the suite already uses.
One trap to avoid, worth stating because the existing distance journeys walk
straight into it: they already skip when no school has a published figure, so
with the flag off they would skip silently and the suite would go green. The
gate must be explicit — **if `/api/flags` reports `admission_distance` on, then
a school with a cut-off must be found**, converting a silent skip into a real
assertion.
## 8. Lifecycle
A flag is temporary scaffolding, and the failure mode of every flag system is
accumulation.
The registry records the date each flag was added, and a backend test fails any
flag older than **90 days**. Removing a flag means deleting the registry entry,
the branches that read it, and the flag in the Unleash UI.
Unleash SDK usage metrics stay enabled, so the UI shows which flags are still
being evaluated — the evidence needed to retire one safely.
## 9. Risks
**Production gains a homelab dependency.** If Unleash is unreachable when a
production container cold-starts with an empty cache, every flag evaluates
`False` and any feature currently switched on disappears. The fcache volume
covers restarts; the 90-day lifecycle rule bounds how long any feature is
exposed to this. It is a real regression risk and the reason flags must be
retired rather than left on indefinitely.
**Flag state is not in git.** `main` no longer tells you what production is
showing. The registry lists what *can* be flagged; only the Unleash UI says
what *is*. This is inherent to the choice of a service.
**A large promotion backlog exists.** Production is running the
pre-SEO-programme build — no place pages, and a sitemap still declaring the
apex host. The first promotion after this work ships that entire backlog. The
flag isolates the distance feature from it and nothing else.
## Out of scope
- Percentage rollouts, user targeting, A/B testing, and Unleash strategies
beyond simple on/off. Flags are booleans.
- Pipeline and dbt flags. Airflow and dbt are not flag consumers.
- Client-side flag evaluation. Flags are server-side only.
- Automatic flag removal. The staleness test reports; a person deletes.
@@ -0,0 +1,294 @@
# School Autosuggest — Design
**Date:** 2026-08-26
**Status:** approved for planning
**Depends on:** the feature-flag layer (PR #125, merged)
## Goal
Suggest schools by name as someone types in the site's main search box, so a
parent who knows the school they want reaches it in one step instead of
searching, scanning a result list, and clicking.
Scope is **schools only**. Places and postcodes were considered and excluded —
see *Out of scope*.
## The finding that shapes everything
The site's rate limiter does not do what it looks like it does.
`limiter = Limiter(key_func=get_remote_address)` with `60/minute` reads
`request.client.host`. In staging and production the backend has no published
ports and sits on the internal `backend` network, so its only caller is the
Next proxy — and `request.client.host` is therefore **the Next container**, for
every browser user on the site.
Measured against staging: 70 concurrent requests to `/api/schools` returned
**60 × 200 and 10 × 429**. One machine consumed the whole site's budget for
that minute.
Autosuggest is the worst possible feature to build on that. One person typing
"st marys primary" produces six to eight debounced requests; **eight concurrent
searchers would 429 the site.** The compare modal's search-as-you-type already
shares this bucket, so the exposure exists today — autosuggest makes it
certain.
Fixing the keying is therefore part of this work, not a follow-up.
## 1. Rate-limit keying
Both environments sit behind Cloudflare (`server: cloudflare`, `cf-ray` present
on staging and production). Cloudflare sets `CF-Connecting-IP` on every request
to the origin and **overwrites any client-supplied value**, which makes it
trustworthy in a way a parsed `X-Forwarded-For` chain is not.
```python
def client_key(request: Request) -> str:
"""Rate-limit bucket: the real caller, not the proxy in front of them."""
cf = request.headers.get("cf-connecting-ip")
if cf:
return cf.strip()
xff = request.headers.get("x-forwarded-for")
if xff:
return xff.split(",")[0].strip()
return get_remote_address(request)
```
`nextjs-app/app/api/[...path]/route.ts` already forwards every inbound header
except `host` and `connection`, so `CF-Connecting-IP` reaches the backend with
no proxy change.
**This header is trustworthy only for traffic that actually passed through
Cloudflare, and nothing in the application can verify that it did.** An earlier
draft of this section claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the `X-Forwarded-For` fallback was forgeable.
That was wrong. Cloudflare does overwrite the header *on requests it handles* —
but a caller reaching the origin directly sets whatever it likes, and this
process cannot distinguish an edge-set header from an attacker-set one. Both
headers are equally forgeable in that scenario.
The consequence is sharper than a weakened defence. An attacker rotating
`CF-Connecting-IP` per request mints a fresh rate-limit bucket every time and
evades per-client limits entirely — including on the DataFrame-heavy
`/api/schools`. Against abuse that is *worse* than the shared bucket it
replaced, which at least capped everyone at 60/minute together.
Two mitigations, and they are not interchangeable:
1. **The real fix is at Cloudflare** — Authenticated Origin Pulls, or an origin
firewall that refuses connections not from Cloudflare's ranges. Only the
edge can vouch for its own header. This is infrastructure work and is not
part of this change; it is the thing that makes the header mean anything.
2. **The ceiling in §1.1 bounds what evading the keying can achieve** while
that remains open. It does not make the header trustworthy — it makes
trusting it survivable.
The backend being unreachable from outside the Docker network is a real second
layer, but it depends on the ingress path in front of the frontend, which this
design does not control and should not assume.
### 1.1 The ceiling, which is back
The shared bucket was acting as an accidental global throttle on a
single-process uvicorn backend that filters a 25,000-row DataFrame in-process.
Correct per-user keying removes it: the origin becomes reachable at 60/min *per
user* rather than 60/min in total, and — per above — at an unbounded rate by
anyone willing to rotate a header.
An earlier draft dropped the in-app ceiling, arguing it belonged at Cloudflare.
That argument assumed the keying was sound. It is not, so the ceiling is
load-bearing rather than redundant, and it ships here:
`GlobalRateLimitMiddleware` counts all `/api/` requests in a fixed 60-second
window against `global_rate_limit_per_minute` (3000), independent of any client
identity, and refuses with a 429 that names capacity rather than the client —
an operator has to be able to tell "one noisy client" from "the origin is
saturated". It is registered last so it is outermost: a ceiling that applies
after the expensive work has run is not a ceiling.
slowapi cannot express this. `default_limits` and `application_limits` are both
evaluated with the same `key_func`, making them per-client rather than global,
and `application_limits` only apply with `SlowAPIMiddleware` installed, which
this app does not use. Hence the explicit middleware — about thirty lines, and
obviously correct, which is what a backstop needs to be.
Requests from `127.0.0.1` are exempt. The container healthcheck runs
`curl http://localhost:80/api/data-info` from inside the container, and
starving it would fail the check, restart the container, and turn a load spike
into an outage loop. The exemption keys on the peer address, never the `Host`
header, which the caller sets.
3000/minute is an estimate, not a measurement, and worth revisiting against
real traffic.
### Per-user limits
Per-user fairness and origin protection are different jobs, and this design now
does both separately: the ceiling above for the origin, and per-route limits
for fairness. Conflating them is what produced the original behaviour, where
one bucket served the whole internet.
The existing 60/minute default is unchanged, and `/api/suggest` gets
120/minute. Both are estimates rather than measurements, and are a starting
point to revisit once the keying is correct enough for real per-user traffic to
be visible — which it was not before, because everyone shared one bucket.
## 2. `GET /api/suggest`
A dedicated endpoint, not a mode of `/api/schools`.
The existing search path calls Typesense for URNs and then filters, ranks and
sorts the full in-memory DataFrame — a pandas pass per keystroke, holding the
GIL and blocking other requests in the same worker. Suggestions need none of
it: `urn`, `school_name`, `phase`, `school_type`, `local_authority`,
`postcode` and `ofsted_rating` are all already in the Typesense document
(`pipeline/scripts/sync_typesense.py`).
```
GET /api/suggest?q=<query>&limit=8
→ 200 {"suggestions": [
{"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school"}
]}
```
- **Under two characters** returns `{"suggestions": []}` with 200. The
keystroke path never returns an error for ordinary input.
- **Typesense unavailable** returns `{"suggestions": []}` with 200. There is
deliberately **no DataFrame fallback**: the substring scan `/api/schools`
falls back to is precisely the cost this endpoint exists to avoid, and a
silent 25,000-row scan per keystroke is worse than no suggestions.
- **`limit` is clamped** to 20. It is a public endpoint.
- **Rate limit `120/minute`** per client, not the default 60. A 200 ms
debounce tops out near 5 requests/second while someone is actively typing,
but averages far below that across a real search; 120 leaves headroom for
bursts while still bounding one client.
- **Local authority is part of the payload, not decoration.** There are many
schools called "St Mary's"; a suggestion list without the authority is
unusable for exactly the queries autosuggest is meant to serve.
### Caching
`CACHE_RULES` gains `("/api/suggest", (60, 3600, 86400))`. Prefix queries
repeat enormously across users and school names change once a year.
The client fetch must **not** use `cache: "no-store"`. The compare modal does,
and copying that pattern would throw away both the browser cache and the ETag
304s the existing `CacheAndETagMiddleware` already provides.
Both environments currently report `cf-cache-status: DYNAMIC` — Cloudflare
ignores the `Cache-Control` the API already sends, because it does not cache
dynamic paths by default. **A Cloudflare Cache Rule for `/api/suggest*` would
let the edge absorb most of this traffic and never reach the origin.** That is
a dashboard change, it is optional, and nothing here depends on it.
## 3. The combobox
This is an ARIA combobox, not a text input with a list underneath.
**Files.** `FilterBar.tsx` is already long. The work splits three ways:
`hooks/useSchoolSuggest.ts` owns fetching, debouncing and cancellation;
`components/SuggestList.tsx` owns rendering and ARIA; `FilterBar.tsx` wires
them to the existing input and form.
**Fetching.** 200 ms debounce; minimum two characters; an `AbortController`
cancels the superseded request on every keystroke. Cancellation is not an
optimisation — without it, a slow response for `"st"` can land after the fast
one for `"st marys"` and replace a correct list with a stale one.
**Suppressed during postcode entry.** The box takes a school name *or* a
postcode, and `isValidPostcode` already distinguishes them. Suggestions do not
appear once the value parses as a postcode.
**Keyboard.** `ArrowDown`/`ArrowUp` move the active option, `Escape` closes and
keeps the typed text, `Tab` closes. `Enter` **with an option active** navigates
to that school's page. `Enter` **with none active** submits the free-text
search exactly as it does today — the existing behaviour is preserved, not
replaced.
**ARIA.** `role="combobox"` with `aria-expanded` and `aria-controls` on the
input, `aria-activedescendant` pointing at the active option, `role="listbox"`
on the list and `role="option"` on each row.
**Both instances get it.** `HomeView` renders `FilterBar` twice — hero and
sticky — from one component, so there is one implementation.
## 4. Behind a flag
Flag `school_autosuggest`, declared in `backend/flags.py`, default off.
This is the most-used control on the site and the first change to it in a
while. `app/page.tsx` is an async server component, so it reads the flag and
threads it to `FilterBar` through `HomeView` — two prop hops, explicit, no
client-side flag read.
Off means the input behaves exactly as it does today: no listener, no fetch, no
markup. Not a rendered-then-hidden dropdown.
The rate-limit keying is **not** flagged. It is a correctness fix that should
apply whether or not autosuggest is on, and flagging it would mean shipping a
known-wrong limiter into production deliberately.
## 5. Analytics
`search_submitted` already carries `via: 'input'`. Selecting a suggestion fires
it with `via: 'suggestion'` plus the chosen `urn`, so the obvious question —
does this actually help, or do people ignore it — has an answer in the data
rather than an opinion.
## 6. Testing
**Backend.** `client_key` prefers `CF-Connecting-IP`, falls back through
`X-Forwarded-For` to the remote address, and two different values get two
different buckets. `/api/suggest` returns matches, returns empty below two
characters, returns empty and 200 when Typesense is unavailable, and clamps
`limit`. That it never touches the DataFrame is asserted by making
`load_school_data` raise and requiring the endpoint to answer anyway.
**Frontend.** The hook debounces, aborts superseded requests, and drops a
late-arriving response for a stale query. The list renders the ARIA
attributes. Keyboard navigation moves the active option; `Enter` on an option
navigates; `Enter` on none submits the search.
**E2E.** With the flag on, typing a known school name shows it and selecting it
lands on that school's page. With the flag off, no combobox markup exists.
Gated on the flag the same way the distance journeys are — read the observable
effect, since `/api/flags` is denied to the public.
## 7. Risks
**Removing the accidental throttle.** Covered in §1. Correct per-user keying
means the origin is reachable at 60/minute *per user* where it was 60/minute
in total, and no in-app global cap replaces it — that job goes to Cloudflare,
which is not done as part of this change. Until it is, a determined caller
with many source addresses can put more load on a single-process origin than
they can today. Against this site's traffic that is a theoretical risk rather
than a live one, but it is a real one and it is the price of the fix.
**Cloudflare bypass — the open one.** If the origin is reachable without
passing through Cloudflare, `CF-Connecting-IP` is attacker-controlled, and
rotating it per request defeats per-client limits on every endpoint. The
ceiling in §1.1 bounds the damage to the origin's total capacity; it does not
restore per-client fairness under attack, and it cannot. Closing this properly
means Authenticated Origin Pulls or an origin firewall restricted to
Cloudflare's published ranges — infrastructure work, outside this change, and
the single most valuable follow-up here.
**Typesense becomes user-visible.** Today a Typesense outage degrades search to
a slow substring match. With autosuggest it also means the dropdown silently
stops appearing. That is the correct failure — quiet, not broken — but it makes
Typesense health worth monitoring in a way it was not before.
## Out of scope
- **Place suggestions.** The 2,646 town, authority and outcode pages are a
strong candidate and would route people onto the pages W2 built, but they
live in the place registry rather than Typesense, so it is a second index and
a ranking rule for comparing two kinds of result. Worth its own change.
- **Postcode completion.** Would put postcodes.io in the keystroke path, with
its own latency and rate limits.
- **The compare modal.** It already has search-as-you-type. Converting it to
this component is a reasonable follow-up, not part of this.
- **Recent or popular searches.** No storage for either, and no evidence yet
that they are wanted.
@@ -0,0 +1,399 @@
# Destination Measures — Design
**Date:** 2026-08-28
**Status:** awaiting review
**Scope:** secondary school detail pages only
## Goal
Say what happened to a school's leavers after they left. Two sections on the
secondary template:
- **After Year 11** — every secondary, from the KS4 destination measures
- **After the sixth form** — sixth-form schools only, from the 16-18 measures
This replaces the "Post-16 destination data coming soon" placeholder standing in
`nextjs-app/components/school/SecondaryAdmissionsSection.tsx:117` since the exam
phase taxonomy work, and fills the `ks5_destinations_pct` slot specified but
never built in `2026-07-07-exam-phase-taxonomy-design.md:201`.
Mockup, with all three data states live:
<https://claude.ai/code/artifact/5be149d6-252f-473c-9a4f-4c36b05161b0>
## The finding that shapes everything
**Suppression is per cell, and the cells sum to the cohort.**
DfE withholds a figure it considers disclosive by writing `c`. It does this at
the level of an individual destination category, not the whole school, and it
publishes the cohort total alongside. The categories form a clean partition. So
where exactly one category is suppressed, subtracting the published ones from the
cohort recovers it exactly.
Verified against three real schools in the 2022/23 file:
| School | URN | Withheld | Recovers to |
|---|---|---|---|
| North East Futures UTC | 145900 | School sixth form | **3 pupils** |
| Whitley Bay High School | 108638 | Further education | **18 pupils** |
| St Matthew's RC High School | 148389 | School sixth form | **4 pupils** |
Those are the precise numbers the `c` exists to hide, and in a random 400-school
sample **22% of mainstream secondaries** have exactly one suppressed category in
their disadvantaged group. This is the normal case, not an edge case.
Three rules follow, and everything else in this document is downstream of them.
**R1 — Never *publish* enough to derive a remainder.**
An earlier draft of this rule said "never *render* a derived remainder", and
that was the defect code review caught in PR #137. Not drawing a number does
nothing to stop it being computed: `GET /api/schools/{urn}` is public and
unauthenticated, so anything in the payload is published whatever the UI
chooses to draw. The rendering guards shipped; the payload still carried the
cohort and every published category, and `cohort - sum(published)` returned
Whitley Bay's withheld figure exactly.
The rule is therefore about the serialiser, and the UI guards are a second line
of defence behind it. Two identities have to be closed:
- within a pupil group the categories sum to the cohort, so a group with
exactly **one** suppressed category gives it away;
- across groups, disadvantaged + other = all for every category, so a category
suppressed in exactly **one** of the three gives itself away.
`_mask_for_disclosure` applies DfE's own answer — secondary suppression —
withholding a companion cell until every row and every column hides either none
or at least two. It iterates, because each new suppression can break the other
identity, and terminates because cells are only ever added.
The companion must carry pupils. Suppressing a zero looks like secondary
suppression and protects nothing: the residual still equals the original
withheld figure.
Where no companion can do the job — a sparse cohort whose every other category
is `not_applicable`, routine in special schools and alternative provision — the
pupil group is **dropped from the payload entirely**. A first version simply
returned at that point with the violation intact and no signal, which review
caught: a disclosure-control pass that fails silently is worse than none,
because everything downstream trusts it. The function now cannot terminate
except in a state where `disclosure_invariant_holds()` is true, and an
exhaustive test sweeps all 81 suppression patterns of a four-category group to
prove it.
Measured cost on the 400-school sample: the all-pupils bar survives on **94%**
of mainstream secondaries rather than 100%. That is the price of not
republishing what DfE withheld.
**R2 — Never aggregate across a suppression boundary.** Summing published
components to fill a gap is R1 with extra steps.
DfE's own aggregates (`Sustained education destination`, `Sustained education,
employment & apprenticeships`) are ingested but **not served**. An aggregate
spanning exactly one suppressed component names it, and nothing renders them
today — an unused field that leaks is not a trade-off worth carrying. They can
be re-added with their own guard if the fallback ladder is ever built.
**R3 — The three pupil groups are one disclosure surface, not three.**
Disadvantaged and Not-known-to-be-disadvantaged partition All pupils, so
rendering any *two* of them recovers the third. Where a category is suppressed in
the disadvantaged group, it must therefore also be withheld from **all other
pupils** — the all-pupils view is the primary one and keeps it.
This costs almost nothing, because DfE already applies the same masking: across
the sample, 493 of 498 suppressed disadvantaged cells were suppressed in the
other group too. The mart enforces the remaining 5, which fell on 2 schools of
262. **The all-pupils bar is unaffected** — masking the whole page wherever the
disadvantaged group is thin would remove the bar from 80% of schools, and is not
what this rule says.
R1 and R2 both hold within a group and still leak across the switch, which is why
R3 is stated separately.
### The convention that would break this quietly
`macros/safe_numeric.sql` coerces every EES sentinel — `z`, `c`, `x`, `q`, `u` —
to `NULL`, deliberately and correctly for attainment, where "suppressed" and "no
data" are equally unrenderable. Here they are not the same thing: one must print
*withheld*, the other must print nothing at all, and the difference is what keeps
R1 enforceable.
**`safe_numeric` must not be used on destination counts.** The staging model
keeps the sentinel in a companion status column. This is the single most likely
way for this feature to regress into a disclosure, so it gets its own dbt test.
## What is actually available
Measured against the EES public API (open, no key). Both datasets carry
`geographicLevel: School` with `urn` on every location option, so the join to
`dim_school` is direct.
| | KS4 | 16-18 |
|---|---|---|
| Dataset id | `019d4f41-22d1-71b2-a1a7-f3b91026815b` | `019d4e73-6440-7523-b60c-bfab1ad4a30d` |
| Rows | 1,871,739 | 3,862,658 |
| Institutions | 4,946 | 3,065 |
| Time periods | 2009/10–2022/23 | 2016/17–2022/23 |
**Destination categories (KS4).** School sixth form · Sixth form college ·
Further education · Other education destination · Sustained apprenticeships (with
level breakdown) · Sustained employment destination · Not recorded as a sustained
destination · Activity not captured. Plus the aggregates `Sustained education
destination` and `Sustained education, employment & apprenticeships`.
**16-18 adds** UK higher education institution and FE split by level, which is
what makes the post-16 section worth having.
**Breakdowns.** `Disadvantage Status` gives Disadvantaged / Not known to be
disadvantaged / Total — exactly the three-way switch. Sex, ethnicity, FSM status,
prior attainment and SEN provision also travel in the same table; we ingest none
of them.
**Indicators.** Both counts and percentages, plus the cohort size. Bar widths use
the counts — the published percentages do not sum to 100.
### Coverage, and what degrades
Random 400-school sample, 2022/23, mainstream secondaries (n=262):
| View | As published by DfE | After R1–R3 masking | Consequence |
|---|---|---|---|
| All pupils, all categories | 100% | **94%** | Bar works nearly everywhere |
| Disadvantaged, headline rate | 95% | 95% | Gap panel works |
| Disadvantaged, three grouped cards | 68% | 68% | Degrades card by card |
| Disadvantaged, all six categories | 20% | **20%** | Bar unusable for this group |
The middle column is what the site actually serves. Masking costs the
all-pupils bar on 6% of mainstream secondaries — those are schools where a
category was suppressed in exactly one pupil group and no non-zero companion
existed below the all-pupils row.
Special schools and alternative provision are far worse: 13% and 41% respectively
have the whole cohort suppressed even for all pupils. The empty state is
load-bearing, not defensive.
## The display
Question-led. Three cards over one bar, with the cards acting as a lens on the
bar rather than a summary beside it — hovering a card dims the bar, table and
England reference to the categories that card is built from. The full mockup is
linked above; what matters for implementation:
**The headline is not the sustained rate.** That figure sits between 92% and 97%
for nearly every school in England. The mix is what varies, so the mix leads.
**The grouping is ours, not DfE's.** "Academic route" = school sixth form +
sixth-form college; "College" = FE and other colleges; "Work" = apprenticeship +
employment. This is the most arguable thing on the page, so it lives in one place
in `lib/destinations.ts`, is explained in a tooltip, and is reversible in one
edit.
**The absence is hatched neutral, never a colour.** "Activity not captured" means
no record in the sources DfE holds — it includes independent schools, moving
abroad and private training. Colouring it as a bad outcome would be a factual
error rendered in CSS. The hatch also fixes a real contrast problem: neutral
against the employment blue failed CVD separation at ΔE 7.6, and texture is the
secondary encoding that rescues it. Every other adjacent pair clears ΔE 10.9
under protanopia.
**Colour tokens.** Education is one hue in three steps (school-like to
college-like); apprenticeship and employment are separate hues. Six new tokens in
`globals.css`, defined in both themes, per the existing token discipline.
**The disadvantage split rides the same control.** One visualisation serving
three cohorts, with the England reference repointing to the matching national
group. The gap statement stays visible below the bar whatever is selected,
because a gap nobody clicks on is a gap nobody sees.
## Data model
### Extraction
A new `tap-uk-ees-destinations` extractor, separate from `tap-uk-ees`. The
existing tap downloads a release ZIP and reads a CSV inside it; the destinations
files are far larger than we need and the query API filters server-side, so this
one POSTs to `/v1/data-sets/{id}/query` and pages through results.
With every dimension pinned — destination measures, disadvantage status, sex
Total, characteristic topic Total — one year returns **252,610 rows** across all
geographic levels. Three school-level years is comfortably tractable.
Pinning is mandatory, not an optimisation: leaving the characteristic dimensions
unconstrained returned 45 rows where 9 were wanted, because every breakdown
shares one table.
The tap emits the raw value as text. **It does not coerce `c`.**
### Staging
`stg_ees_ks4_destinations` / `stg_ees_ks5_destinations`. Each raw value becomes
two columns:
```sql
case when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then raw::numeric end as pupils,
case
when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then 'published'
when lower(trim(raw)) = 'c' then 'suppressed'
else 'not_applicable'
end as status
```
### Marts
`fact_ks4_destinations` and `fact_ks5_destinations`, **long format**:
```
urn, year, pupil_group, destination_category, cohort_pupils, pupils, percentage, status
```
This departs from the wide house pattern (`fact_ks4_performance` and friends) on
purpose. `pupil_group` is a genuine third dimension; going wide would need three
sets of every column, and R2 is far easier to test on rows than on columns.
Roughly 8 categories × 3 groups × 4,946 schools × 3 years ≈ 356k rows.
`fact_destination_national` carries the same grain for England, so the page's
England reference repoints with the switch.
### dbt tests
- `assert_destinations_no_derived_remainder` — for every (urn, year,
pupil_group) with exactly one suppressed category, assert no aggregate row
exists that would let the residual be recovered. **This is the R1 guard.**
- `assert_destinations_group_masking` — for every (urn, year, category), if the
disadvantaged group carries `suppressed`, so does the other-pupils group.
**This is the R3 guard**, applied in the mart so no consumer can reach an
unmasked combination.
- `assert_destination_status_null_agreement` — `pupils is null` wherever
`status != 'published'`, and never null where it is.
- `assert_destinations_join_dim_school` — no orphaned URNs, matching the
existing `assert_no_orphaned_facts` pattern.
## API
`GET /api/schools/{urn}` gains a `destinations` block:
```json
{
"ks4": {
"cohort_year": "2022/23",
"published": "2026-04",
"groups": {
"all": { "cohort": 180, "categories": [ … ], "aggregates": { … } },
"disadvantaged": { … },
"other": { … }
}
},
"ks5": { … }
}
```
Each category carries `pupils`, `percentage` and `status`. **The serialiser never
emits a computed remainder**, and a backend test asserts that a group containing a
suppressed category serialises no total that closes the gap.
`null` for the whole block where nothing is published — the frontend renders the
empty state from its absence, not from a sentinel.
## Frontend
| File | Kind | Job |
|---|---|---|
| `lib/destinations.ts` | pure | Category list, the academic/college/work grouping, `canAggregate()` enforcing R2, percentage derivation from counts |
| `components/school/DestinationsSection.tsx` | server | Section shell, renders **all pupils** into the HTML |
| `components/school/DestinationsView.tsx` | client | Cohort switch, card↔bar linkage |
| `components/school/Post16DestinationsSection.tsx` | server | Year 13 section, sixth-form schools only |
| `app/globals.css` | tokens | Six destination colours, both themes |
Server-first matches the directory's existing discipline — every component in
`components/school/` is a server component except `AdmissionsViewToggle`, which
is the precedent this follows. All-pupils figures are in the HTML for crawlers
and for no-JS; only the switch and the hover linkage need the client.
`lib/schoolSections.ts` gains `hasKs4Destinations` / `hasKs5Destinations` flags
and the nav items, following the existing `computeSchoolFlags` pattern.
**Placement** on the secondary template: GCSE results → After Year 11 → After the
sixth form → admissions. Destinations follow attainment because they answer "and
then what happened".
**Dating.** The latest destination year is 2022/23, published April 2026, while
the site's newest KS4 year is 2024/25. The section header states its own cohort
year, or it reads as stale data next to the GCSE section above it.
## Edge states
| State | Frequency | Behaviour |
|---|---|---|
| Whole cohort suppressed | 13% of special, 41% of AP | Section renders the explanation, no chart |
| Some categories withheld | 80% of disadvantaged views | Cards degrade individually; **no bar**; table marks withheld rows |
| Disadvantaged group suppressed entirely | 5% | Switch drops to two options, gap panel not rendered |
| No sixth form | — | Post-16 section not rendered at all — absence is correct, a "no data" placeholder would imply something is missing |
| School too new | — | "First figures expected in 2026", not a bare no |
## Testing
Per CLAUDE.md, user-facing behaviour extends `e2e/` in the same PR.
**Unit** — `lib/destinations.ts` is where R1 and R2 live, so it carries the
heaviest tests: `canAggregate()` refuses a group containing one suppressed cell,
allows one spanning two, and the bar builder refuses to emit segments for any
group with suppression. These are the tests that must fail loudly if someone
later "fixes" a gap in the chart.
**dbt** — the three tests above.
**Backend** — the serialiser emits no closing total for a partially suppressed
group.
**E2E** — a school with full data renders three cards and a bar; a school with a
partially suppressed disadvantaged group renders the withheld state and **no bar
element**; a suppressed school renders the explanation; a school with no sixth
form renders no post-16 section.
Note the staging caveat: mart changes are inert until the Airflow pipeline runs,
and the staging E2E gate runs post-merge.
## Out of scope
- **Compare view and rankings.** The long mart shape supports both; neither is
built here. Flagged because "% to a school sixth form" is a plausible rankings
metric and the mart shape should not have to change to allow it.
- **Ethnicity, sex, SEN and prior-attainment breakdowns.** Available in the same
file, ingested deliberately not at all — each is a separate editorial decision
about what a school page should assert.
- **Longer term destinations** (3 and 5 years out) and **Progression to higher
education** — separate publications, worth a later look for sixth forms.
- **Primary schools.** No KS2 destination measures publication exists; DfE
tracking starts at KS4. Naming the secondaries a primary's leavers go to needs
the National Pupil Database, which is not publishable at that grain.
## Risks
**A later change reintroduces the disclosure.** The likeliest routes are
applying `safe_numeric` to a destination column for consistency, adding a
`coalesce` in a mart, or — as happened in review — enforcing a disclosure rule
at the rendering layer instead of the publishing layer. Mitigation is the dbt
tests plus `backend/tests/test_destinations_api.py`, which reconstructs the
residual the way an attacker would and asserts it no longer resolves.
**The two-year lag reads as staleness.** Mitigated by dating the cohort in the
section header rather than only in a tooltip.
**Sixth-form retention will be misread.** "41% went to a school sixth form" says
nothing about *which* school. The published file reports destination type, never
destination institution. Copy must never imply "stayed on here", and the tooltip
should say so.
**Section length.** The secondary template is already long and this adds two
sections. If it becomes a problem the post-16 section is the one to collapse
behind a disclosure, not the Year 11 one.
## Open questions
1. Is the disadvantage split its own section or a sub-block inside the
destinations section? Modelled as a sub-block; it is the most differentiating
figure on the page and the most easily misread on a small cohort.
2. Do we ingest the apprenticeship level breakdown (intermediate / advanced /
higher) now, or collapse to one apprenticeship figure and revisit? Collapsed
in this design.
@@ -0,0 +1,368 @@
# Giving schoolcompare a human author: an About page and a blog
**Date:** 2026-09-02
**Status:** Design — awaiting review
**Scope:** A named author for the site, an `/about` page, and a Payload-CMS-backed
blog at `/blog`.
## Why
The site reads as synthetic. Not because of its tone, but because of three
specific absences:
1. **Nobody is accountable for the numbers.** There is no author, no statement
of why the site exists, and no one who can be wrong. The only human trace on
the entire site is `contact@schoolcompare.co.uk` in the footer.
2. **No visible judgement.** Every figure is presented as though it fell out of
a machine. Hundreds of editorial decisions went into this codebase — which
metrics to show, when a benchmark is invalid, what to suppress — and not one
of them is visible to a reader. `isSpecialSchool()` silently drops the
England comparison for special schools and PRUs because that comparison is
meaningless; nowhere does the site *say* so.
3. **The voice is institutional third person.** "schoolcompare brings it all
into one place." "Built for parents, governors, journalists." That is
brochure register, and it is precisely the register that machine-generated
content defaults to.
There is a second, independent reason. The SEO programme
(`2026-08-20-seo-programme-design.md`) defines eight workstreams and none of
them address E-E-A-T or authorship. School performance data is YMYL territory;
an anonymous site republishing DfE figures has no authorship signal at all. This
work fills that hole, and the blog gives W6 (explainer content) somewhere to
live.
### The failure mode to avoid
The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about
education!" — reads as *more* synthetic than the current coldness. Manufactured
warmth is a stronger machine-tell than plain institutional voice. Everything
here has to be specific, occasionally awkward, and willing to be unflattering,
or it makes the problem worse.
## Positioning
The author is **Tudor**: first name only, real photograph, no surname, no
employer named.
The credibility claim is deliberately **not** educational expertise. The About
page states plainly: *"I'm not an education expert."* Authority comes from two
things that are actually true:
- **Experience.** A parent going through primary admissions in south-west London
right now. Google's E-E-A-T leads with Experience, and lived experience of the
thing is exactly what the DfE's own service lacks.
- **Method.** Every number's provenance is stated, so a reader can check the
site rather than trust it.
This is more durable than borrowed expertise: it cannot be undermined by someone
noticing the author has no teaching qualification.
**Consequence for the design.** A `Person` entity with no surname is a weak
search signal and cannot be corroborated off-site. The credibility load
therefore shifts onto the methodology being visibly rigorous. That is a design
constraint, not a caveat — it is why the About page carries a substantial
"how this is built and where it can be wrong" section rather than a short bio.
### Voice rules
Applied to About and every post. Recorded here so the voice does not drift.
- First person singular. "I built", not "we provide".
- Concrete over general. "when we were looking at schools in Wandsworth" beats
any amount of stated warmth.
- State limits before someone else finds them. Every post that presents a
metric says what it does not show.
- No mission statements, no "passionate about", no invented team.
- No em dashes. One of the clearest tells of machine-written prose, which is
the exact problem this work exists to fix.
- Short sentences. The existing code comments in this repo are already written
this way; the prose should match.
## Scope
**In:**
- `/about` — a coded page (not CMS-managed).
- `/blog` and `/blog/[slug]` — Payload-backed, with an index and post pages.
- Payload CMS installed into the existing Next application.
- Footer and navigation links to both.
- `Person`, `Organization`, `BlogPosting`, `BreadcrumbList` JSON-LD.
- RSS feed and sitemap integration.
- One first post, so the blog does not launch empty.
**Out (deliberately):**
- Rewriting existing homepage/how-it-works copy into first person. Worth doing,
but it would double the review surface of this PR. Separate change.
- In-product signed notes on school pages (the "distributed humanity" idea).
Revisit once About and the blog exist.
- Comments, newsletter, author accounts beyond one.
- A team page. There is no team.
## Architecture
### Topology
Payload 3 installs **into the existing Next application** and serves `/admin`
from the same container. One image, one deploy, no new service. This is
Payload 3's native model and it makes on-demand revalidation trivial, because
the CMS hooks run in the same process as the Next cache.
Accepted costs: the public site's image now carries Payload, so a CMS security
patch redeploys the whole site; and the image grows substantially.
### Two collisions that must be handled
**1. `/api` is already taken.** `app/api/[...path]/route.ts` is a catch-all that
proxies `/api/*` to FastAPI at runtime. Payload's default API route is also
`/api`. Left alone, these fight, and the failure is not clean — the catch-all
would swallow Payload's admin API calls and forward them to FastAPI.
Payload's API route is therefore remapped:
```ts
routes: { api: '/cms-api', admin: '/admin' }
```
with its route group at `app/(payload)/cms-api/[...slug]/route.ts`. The
`/cms-api` prefix must also be added to the FastAPI proxy's excluded-paths list
as a defensive second line.
**2. `next.config.js` is CommonJS.** Payload's `withPayload()` wrapper is ESM
only. The config must become `next.config.mjs`, converting `module.exports` to
`export default` and wrapping the export. All existing content — the standalone
output, `outputFileTracingIncludes`, the staging `X-Robots-Tag` header block,
the CSP — carries over unchanged. This is mechanical but it touches the file
that controls staging's noindex, so it needs care and an explicit test.
### Database
Payload uses the existing `sc_database` Postgres instance, in its **own
`payload` schema**:
```ts
db: postgresAdapter({
pool: { connectionString: process.env.DATABASE_URL },
schemaName: 'payload',
})
```
The frontend container is already on the `backend` Docker network, so it can
reach `sc_database:5432` with no networking change. It needs a new
`DATABASE_URL` environment variable.
Schema isolation is not cosmetic. `public` currently holds the application
tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists
to drop and reimport. Blog content living in its own schema means no data
pipeline operation can destroy it.
**Verified 2026-09-02** (this was an open question when the spec was written).
`--drop` calls `run_full_migration()` in `backend/migration.py`, which drops
exactly two tables by name:
```python
ks2_tables = ["school_results", "schools"]
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
```
There is no `Base.metadata.drop_all()` anywhere in `backend/`, and no
`DROP SCHEMA`. The only other drop is `_apply_schema_drops()`, a single
schema-qualified `DROP TABLE IF EXISTS marts.fact_parent_view CASCADE`.
Nothing sets `search_path`, so the SQLAlchemy metadata resolves to `public`,
and `inspector.get_table_names()` does not even enumerate other schemas.
So the guarantee is stronger than schema isolation alone: `--drop` targets two
named tables that Payload does not have, and would not reach `posts`, `media`
or `users` even if they shared a schema. The `payload` schema remains the right
choice — it protects against a *future* broadening of that script rather than
today's behaviour — but the safety claim rests on verified code, not on
assumption.
Putting CMS tables in this instance is consistent with existing practice —
Airflow already stores its metadata there.
### Migrations
Payload's Postgres adapter auto-pushes schema in development and requires
explicit migrations in production. Use `prodMigrations`, which runs pending
migrations during server initialisation:
```ts
db: postgresAdapter({ /* ... */, prodMigrations: migrations })
```
This is preferred over a one-shot init container (the `airflow-init` pattern)
because the app is a single long-running process and there is no ordering
problem to solve. Migration files are generated with `payload migrate:create`
and committed, so schema changes travel through the same PR and staging gate as
code.
### Media
Uploads go to a Docker named volume, consistent with `postgres_data`,
`typesense_data` and `airflow_logs`.
- `staticDir` must be an **absolute** path in Payload 3: `/app/media`.
- The container runs as `nextjs` (uid 1001). The Dockerfile must
`mkdir -p /app/media && chown nextjs:nodejs /app/media` **before** the volume
is mounted, or Docker will create the mountpoint root-owned and every upload
will fail with EACCES.
- `sharp` moves from `devDependencies` to `dependencies` — Payload needs it at
runtime to generate `imageSizes`.
- The volume must be added to the backup routine alongside Postgres. A blog
post's images are not reproducible from the pipeline.
### Rendering
**Constraint:** CI builds the image with no database reachable. Blog pages
therefore cannot use build-time `generateStaticParams` — that would either fail
the build or bake in an empty post list.
Instead: ISR. Post and index pages declare a `revalidate` window and render on
first request, with Payload `afterChange` / `afterDelete` hooks calling
`revalidatePath('/blog')` and `revalidatePath('/blog/' + slug)` for immediate
publication. Because Payload runs in the same process, the hook calls
`revalidatePath` from `next/cache` directly — no webhook, no shared secret.
The ISR cache lives on container disk and is cleared by a redeploy. For a
single container serving a handful of posts this is fine.
### Collections
- **`posts`** — `title`, `slug`, `publishedAt`, `excerpt`, `heroImage`
(relation to `media`), `content` (Lexical rich text), `seo` group
(`metaTitle`, `metaDescription`), `_status` (drafts enabled).
- **`media`** — upload collection, `alt` required, `imageSizes` for thumbnail
and hero widths, public read access.
- **`users`** — Payload's auth collection. One account. Public creation
disabled.
Drafts are enabled so posts can be written over several sittings and previewed
before publication.
**Payload Blocks** are how posts embed live product components — a real trend
chart or comparison table inside a post, rendered from live data rather than
screenshotted. This is the main thing the CMS has to earn back against
file-based MDX, and it directly serves the goal: showing judgement in context.
Ship with one block (a callout/aside for "what this number doesn't tell you");
add a live-chart block once a post needs it.
### Security
`/admin` is the first authenticated surface on this site. Public, hardened:
- `PAYLOAD_SECRET` — long, random, set in the Portainer stack environment, never
committed. The same variable must exist in staging with a *different* value.
- Strong unique password on the single admin account.
- Login rate limiting via Payload's `maxLoginAttempts` / `lockTime`.
- `X-Robots-Tag: noindex, nofollow` on `/admin/*` and `/cms-api/*`, and a
`robots.ts` disallow. The admin panel must never be indexed.
- Public user creation disabled; no open registration.
- Verify the existing CSP `frame-ancestors` directive does not break the admin
panel.
Residual risk, accepted: a future Payload authentication CVE is live against the
public internet. Mitigation is prompt patching, which the staging→prod pipeline
already supports. If this becomes uncomfortable, restricting `/admin` at the
proxy to LAN/VPN is a one-line change later.
Staging note: staging runs the same image on `stx.`, so it gets its own admin
panel and its own database. It must have its own `PAYLOAD_SECRET` and its own
credentials — never production's.
## Deployment changes
- `nextjs-app/Dockerfile` — create and chown `/app/media`; ensure Payload's
admin bundle and `sharp` survive standalone output file tracing.
- `docker-compose.portainer.yml` and the staging equivalent — add
`DATABASE_URL` and `PAYLOAD_SECRET` to the `frontend` service, add a
`payload_media` volume mounted at `/app/media`, and add
`depends_on: sc_database`.
- Document both new environment variables in the compose header comment block,
which is where this stack records its configuration.
## SEO
- `Person` (Tudor, with photo) and `Organization` JSON-LD on `/about`.
- `BlogPosting` + `BreadcrumbList` on post pages, with `author` referencing the
same `Person`.
- Canonical URLs on `/blog` and every post.
- Posts and `/about` added to the existing sitemap (`app/sitemap.xml/route.ts`
and `app/sitemaps/[...parts]`). Post URLs come from Payload at request time.
- RSS feed at `/blog/rss.xml`.
- Footer links to both pages, under a new "About" column.
**Navigation is deliberately left alone.** `Navigation.tsx` renders a bottom tab
bar on mobile that already carries four items (Search, Compare, Rankings,
Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are
both lower-intent than any of the four. Both live in the footer; About
additionally gets a byline link from every post, which is where a reader who
cares actually asks the question. Revisit only if analytics show people hunting
for it.
## Testing
Unit (Jest):
- Post rendering, including a post with no hero image and one with no excerpt.
- Slug generation and collision handling.
- JSON-LD shape for `BlogPosting` and `Person`.
- The `next.config.mjs` conversion preserves the staging `X-Robots-Tag` rule —
this guards the riskiest mechanical change in the plan.
E2E (Playwright, `e2e/`, required by CLAUDE.md for user-facing change):
- `/about` renders, shows the author name and photo, and is reachable from the
footer and nav.
- `/blog` lists at least one post; clicking through reaches the post.
- A post page renders title, date, body and byline.
- `/admin` responds with `noindex` and does not leak a stack trace when
unauthenticated.
Note the known constraint: new journeys cannot be proven in PR checks, because
the staging E2E gate runs post-merge.
## Risks
| Risk | Mitigation |
|---|---|
| `next.config.mjs` conversion silently drops the staging noindex header, making staging a crawlable duplicate | Unit test asserting the header rule; verify on staging before promotion |
| Payload API route collides with the FastAPI `/api` proxy | Remap to `/cms-api`; add to the proxy's exclusion list |
| Media volume mounts root-owned; all uploads fail with EACCES | `mkdir`+`chown` in the Dockerfile before the mount; test an upload on staging |
| Build fails or bakes empty content because CI has no DB | No build-time DB access; ISR only |
| A pipeline `--drop` destroys blog content | Separate `payload` schema; verify `--drop` blast radius before building |
| Media volume not backed up; images unrecoverable | Add `payload_media` to the backup routine |
| Payload auth CVE exposed publicly | Prompt patching; proxy restriction available as a fallback |
| Blog launches empty or goes stale | Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule |
## Sequence
Each step is independently reviewable and mergeable.
1. **Payload foundation** — install, `next.config.mjs` conversion, `payload`
schema, `/cms-api` remap, `users` collection, `/admin` hardening, compose and
Dockerfile changes. No public-facing change yet. Verify on staging that the
site is unchanged and `/admin` works.
2. **`/about`** — coded page, photo, `Person`/`Organization` JSON-LD, footer and
nav links, e2e journey. Independently valuable and does not depend on the
blog.
3. **Blog** — `posts` and `media` collections, `/blog` index and post pages, ISR
plus revalidation hooks, RSS, sitemap, structured data, e2e journeys.
4. **First post** — written in the admin panel, published through the normal
flow, proving the whole path end to end.
Step 1 carries all the infrastructure risk and none of the visible benefit, so
it should be verified on staging carefully before step 2 starts.
## Dependencies on Tudor
- **A photograph.** Blocks step 2. Nothing else in the plan is blocked by it.
- **The first post's subject.** Blocks step 4 only. Suggested: what school
performance data cannot tell you — it demonstrates judgement, is genuinely
useful, and is the kind of thing an anonymous or machine-written site will not
publish.
- ~~Confirmation that `scripts/migrate_csv_to_db.py --drop` is schema-scoped.~~
**Resolved 2026-09-02** — verified in `backend/migration.py`; see the
Database section. No action needed.
@@ -0,0 +1,432 @@
# Other Schools Nearby — Design
**Date:** 2026-09-21, revised 2026-09-22
**Status:** revised after staging review
**Scope:** school detail pages, both phase templates
> **Revision, 2026-09-22.** The first build ranked by intake similarity and used
> distance as a tiebreak. On staging a Catholic primary showed six Catholic
> primaries, none of them close enough to be a real option, and omitted the
> community school down the road. Distance now decides the order and nothing
> else does; the tier system is gone. The reasoning is kept below rather than
> quietly overwritten, because the mistake is the instructive part.
## Goal
Give a school detail page an answer to the question every reader arrives with
after the results tables: *and what else is around here?*
Today a school page links outward to its place pages through
`components/school/NearbyPlaces.tsx` and nowhere else. It never links to another
school. This section adds that edge — up to six nearby schools of the same phase
and a comparable intake, three at a time in a carousel, each a crawlable link and
each addable to the comparison basket in one click.
Mockup, in both themes (drawn against the original tiered design, so its ledes
and chip fallbacks are one revision behind the copy specified below):
<https://claude.ai/artifact/168KdUMcfkUeGWW2FGjuec>
Source of the same page in the repo: `mockups/similar-schools-nearby.html`.
## The constraint that shapes everything
**A nearby school is not automatically a comparable school.**
The section's whole value is that a reader treats what it shows as a shortlist.
That makes every card an implicit claim that the school is a realistic
alternative, and there are three ways that claim goes wrong:
1. A **selective** school beside a non-selective one. Their intakes are
different by construction, so putting their Attainment 8 figures side by side
invites a conclusion the data cannot support.
2. A **special school, PRU or AP** beside a mainstream school. This is the same
error PR #70 fixed for the England benchmark, where Greenmead (URN 101099)
rendered "0% — 62 below England".
3. A **single-sex** school of the opposite sex. Not a weak match — not an option
at all.
So the design separates two kinds of fact, and never confuses them:
- **Hard filters** encode the claims above. They decide eligibility, and are
never relaxed at any distance, even if that means the section does not render.
- **Shared characteristics** — gender, religious character, selectivity —
describe how closely an intake resembles this school's. They are *reported on
the card and never ranked on*, so the reader weighs them rather than having
them weighed for them.
Everything below follows from that split. The revision at the top of this
document is what happens when the second kind is treated as the first.
## Selection algorithm
A backend helper, `_nearby_schools_payload(urn)` in `backend/app.py`, modelled
on the existing `_places_payload(urn)` and delegating to
`backend/nearby_schools.select_nearby(frame, urn)`, which operates on the cached
`load_latest_school_data()` frame — one row per URN, already carrying
`latitude`, `longitude`, `phase`, `gender`, `religious_denomination`,
`admissions_policy`, `school_type` and `status`.
### Hard filters
| Filter | Rule |
|---|---|
| Self | `urn` is excluded |
| Status | GIAS status must be open |
| Coordinates | both `latitude` and `longitude` present on both schools |
| Phase | same phase group via the existing `PHASE_GROUPS` map |
| Provision | special/PRU/AP match only each other |
| Selectivity | selective matches selective; non-selective matches non-selective |
| Gender | Boys never matches Girls; Mixed is compatible with both |
`PHASE_GROUPS` is reused rather than re-derived so an all-through school is
offered correctly on both the primary and secondary sides, exactly as it already
behaves in search.
The provision filter needs a backend counterpart to the frontend's
`isSpecialSchool()` in `nextjs-app/lib/utils.ts:897`, reading the same GIAS
establishment types through `backend/gias_codes.py`. The two must agree: a
school the frontend treats as special for benchmarking but the backend treats as
mainstream for matching would be dropped from its own England comparison and
then offered as a peer to a mainstream school on the next page along.
**Up to six cards, three visible.** Three fit the row; the rest are reached with
the carousel arrows. Two is the minimum that renders at all.
### Order: distance, and nothing else
The nearest eligible schools, closest first. Similarity does not enter the
ranking at any point.
**Why not, having built it the other way first.** The original design ranked by
tiers — same gender and faith within 3 miles, then same gender within 5, then
anything within 10 — and used distance only to order the result. Two things
followed, and both showed up on the first Catholic primary anyone looked at:
- A faith match at 2.9 miles outranked a community school at 0.3 miles. For a
primary, whose catchment is routinely under a mile, the far school is not a
weaker option; it is not an option.
- Because the row filled from the best tier before widening, three Catholic
schools within 3 miles were enough to fill all six slots with Catholic
schools. The stopping rule that produced this had been added to prevent the
*opposite* failure — padding a row with weak distant matches — and made this
one certain.
The premise was backwards. **Distance is a constraint and intake is a
preference.** A parent cannot act on a school outside their reach however well
it matches, and they are perfectly capable of noticing a shared denomination
for themselves if we show it to them. So similarity moved from the ranking to
the card: `shared` reports what a school genuinely has in common, and the reader
applies their own weighting.
The hard filters above were always where the defensibility lived. They are
untouched.
### Reach: a sanity bound, not a target
| Phase | Reach |
|---|---|
| Primary, middle deemed primary, all-through | 2 miles |
| Secondary, middle deemed secondary | 6 miles |
| 16 plus | 10 miles |
Ordering by distance already handles density — a school in inner London fills
all six slots inside a mile and never approaches the cap. The cap decides one
thing: what happens where the area is sparse. It differs by phase because
catchments do, and because people travel furthest for post-16.
**A primary with nothing inside two miles renders no section**, and that is the
intended answer rather than a gap. The alternative is a section headed "nearby"
listing a school four miles from a five-year-old.
**Past the sixth school, the rest are dropped without a count.** The section
does not try to be the list: `NearbyPlaces` sits directly beneath and already
leads to the place pages, which are built for browsing a full set and which the
school page exists to feed.
**Fewer than two results renders nothing.** Not an empty state, not a single
lonely card. The section is absent, the nav item is absent, and the page is
unchanged from before it existed.
### Distance
Straight-line, from the vectorised haversine already used for postcode search at
`backend/app.py:831`, computed over the ~27k-row frame in numpy. Reported to one
decimal place in miles, consistent with the rest of the site.
Straight-line distance is not road distance and is not measured from the
reader's home. The section says so in its disclosure rather than leaving the
reader to assume otherwise.
## API
`/api/schools/{urn}` gains a `similar_schools` array. Each row:
| Field | Notes |
|---|---|
| `urn` | for the link and the compare basket |
| `school_name` | link text |
| `distance_miles` | one decimal place |
| `school_type` | GIAS type, translated, for the card's meta line |
| `age_range` | for the meta line |
| `shared` | what this school genuinely shares with the subject; may be empty |
| `metric_value` | the phase-appropriate headline figure, or null |
| `metric_key` | `rwm_expected_pct` or `attainment_8_score` — see below |
| `metric_year` | the year the figure is from |
The metric follows the subject school's phase side, not the neighbour's own
phase, so a row of cards never mixes two scales. The secondary
side uses `attainment_8_score`; the primary side uses `rwm_expected_pct`. Where
the neighbour has no value for that key, the card reads "Not published" rather
than falling back to the other key.
Which side a school takes is decided once, in
`similar_schools.is_secondary_phase`, by membership of `PHASE_GROUPS["secondary"]`
minus all-through — never by testing for the substring "secondary", which misses
`16 plus` (GIAS phase 6) and hands a sixth-form college the primary bucket.
All-through is the exception in the other direction: `PHASE_GROUPS` lists it on
both sides, but it takes the primary metric.
This is usually the same thing as "the template the page renders", but not
always. `computeSchoolFlags` decides the template with that same substring test,
so a `16 plus` school renders `PrimarySchoolSections` while being matched —
correctly — against secondaries. The section therefore takes its lede noun from
the school's own phase rather than from its template, or it would print "Other
primary schools near <sixth form college>" above a row of secondaries.
Up to six rows of roughly 130 bytes each. It rides in the existing detail payload
rather than a new endpoint because the page already makes exactly one server
fetch for its data, and `/school/[slug]` regenerates at most weekly
(`revalidate = 604800`), so the per-request cost is paid once per school per
week.
**The key is absent, not null, on a backend that does not have this code.** The
frontend treats absent and empty identically, which is what allowed
`NearbyPlaces` to ship without a lockstep deploy of the two images.
`shared` is computed on the backend, beside the data it is derived from, not
re-derived on the frontend. Deriving it twice is how a card comes to claim
something the selection never established. An empty list is a real answer and
renders no chips: a bare card costs a school nothing but the likeness it does
not have, since the order was already settled by distance.
## Frontend
### Components
`components/school/SimilarSchoolsSection.tsx` — a server component wrapped in
the shared `Section` shell from `sectionShared.tsx`. It renders the heading,
the lede, the card grid, the footer CTA and one caption line. Every
card's title is an `<a>` to the school's canonical slug URL via `schoolUrl()`.
`components/school/AddToCompareButton.tsx` — calls `addSchool` from
`ComparisonProvider` and reports the selection with a `from: 'similar_schools'`
attribution, mirroring `addSchoolFromSearch` in `HomeView.tsx:442`.
`components/school/SimilarSchoolsCarousel.tsx` — the scroller and its arrows. It
takes the server-rendered cards as `children` and the server-rendered heading and
lede as a `header` prop, so those stay server components while the client
component owns only the ref, the scroll handler and the arrows' disabled state.
The split matters: the links — the part with SEO value and the part that must
work without JavaScript — are server-rendered into the initial HTML, and only
the basket interaction and the arrows are hydrated.
### The carousel
**Every card is in the initial HTML.** The arrows scroll a list; they never swap
a view. Six `<a>` elements are in the markup whether or not anything is
hydrated, which is the whole reason the section exists — a paginated widget that
mounts cards on click would put four of the six links beyond a crawler and
beyond a reader with no JavaScript.
So the scroller is a plain overflowing `<ul>` with `scroll-snap-type: x
mandatory`, and the arrows call `scrollBy` on it. With no JavaScript it
degrades to a horizontally scrollable row that still works by touch and by
trackpad. Three cards are visible at desktop width and two below 820px.
**Arrows appear only when there is somewhere to go** — that is, only when more
than three schools were found. Each disables itself at its own end of the
travel.
#### Below 640px the arrows go away
This follows [MOBILE.md](../../../MOBILE.md), which makes 360px the design
floor and mobile the primary target at ≥55% of traffic.
Kept in the heading's flex row at 360px, the two arrow buttons take 96px from a
328px card and crush the lede into a four-line column — measured, not guessed.
And swiping already does what they do. So below 640px the header becomes a
single column, the arrows are not rendered, one card shows at 86% width so the
next one peeks, and the affordance is carried by the right-edge scroll-fade that
MOBILE.md documents for exactly this case:
```css
mask-image: linear-gradient(to right, #000 calc(100% - 28px), transparent);
```
The fade lifts at the end of the travel, where there is nothing left to hint
at. That means the at-end state must be computed whether or not an arrow exists
to consume it — on mobile it drives the mask alone.
**Every interactive element clears 44×44px**, per MOBILE.md's iOS HIG check: the
arrow buttons and the add-to-compare button are both 44px, up from the 40px they
were first drawn at. A card title's own box is shorter than that, but its hit
area is the whole card through the `::after` overlay, so it passes on the target
that actually receives the tap.
**The edge test needs a tolerance, and this is not fussiness.** The scroller
carries 2px of padding so focus rings are not clipped, and scroll-snap treats
that padding as the first card's snap position: a scroller sitting at its start
reports `scrollLeft` of 2, not 0. Sub-pixel rounding moves it again at other
zoom levels. Testing `scrollLeft === 0` therefore leaves the back arrow live and
pointing nowhere on first paint — confirmed in the mockup before it was fixed.
Both ends compare against an 8px tolerance.
**Selecting a school must not move the row.** Adding to the basket re-renders
the footer; the scroll offset lives in the DOM rather than in React state, so
the carousel must not remount or reset on that render. A reader who ticks the
fifth school and is thrown back to the first has been punished for using the
feature.
### Placement and navigation
Rendered as the last section **inside** `SchoolDetailShell`, from both
`PrimarySchoolSections` and `SecondarySchoolSections`. Inside, not after, because
the sticky nav's scroll-spy locates sections with `document.getElementById` and
can only reach a section that lives in the shell.
`NearbyPlaces` stays where it is, outside the shell, immediately below. The
resulting order — this school, then similar schools, then the places containing
them — narrows before it widens, which is the order a reader leaves a page in.
`buildNavItems` and `buildSecondaryNavItems` both gain
`{ id: 'similar', label: 'Similar schools' }`, gated on the section rendering.
The id must match the `Section` id or the scroll-spy silently breaks.
### The comparison CTA
A plain `<a href="/compare?urns=…">`, built from this school's URN plus the
selected ones. `/compare` already parses `urns` from the query string
(`app/(frontend)/compare/page.tsx:55`), so this needs no new compare plumbing.
With nothing selected the CTA is disabled; the button also adds to the shared
basket so the site-wide comparison state stays consistent with what the page
shows.
## Copy, and what the section is allowed to claim
**The lede never claims an intake.** It reads "Other primary schools near X." —
one sentence, no variants. The earlier version varied the wording by tier, which
only existed to soften a claim the section should not have been making.
**The heading is "Other schools nearby", not "Similar schools nearby".** The
hard filters do guarantee a comparable set — same phase, same selectivity,
mainstream never beside special — but nothing ranks on likeness, so the heading
does not say it does. The nav item reads "Nearby schools" and the section id is
`nearby`.
**Chips state only what is shared, and may be absent entirely.** A card with
nothing in common renders no chip row rather than falling back to a filler.
Since chips no longer affect the order, an empty one costs that school nothing
except a claim it cannot support — and a Catholic parent scanning the row still
spots "Roman Catholic" on the card that carries it, and weighs it themselves.
**The neighbour's metric carries no valence colour.** Green and terracotta are
reserved site-wide for comparison against the England average. Colouring a
neighbour's figure against this school's would read as ranking the neighbours
against each other, which is precisely the endorsement this section must not
make. The figure sits in neutral ink above a plain "72% at this school"
reference line, and the reader draws their own conclusion.
**A missing figure reads "Not published".** Never 0, never blank, never an
em dash. This follows the same rule the rest of the detail page uses: a school
with no published result has not scored zero.
**There is no "how these are chosen" disclosure.** The method is visible in what
the section already shows — the phase in the lede, the shared characteristics on
each card, the distance above each name — and a collapsed panel restating it
earns less than the space it costs.
**One caption line survives, and only one:** that distances are straight-line
from the school and not road distance. This is not a method note. A reader who
sees "0.6 miles away" and takes it for the walk has been misled by us, and no
other element on the card corrects that. The remaining notes — that listing is
not a recommendation, that special schools only meet special schools — are
statements the selection rules already keep true without being narrated.
## Degradation
| Condition | Behaviour |
|---|---|
| `similar_schools` absent (older backend image) | no section, no nav item |
| fewer than 2 qualifying schools | no section, no nav item |
| this school has no coordinates | no section |
| the helper raises | returns `[]`; the page renders without the section |
The helper is wrapped so a failure inside it never 500s a page that is otherwise
complete — the posture `get_supplementary_data` already takes for its own
queries.
## Testing
**Backend**, in a new `backend/tests/test_similar_schools.py`, against a
synthetic frame rather than live marts:
- a selective school never returns a non-selective one, and vice versa
- a special school returns only special schools; a mainstream school returns none
- a Boys school never returns a Girls school; Mixed matches both
- closed schools and schools without coordinates are never returned
- results are ordered by distance ascending, always
- a faith match never outranks a closer school (the staging defect, pinned)
- the nearest eligible school is always present
- more than six qualifying schools returns the six nearest
- reach is capped per phase, and a primary beyond two miles returns `[]`
- an all-through school is offered on both phase sides
- a `16 plus` school is matched against secondaries and colleges, never primaries
- `is_secondary_phase` and `PHASE_GROUPS` agree on every GIAS phase value
- fewer than two qualifying schools returns `[]`
- distances match a hand-computed haversine for a known pair
**Frontend**, in `nextjs-app/__tests__`:
- the section renders nothing for absent, empty and single-row inputs
- the lede never claims a similar intake
- an empty `shared` renders no chips rather than a filler
- a null metric renders "Not published"
- the nav item appears only alongside the section
- every card is in the DOM, including the ones scrolled out of view
- arrows render only when more than three schools were found
jsdom has no layout, so `scrollWidth` and `clientWidth` are both 0 there and the
arrows' disabled state cannot be meaningfully asserted in Jest. That behaviour is
covered in the journey instead, against a real engine, rather than by a unit test
that would pass on a measurement that does not exist.
**E2E**, added to the existing journeys in `e2e/tests` in the same PR, per the
repository's rule on user-facing behaviour:
- the section renders on a known staging URN, with resolving links
- where arrows are present, the back arrow starts disabled and the forward arrow
moves the row
- selecting a school does not reset the scroll position
- add-to-compare reaches `/compare` with the expected `urns`
- at 360, 390 and 430px: no horizontal overflow, every interactive element in the
section clears 44×44px, and no arrows are rendered
The E2E gate runs after merge on this project, so these journeys are not
provable in the PR checks; the PR is verified on the unit tests, and the
journeys are confirmed on the post-merge staging run.
## Out of scope
- A map of the nearby schools. The section is a list; the page already has a map.
- Autoplay, dots, or an infinite loop on the carousel. It is a short list a
reader scans deliberately, not a banner competing for attention, and a row
that moves on its own is a row that moves while someone is reading it.
- Statistical neighbours on deprivation, size or cohort profile. If plain
distance proves too blunt, that is the trigger to move this computation into a
dbt mart — `select_nearby` is a deliberate seam for exactly that swap.
- Precomputing neighbours in `marts.*`. Rejected for now: a new mart is inert
until Airflow runs, so the feature would ship dark, and every tuning change to
the rules would become a pipeline round-trip instead of a deploy. The revision
at the top of this document is the argument for keeping that loop short.
- Any change to `/api/compare`, the compare page, or the comparison basket.
File diff suppressed because it is too large. Load diff
+41
View File
@@ -0,0 +1,41 @@
import { test, expect, Route } from '@playwright/test';
test('the deployed frontend and backend report the tested build', async ({ request }) => {
const response = await request.get('/release.json');
expect(response.ok()).toBeTruthy();
expect(response.headers()['cache-control']).toContain('no-store');
const identity = await response.json();
expect(identity.frontend).toEqual(identity.backend);
expect(identity.frontend.sha).toMatch(/^[a-f0-9]{40}$/);
expect(identity.frontend.build_id).toMatch(/^[a-f0-9]{32}$/);
if (process.env.EXPECTED_SHA) expect(identity.frontend.sha).toBe(process.env.EXPECTED_SHA);
if (process.env.EXPECTED_BUILD_ID) expect(identity.frontend.build_id).toBe(process.env.EXPECTED_BUILD_ID);
});
test('changing search while loading another page does not append old results', async ({ page }) => {
await page.goto('/?phase=primary');
await expect(page.getByRole('button', { name: 'Load more schools' })).toBeVisible();
let received!: (route: Route) => void;
const pending = new Promise<Route>(resolve => { received = resolve; });
await page.route('**/api/schools?**', async route => {
if (new URL(route.request().url()).searchParams.get('page') === '2') {
received(route);
return;
}
await route.continue();
});
await page.getByRole('button', { name: 'Load more schools' }).click();
const oldRequest = await pending;
const search = page.getByPlaceholder('School name or postcode').first();
await search.fill('secondary');
await search.press('Enter');
await page.waitForURL(/search=secondary/);
// A cancelled fetch may prevent route fulfilment altogether; either way,
// this deliberately late response must not become part of the new results.
await oldRequest.fulfill({ json: {
schools: [{ urn: 999998, school_name: 'P1 stale result sentinel', phase: 'Primary' }],
total: 2, page: 2, page_size: 1, total_pages: 2,
} }).catch(() => {});
await expect(page.getByText('P1 stale result sentinel')).toHaveCount(0);
await expect(page.getByRole('button', { name: 'Loading...' })).toHaveCount(0);
});
+591
View File
@@ -0,0 +1,591 @@
<title>Last distance offered — detail page mockup</title>
<style>
:root{
--bg-primary:#faf7f2; --bg-secondary:#f3ede4; --bg-card:#fff;
--text-primary:#1a1612; --text-secondary:#5c564d; --text-muted:#6d685f;
--accent-coral:#e07256; --accent-coral-dark:#b04a2e;
--accent-teal:#296f6f; --accent-teal-light:#3a9e9e;
--accent-gold:#c9a227; --accent-gold-text:#7a6800;
--coral-bg:rgba(224,114,86,.12); --teal-bg:rgba(45,125,125,.12); --gold-bg:rgba(201,162,39,.12);
--border:#e5dfd5; --shadow:0 2px 8px rgba(26,22,18,.06); --shadow-md:0 4px 20px rgba(26,22,18,.1);
--radius-sm:4px; --radius-md:8px; --radius-lg:16px;
--serif:'Playfair Display',Georgia,'Iowan Old Style',serif;
--sans:'DM Sans',-apple-system,BlinkMacSystemFont,'Segoe UI',sans-serif;
}
/* Mockup is a fixed light artefact — it mirrors the live app, which is light-only. */
*{margin:0;padding:0;box-sizing:border-box}
body{font-family:var(--sans);background:var(--bg-primary);color:var(--text-primary);line-height:1.6;
padding:44px 20px 110px;font-variant-numeric:tabular-nums}
.wrap{max-width:820px;margin:0 auto;display:flex;flex-direction:column;gap:0}
.pagehead h1{font-family:var(--serif);font-weight:600;font-size:clamp(26px,4vw,34px);letter-spacing:-.015em;text-wrap:balance}
.pagehead p{color:var(--text-secondary);margin-top:10px;font-size:15px;max-width:64ch}
.pagehead p + p{margin-top:8px}
.step{margin:60px 0 6px;display:flex;align-items:baseline;gap:10px;flex-wrap:wrap}
.step h2{font-family:var(--serif);font-size:20px;font-weight:600;letter-spacing:-.01em}
.step .where{font-size:12px;font-weight:600;letter-spacing:.05em;text-transform:uppercase;color:var(--accent-teal);
background:var(--teal-bg);padding:3px 9px;border-radius:999px}
.stepdesc{color:var(--text-muted);font-size:14px;margin-bottom:18px;max-width:66ch}
.stepdesc code{font-family:ui-monospace,SFMono-Regular,Menlo,monospace;font-size:12.5px;background:var(--bg-secondary);padding:1px 5px;border-radius:var(--radius-sm)}
/* ---- card shell mirrors SchoolDetailView .card ---- */
.card{background:var(--bg-card);border:1px solid var(--border);border-radius:var(--radius-lg);
box-shadow:var(--shadow);padding:26px 26px 24px}
.cardhead{display:flex;align-items:center;justify-content:space-between;gap:16px;flex-wrap:wrap;margin-bottom:8px}
.sectionTitle{font-family:var(--serif);font-weight:600;font-size:22px;letter-spacing:-.01em}
.sectionSub{color:var(--text-secondary);font-size:14px;margin-bottom:18px}
.seg{display:inline-flex;background:var(--bg-secondary);border-radius:999px;padding:3px;gap:2px;flex:none}
.seg button{appearance:none;border:none;background:none;cursor:pointer;font:inherit;font-size:13px;font-weight:600;
color:var(--text-muted);padding:6px 13px;border-radius:999px;white-space:nowrap;transition:background .15s,color .15s}
.seg button[aria-pressed="true"]{background:var(--bg-card);color:var(--text-primary);box-shadow:var(--shadow)}
.seg button:focus-visible{outline:2px solid var(--accent-teal);outline-offset:2px}
/* ---- tiles ---- */
.tiles{display:grid;grid-template-columns:repeat(auto-fit,minmax(146px,1fr));gap:10px;margin-top:4px}
.tile{background:var(--bg-secondary);border-radius:var(--radius-md);padding:14px 15px 13px}
.tile .num{font-family:var(--serif);font-size:27px;font-weight:600;line-height:1.15;letter-spacing:-.01em;display:block}
.tile .num .unit{font-family:var(--sans);font-size:15px;font-weight:600;margin-left:2px}
.tile .num .sub{display:block;font-family:var(--sans);font-size:12.5px;font-weight:400;color:var(--text-muted);margin-top:1px}
.tile .lbl{display:block;font-size:12.5px;color:var(--text-secondary);margin-top:5px;line-height:1.35}
.tile.accent{background:var(--coral-bg)}
.tile.accent .num{color:var(--accent-coral-dark)}
.tile.newtile{background:var(--coral-bg);box-shadow:inset 0 0 0 1.5px var(--accent-coral)}
.tile.newtile .num{color:var(--accent-coral-dark)}
.newflag{display:inline-block;font-size:10px;font-weight:700;letter-spacing:.08em;text-transform:uppercase;
color:var(--accent-coral-dark);background:rgba(224,114,86,.2);padding:1px 6px;border-radius:999px;margin-bottom:6px}
/* ---- verdict banner ---- */
.verdict{display:flex;gap:13px;align-items:flex-start;padding:15px 17px;border-radius:var(--radius-md);margin:2px 0 20px}
.verdict.hard{background:var(--coral-bg)}
.verdict .vic{flex:none;width:22px;height:22px;border-radius:50%;display:grid;place-items:center;margin-top:1px;
background:var(--accent-coral-dark);color:#fff;font-size:12px;font-weight:700}
.verdict .vhead{font-weight:700;font-size:16.5px;letter-spacing:-.01em;color:var(--accent-coral-dark);text-wrap:balance}
.verdict .vsub{color:var(--text-secondary);font-size:13.5px;margin-top:3px}
/* ---- chart ---- */
.chartwrap{overflow-x:auto}
.chart{display:block;width:100%;min-width:460px;height:auto}
.axtxt{font-family:var(--sans);font-size:11px;fill:var(--text-muted)}
.ptlbl{font-family:var(--sans);font-size:12px;font-weight:700;fill:var(--accent-coral-dark)}
.keyrow{display:flex;flex-wrap:wrap;gap:16px;margin-top:12px;font-size:12.5px;color:var(--text-secondary)}
.keyrow span{display:inline-flex;align-items:center;gap:7px}
.kdot{width:11px;height:11px;border-radius:50%;flex:none}
.kdot.line{background:var(--accent-coral)}
.kdot.open{background:#fff;box-shadow:inset 0 0 0 2px var(--accent-teal)}
.kdot.gap{background:repeating-linear-gradient(90deg,var(--text-muted) 0 2px,transparent 2px 4px);border-radius:0;height:2px}
/* ---- table ---- */
.tblwrap{overflow-x:auto;margin-top:4px}
table{width:100%;border-collapse:collapse;font-size:14px;min-width:420px}
thead th{text-align:right;font-size:11px;font-weight:600;letter-spacing:.04em;text-transform:uppercase;
color:var(--text-muted);padding:0 10px 9px;border-bottom:1px solid var(--border)}
thead th:first-child{text-align:left}
tbody td{text-align:right;padding:11px 10px;border-bottom:1px solid var(--border)}
tbody td:first-child{text-align:left;font-weight:600}
tbody tr:last-child td{border-bottom:none}
tbody tr.now{background:var(--bg-secondary)}
td.miss{color:var(--text-muted);font-weight:400}
.pill{display:inline-block;font-size:11px;font-weight:600;padding:2px 9px;border-radius:999px;white-space:nowrap}
.pill.over{background:var(--coral-bg);color:var(--accent-coral-dark)}
.pill.ok{background:var(--teal-bg);color:var(--accent-teal)}
.pill.na{background:var(--bg-secondary);color:var(--text-muted)}
/* ---- map ---- */
.mapfig{border-radius:var(--radius-md);overflow:hidden;border:1px solid var(--border);background:#eef2ec}
.mapsvg{display:block;width:100%;height:auto}
.maplegend{display:flex;flex-wrap:wrap;gap:14px 20px;margin-top:12px;font-size:12.5px;color:var(--text-secondary)}
.maplegend span{display:inline-flex;align-items:center;gap:7px}
.swatch{width:14px;height:14px;border-radius:50%;flex:none}
.swatch.now{background:rgba(224,114,86,.18);box-shadow:inset 0 0 0 2px var(--accent-coral)}
.swatch.past{background:transparent;box-shadow:inset 0 0 0 1.5px rgba(176,74,46,.4)}
.swatch.you{background:var(--accent-teal);border-radius:2px;transform:rotate(45deg);width:11px;height:11px}
/* ---- postcode check ---- */
.checkbox{margin-top:20px;border-top:1px solid var(--border);padding-top:18px}
.checkhead{font-weight:700;font-size:15px;margin-bottom:3px}
.checksub{font-size:13.5px;color:var(--text-muted);margin-bottom:12px}
.checkform{display:flex;gap:8px;flex-wrap:wrap}
.checkform input{font:inherit;font-size:15px;padding:10px 13px;border:1px solid var(--border);border-radius:var(--radius-md);
background:var(--bg-card);color:var(--text-primary);min-width:150px;flex:1 1 150px;text-transform:uppercase}
.checkform input:focus-visible{outline:2px solid var(--accent-teal);outline-offset:1px;border-color:var(--accent-teal)}
.checkform button{font:inherit;font-weight:600;font-size:15px;padding:10px 20px;border:none;border-radius:var(--radius-md);
background:var(--accent-coral-dark);color:#fff;cursor:pointer;transition:background .15s}
.checkform button:hover{background:#9c3f26}
.checkform button:focus-visible{outline:2px solid var(--accent-teal);outline-offset:2px}
.result{margin-top:14px;background:var(--teal-bg);border-radius:var(--radius-md);padding:15px 17px}
.result .rhead{font-weight:700;font-size:16px;color:var(--accent-teal);letter-spacing:-.01em;text-wrap:balance}
.result .rsub{font-size:13.5px;color:var(--text-secondary);margin-top:4px}
.yearstrip{display:flex;gap:4px;margin-top:12px;flex-wrap:wrap}
.yr{font-size:11px;font-weight:600;padding:3px 7px;border-radius:var(--radius-sm);white-space:nowrap}
.yr.in{background:rgba(45,125,125,.22);color:var(--accent-teal)}
.yr.out{background:var(--coral-bg);color:var(--accent-coral-dark)}
.yr.none{background:var(--bg-secondary);color:var(--text-muted)}
/* ---- notes / disclosure ---- */
.note{display:flex;gap:9px;margin-top:16px;font-size:13px;color:var(--text-muted);line-height:1.55}
.note .i{flex:none;width:17px;height:17px;border-radius:50%;background:var(--bg-secondary);color:var(--text-muted);
font-size:11px;font-weight:700;display:grid;place-items:center;margin-top:2px}
.disclosure{margin-top:16px;border-top:1px solid var(--border);padding-top:4px}
.disclosure>summary{list-style:none;cursor:pointer;display:flex;align-items:center;gap:8px;padding:10px 0;
font-weight:600;font-size:14px;color:var(--accent-teal)}
.disclosure>summary::-webkit-details-marker{display:none}
.disclosure>summary:focus-visible{outline:2px solid var(--accent-teal);outline-offset:2px;border-radius:var(--radius-sm)}
.disclosure .chev{transition:transform .2s ease}
.disclosure[open]>summary .chev{transform:rotate(90deg)}
.disclosure .body{padding:2px 0 10px;font-size:13.5px;color:var(--text-secondary);display:flex;flex-direction:column;gap:10px}
.disclosure .body b{color:var(--text-primary)}
/* ---- viewport stack (keeps card height stable across views) ---- */
.viewport{display:grid}
.viewport>.view{grid-area:1/1}
.viewport>.view[hidden]{display:block;visibility:hidden;pointer-events:none}
/* ---- edge cases ---- */
.cases{display:grid;grid-template-columns:repeat(auto-fit,minmax(250px,1fr));gap:14px}
.case{background:var(--bg-card);border:1px solid var(--border);border-radius:var(--radius-md);padding:16px 17px}
.case h3{font-size:12px;font-weight:700;letter-spacing:.05em;text-transform:uppercase;color:var(--text-muted);margin-bottom:10px}
.case .body{font-size:14px;color:var(--text-secondary);line-height:1.5}
.case .body strong{color:var(--text-primary)}
.emptybox{background:var(--bg-secondary);border-radius:var(--radius-md);padding:13px 15px;font-size:13.5px;color:var(--text-secondary)}
/* ---- phone ---- */
.phonerow{display:flex;gap:24px;flex-wrap:wrap;align-items:flex-start}
.phone{width:330px;max-width:100%;border:9px solid #1a1612;border-radius:34px;overflow:hidden;box-shadow:var(--shadow-md);background:var(--bg-primary)}
.phonebody{padding:14px 13px 20px;display:flex;flex-direction:column;gap:12px}
.phone .card{padding:17px 16px 16px;border-radius:var(--radius-md)}
.phone .sectionTitle{font-size:18px}
.phone .tiles{grid-template-columns:1fr 1fr;gap:8px}
.phone .tile{padding:11px 12px}
.phone .tile .num{font-size:22px}
.phone .chart{min-width:0}
.phone .chartwrap{overflow:visible}
.phonenote{font-size:13px;color:var(--text-muted);flex:1 1 240px;min-width:220px}
.phonenote h3{font-family:var(--serif);font-size:17px;color:var(--text-primary);margin-bottom:8px;font-weight:600}
.phonenote ul{padding-left:18px;display:flex;flex-direction:column;gap:7px}
@media (prefers-reduced-motion:reduce){*{transition:none!important;animation:none!important}}
</style>
<div class="wrap">
<div class="pagehead">
<h1>Last distance offered — school detail page</h1>
<p>Adds the final-offer cut-off distance to the existing <b>Admissions</b> card, plus a catchment
view on the map. Data covers one to ten years depending on the school and local authority, so every
screen here is built around partial coverage rather than assuming a full run.</p>
<p>Sample school: <b>Fairlawn Primary School</b>, Lewisham — 8 years of distance data out of 10 years of admissions data.</p>
</div>
<!-- ============ 1. ADMISSIONS CARD ============ -->
<div class="step">
<h2>1. A distance tile joins the admissions tiles</h2>
<span class="where">SchoolDetailView · #admissions</span>
</div>
<p class="stepdesc">No new section and no new nav entry — the number a parent actually asks for
("how close do we need to live?") sits with the rest of the intake story. The segmented control gains a
third view, <code>Distance</code>.</p>
<div class="card">
<div class="cardhead">
<h2 class="sectionTitle">Admissions</h2>
<div class="seg" role="group" aria-label="Admissions view">
<button type="button" aria-pressed="true" data-view="year">This year</button>
<button type="button" aria-pressed="false" data-view="trend">10-year trend</button>
<button type="button" aria-pressed="false" data-view="dist">Distance</button>
</div>
</div>
<p class="sectionSub">Reception entry, September 2025.</p>
<div class="viewport">
<!-- view: this year -->
<div class="view" id="v-year">
<dl class="tiles">
<div class="tile">
<dd class="num">60</dd><dt class="lbl">Places offered</dt>
</div>
<div class="tile">
<dd class="num">142</dd><dt class="lbl">Wanted it first</dt>
</div>
<div class="tile accent">
<dd class="num">54<span class="sub">of 142 · 38%</span></dd>
<dt class="lbl">Got their first choice</dt>
</div>
<div class="tile newtile">
<span class="newflag">New</span>
<dd class="num">0.31<span class="unit">mi</span><span class="sub">≈ 500 m · 6 min walk</span></dd>
<dt class="lbl">Last distance offered</dt>
</div>
</dl>
<div class="note">
<span class="i" aria-hidden="true">i</span>
<span>The furthest home offered a place once siblings, faith and EHCP priority were applied.
It is not a fixed catchment — it moves every year with the number of applications.</span>
</div>
</div>
<!-- view: distance -->
<div class="view" id="v-dist" hidden>
<div class="verdict hard">
<span class="vic" aria-hidden="true">↓</span>
<div>
<div class="vhead">The catchment has halved in nine years</div>
<div class="vsub">0.62 mi in 2016 → 0.31 mi in 2025. Four of the last five years tightened.</div>
</div>
</div>
<div class="chartwrap">
<svg class="chart" viewBox="0 0 700 250" role="img"
aria-label="Last distance offered by year: 0.62 miles in 2016, 0.55 in 2017, not published in 2018, 0.48 in 2019, 0.51 in 2020, 0.44 in 2021, no cut-off needed in 2022, 0.39 in 2023, 0.35 in 2024, 0.31 in 2025.">
<!-- grid -->
<g stroke="#e5dfd5" stroke-width="1">
<line x1="46" y1="30" x2="686" y2="30"/>
<line x1="46" y1="80" x2="686" y2="80"/>
<line x1="46" y1="130" x2="686" y2="130"/>
<line x1="46" y1="180" x2="686" y2="180"/>
</g>
<line x1="46" y1="206" x2="686" y2="206" stroke="#d8d0c3" stroke-width="1.5"/>
<g class="axtxt" text-anchor="end">
<text x="38" y="34">0.8</text><text x="38" y="84">0.6</text>
<text x="38" y="134">0.4</text><text x="38" y="184">0.2</text>
</g>
<text class="axtxt" x="46" y="16" text-anchor="start">miles</text>
<!-- gap segments (dashed = no figure published) -->
<g fill="none" stroke="#6d685f" stroke-width="1.5" stroke-dasharray="4 4" opacity=".55">
<path d="M114 92 L182 110"/>
<path d="M318 122 L386 132.5"/>
</g>
<!-- solid series -->
<polyline fill="none" stroke="#e07256" stroke-width="2.5" stroke-linejoin="round" stroke-linecap="round"
points="46,75 114,92"/>
<polyline fill="none" stroke="#e07256" stroke-width="2.5" stroke-linejoin="round" stroke-linecap="round"
points="182,110 250,102 318,122"/>
<polyline fill="none" stroke="#e07256" stroke-width="2.5" stroke-linejoin="round" stroke-linecap="round"
points="386,132.5 454,142.5 522,152.5"/>
<!-- points -->
<g fill="#e07256" stroke="#fff" stroke-width="2">
<circle cx="46" cy="75" r="5"/><circle cx="114" cy="92" r="5"/>
<circle cx="182" cy="110" r="5"/><circle cx="250" cy="102" r="5"/>
<circle cx="318" cy="122" r="5"/><circle cx="386" cy="132.5" r="5"/>
<circle cx="454" cy="142.5" r="5"/>
</g>
<!-- latest, emphasised -->
<circle cx="522" cy="152.5" r="7.5" fill="#b04a2e" stroke="#fff" stroke-width="2.5"/>
<!-- "no cut-off needed" marker, 2022 -->
<circle cx="352" cy="46" r="6" fill="#fff" stroke="#296f6f" stroke-width="2.5"/>
<text class="axtxt" x="352" y="34" text-anchor="middle" fill="#296f6f" font-weight="600">all offered</text>
<line x1="352" y1="54" x2="352" y2="196" stroke="#296f6f" stroke-width="1.5" stroke-dasharray="3 4" opacity=".45"/>
<!-- endpoint labels -->
<text class="ptlbl" x="46" y="63" text-anchor="start">0.62</text>
<text class="ptlbl" x="530" y="157" text-anchor="start">0.31 mi</text>
<!-- year axis -->
<g class="axtxt" text-anchor="middle">
<text x="46" y="226">2016</text><text x="114" y="226">2017</text>
<text x="182" y="226">2019</text><text x="250" y="226">2020</text>
<text x="318" y="226">2021</text><text x="386" y="226">2023</text>
<text x="454" y="226">2024</text><text x="522" y="226">2025</text>
</g>
<text class="axtxt" x="148" y="243" text-anchor="middle" fill="#6d685f">2018 not published</text>
<text class="axtxt" x="352" y="243" text-anchor="middle" fill="#296f6f">2022 undersubscribed</text>
</svg>
</div>
<div class="keyrow">
<span><i class="kdot line" aria-hidden="true"></i>Distance of the last place offered</span>
<span><i class="kdot open" aria-hidden="true"></i>No cut-off needed — every applicant offered</span>
<span><i class="kdot gap" aria-hidden="true"></i>Not published by the local authority</span>
</div>
<div class="tblwrap" style="margin-top:22px">
<table>
<caption class="sr-only" style="position:absolute;width:1px;height:1px;overflow:hidden;clip:rect(0 0 0 0)">Last distance offered by year</caption>
<thead>
<tr><th scope="col">Year</th><th scope="col">Last distance</th><th scope="col">Places</th><th scope="col">Status</th></tr>
</thead>
<tbody>
<tr class="now"><td>2025</td><td>0.31 mi</td><td>60</td><td><span class="pill over">Oversubscribed</span></td></tr>
<tr><td>2024</td><td>0.35 mi</td><td>60</td><td><span class="pill over">Oversubscribed</span></td></tr>
<tr><td>2023</td><td>0.39 mi</td><td>60</td><td><span class="pill over">Oversubscribed</span></td></tr>
<tr><td>2022</td><td class="miss">No cut-off needed</td><td>60</td><td><span class="pill ok">All offered</span></td></tr>
<tr><td>2021</td><td>0.44 mi</td><td>60</td><td><span class="pill over">Oversubscribed</span></td></tr>
<tr><td>2020</td><td>0.51 mi</td><td>60</td><td><span class="pill over">Oversubscribed</span></td></tr>
<tr><td>2019</td><td>0.48 mi</td><td>60</td><td><span class="pill over">Oversubscribed</span></td></tr>
<tr><td>2018</td><td class="miss">—</td><td>60</td><td><span class="pill na">Not published</span></td></tr>
<tr><td>2017</td><td>0.55 mi</td><td>60</td><td><span class="pill over">Oversubscribed</span></td></tr>
<tr><td>2016</td><td>0.62 mi</td><td>60</td><td><span class="pill over">Oversubscribed</span></td></tr>
</tbody>
</table>
</div>
<details class="disclosure">
<summary><span class="chev" aria-hidden="true">›</span>What this number does and doesn't tell you</summary>
<div class="body">
<span><b>It is a result, not a rule.</b> It records how far the last successful applicant lived
in a given year. Move one large sibling cohort and the figure shifts.</span>
<span><b>Distance is the final tiebreak.</b> Children in care, EHCP places, siblings and — at faith
schools — the faith criteria are ranked first. A family inside the distance can still miss out.</span>
<span><b>Lewisham measures straight-line distance</b> from home to the school's main gate. Other
authorities use walking routes, which are always longer for the same home.</span>
<span><b>Gaps are normal.</b> An authority may not publish a figure, or the school may not have
needed a distance cut-off that year. Both are shown here rather than hidden.</span>
</div>
</details>
</div>
</div>
</div>
<!-- ============ 2. MAP ============ -->
<div class="step">
<h2>2. The cut-off drawn on the map</h2>
<span class="where">SchoolHeroMap · catchment layer</span>
</div>
<p class="stepdesc">"0.31 miles" is abstract until you see it over your own streets. The latest year is a
filled ring; earlier years sit behind it as hairlines, so the tightening reads instantly as a set of
shrinking circles. Straight-line rings only — they're an illustration of the number, not a boundary.</p>
<div class="card">
<div class="cardhead" style="margin-bottom:14px">
<h2 class="sectionTitle">Where the last place went</h2>
<div class="seg" role="group" aria-label="Map years">
<button type="button" aria-pressed="true">2025 only</button>
<button type="button" aria-pressed="false">Last 5 years</button>
</div>
</div>
<figure class="mapfig">
<svg class="mapsvg" viewBox="0 0 700 400" role="img"
aria-label="Map showing concentric catchment rings around Fairlawn Primary School: 0.62 miles in 2016 shrinking to 0.31 miles in 2025, with a marker for a sample home 0.24 miles away, inside the current ring.">
<rect width="700" height="400" fill="#eef2ec"/>
<!-- park -->
<path d="M470 20 h230 v150 h-160 q-70 -20 -70 -80 z" fill="#dfe9dc"/>
<!-- water -->
<path d="M0 330 q120 -40 250 -10 t260 -20 l190 -30 v130 H0 z" fill="#dbe6ee"/>
<!-- street grid -->
<g stroke="#fff" stroke-width="7" stroke-linecap="round" opacity=".95">
<path d="M-10 90 H710"/><path d="M-10 200 H710"/><path d="M-10 300 H710"/>
<path d="M120 -10 V410"/><path d="M300 -10 V410"/><path d="M470 -10 V410"/><path d="M620 -10 V410"/>
</g>
<g stroke="#fff" stroke-width="3.5" opacity=".8">
<path d="M-10 145 H710"/><path d="M-10 250 H710"/><path d="M210 -10 V410"/><path d="M385 -10 V410"/><path d="M550 -10 V410"/>
</g>
<!-- building blocks -->
<g fill="#e4e2dc" opacity=".85">
<rect x="132" y="102" width="60" height="30" rx="2"/><rect x="222" y="102" width="64" height="30" rx="2"/>
<rect x="132" y="212" width="60" height="26" rx="2"/><rect x="222" y="212" width="64" height="26" rx="2"/>
<rect x="400" y="102" width="56" height="30" rx="2"/><rect x="400" y="212" width="56" height="26" rx="2"/>
</g>
<!-- historic rings (hairline) -->
<g fill="none" stroke="#b04a2e" opacity=".38" stroke-width="1.5">
<circle cx="300" cy="200" r="150"/>
<circle cx="300" cy="200" r="126"/>
<circle cx="300" cy="200" r="106"/>
<circle cx="300" cy="200" r="90"/>
</g>
<text class="axtxt" x="300" y="44" text-anchor="middle" fill="#b04a2e" font-weight="600">2016 · 0.62 mi</text>
<!-- current ring -->
<circle cx="300" cy="200" r="75" fill="rgba(224,114,86,.18)" stroke="#e07256" stroke-width="3"/>
<text class="axtxt" x="300" y="118" text-anchor="middle" fill="#b04a2e" font-weight="700" font-size="12.5">2025 · 0.31 mi</text>
<!-- school pin -->
<circle cx="300" cy="200" r="11" fill="#b04a2e" stroke="#fff" stroke-width="3"/>
<text class="axtxt" x="300" y="232" text-anchor="middle" fill="#1a1612" font-weight="700" font-size="12">Fairlawn Primary</text>
<!-- your home -->
<g transform="translate(246,158)">
<rect x="-7" y="-7" width="14" height="14" rx="2" fill="#296f6f" stroke="#fff" stroke-width="2.5" transform="rotate(45)"/>
</g>
<text class="axtxt" x="246" y="140" text-anchor="middle" fill="#296f6f" font-weight="700" font-size="12">Your home · 0.24 mi</text>
</svg>
</figure>
<div class="maplegend">
<span><i class="swatch now" aria-hidden="true"></i>2025 cut-off — 0.31 mi</span>
<span><i class="swatch past" aria-hidden="true"></i>Earlier years, 2016–2024</span>
<span><i class="swatch you" aria-hidden="true"></i>Your home</span>
</div>
<div class="checkbox">
<div class="checkhead">Would you have got in?</div>
<p class="checksub">We measure straight-line distance from your postcode, the same way Lewisham does.</p>
<form class="checkform" onsubmit="return false">
<label class="sr-only" for="pc" style="position:absolute;width:1px;height:1px;overflow:hidden;clip:rect(0 0 0 0)">Your postcode</label>
<input id="pc" type="text" value="SE23 3NA" autocomplete="postal-code" spellcheck="false">
<button type="submit">Check</button>
</form>
<div class="result">
<div class="rhead">0.24 miles away — inside the cut-off in all 8 years on record</div>
<div class="rsub">That's 0.07 miles of headroom on 2025, the tightest year so far. 2018 has no published
figure, and in 2022 every applicant was offered a place.</div>
<div class="yearstrip">
<span class="yr in">2016 ✓</span>
<span class="yr in">2017 ✓</span>
<span class="yr none">2018 –</span>
<span class="yr in">2019 ✓</span>
<span class="yr in">2020 ✓</span>
<span class="yr in">2021 ✓</span>
<span class="yr in">2022 ✓</span>
<span class="yr in">2023 ✓</span>
<span class="yr in">2024 ✓</span>
<span class="yr in">2025 ✓</span>
</div>
</div>
<div class="note">
<span class="i" aria-hidden="true">!</span>
<span>An indication only. Distance is applied after siblings, faith and EHCP priority, and next year's
cut-off depends on next year's applicants. Always check the school's own admissions policy.</span>
</div>
</div>
</div>
<!-- ============ 3. COVERAGE STATES ============ -->
<div class="step">
<h2>3. Coverage states</h2>
<span class="where">Partial data is the normal case</span>
</div>
<p class="stepdesc">Coverage runs from ten years to none. Each state says something true rather than
falling back on a generic "no data" — the reason a figure is absent is itself useful to a parent.</p>
<div class="cases">
<div class="case">
<h3>4+ years</h3>
<div class="body">Full treatment: verdict banner, chart, table, map rings. <strong>The verdict line only
appears with 4+ points</strong> — below that a two-year swing isn't a trend.</div>
</div>
<div class="case">
<h3>2–3 years</h3>
<div class="body">Tile and table, no verdict banner, single map ring for the latest year.
<strong>"Only 3 years available"</strong> sits under the table.</div>
</div>
<div class="case">
<h3>1 year</h3>
<div class="body">Tile plus one map ring. No <code>Distance</code> tab — the tile and its footnote are
the whole story.</div>
</div>
<div class="case">
<h3>Never oversubscribed</h3>
<div class="body" style="margin-bottom:10px">Not missing data — good news, and it should read that way.</div>
<div class="emptybox"><b style="color:var(--accent-teal)">Everyone who applied was offered a place</b>
in each of the last 6 years, so no distance cut-off was needed.</div>
</div>
<div class="case">
<h3>No data at all</h3>
<div class="body" style="margin-bottom:10px">Replaces the current placeholder text in
<code>SecondarySchoolDetailView.tsx:780</code>.</div>
<div class="emptybox">Lewisham hasn't published cut-off distances for this school.
<a href="#" style="color:var(--accent-teal);font-weight:600">The council's admissions page</a> may list them.</div>
</div>
<div class="case">
<h3>Not distance-ranked</h3>
<div class="body" style="margin-bottom:10px">Grammar and some faith schools rank on test score or faith
practice, so a distance figure would mislead.</div>
<div class="emptybox">Places here are ranked by the entrance test, not by distance.</div>
</div>
</div>
<!-- ============ 4. MOBILE ============ -->
<div class="step">
<h2>4. Mobile</h2>
<span class="where">≤ 640 px</span>
</div>
<p class="stepdesc">Tiles fall to two columns and the distance tile takes the full width beneath them, so the
headline number survives the reflow. The chart drops its table on mobile behind a "See all years" disclosure.</p>
<div class="phonerow">
<div class="phone">
<div class="phonebody">
<div class="card">
<div class="cardhead" style="margin-bottom:6px">
<h2 class="sectionTitle">Admissions</h2>
</div>
<p class="sectionSub" style="font-size:13px;margin-bottom:12px">Reception, Sept 2025</p>
<div class="seg" style="margin-bottom:14px">
<button type="button" aria-pressed="true">Year</button>
<button type="button" aria-pressed="false">Trend</button>
<button type="button" aria-pressed="false">Distance</button>
</div>
<dl class="tiles">
<div class="tile"><dd class="num">60</dd><dt class="lbl">Places offered</dt></div>
<div class="tile"><dd class="num">142</dd><dt class="lbl">Wanted it first</dt></div>
</dl>
<div class="tile newtile" style="margin-top:8px">
<span class="newflag">New</span>
<dd class="num" style="font-size:26px">0.31<span class="unit">mi</span>
<span class="sub">≈ 500 m · 6 min walk</span></dd>
<dt class="lbl">Last distance offered, 2025</dt>
</div>
</div>
<div class="card">
<h2 class="sectionTitle" style="margin-bottom:12px">Distance</h2>
<div class="verdict hard" style="padding:12px 13px;margin:0 0 14px">
<span class="vic" aria-hidden="true">↓</span>
<div><div class="vhead" style="font-size:15px">Halved in nine years</div>
<div class="vsub" style="font-size:12.5px">0.62 mi → 0.31 mi</div></div>
</div>
<div class="chartwrap">
<svg class="chart" viewBox="0 0 300 150" role="img" aria-label="Last distance offered falling from 0.62 miles in 2016 to 0.31 miles in 2025.">
<g stroke="#e5dfd5" stroke-width="1">
<line x1="26" y1="20" x2="292" y2="20"/><line x1="26" y1="62" x2="292" y2="62"/><line x1="26" y1="104" x2="292" y2="104"/>
</g>
<g class="axtxt" text-anchor="end" font-size="9">
<text x="21" y="23">0.8</text><text x="21" y="65">0.5</text><text x="21" y="107">0.2</text>
</g>
<path d="M26 55 L64 69" fill="none" stroke="#e07256" stroke-width="2.5" stroke-linecap="round"/>
<path d="M64 69 L102 84" fill="none" stroke="#6d685f" stroke-width="1.5" stroke-dasharray="3 3" opacity=".55"/>
<path d="M102 84 L140 78 L178 95" fill="none" stroke="#e07256" stroke-width="2.5" stroke-linejoin="round" stroke-linecap="round"/>
<path d="M178 95 L216 109" fill="none" stroke="#6d685f" stroke-width="1.5" stroke-dasharray="3 3" opacity=".55"/>
<path d="M216 109 L254 116 L280 120" fill="none" stroke="#e07256" stroke-width="2.5" stroke-linejoin="round" stroke-linecap="round"/>
<g fill="#e07256" stroke="#fff" stroke-width="1.8">
<circle cx="26" cy="55" r="4"/><circle cx="64" cy="69" r="4"/><circle cx="102" cy="84" r="4"/>
<circle cx="140" cy="78" r="4"/><circle cx="178" cy="95" r="4"/><circle cx="216" cy="109" r="4"/><circle cx="254" cy="116" r="4"/>
</g>
<circle cx="280" cy="120" r="6" fill="#b04a2e" stroke="#fff" stroke-width="2"/>
<circle cx="197" cy="30" r="4.5" fill="#fff" stroke="#296f6f" stroke-width="2"/>
<g class="axtxt" font-size="9" text-anchor="middle">
<text x="26" y="132">'16</text><text x="140" y="132">'20</text><text x="280" y="132">'25</text>
</g>
<text class="ptlbl" x="280" y="140" text-anchor="middle" font-size="11">0.31 mi</text>
</svg>
</div>
<details class="disclosure" style="margin-top:10px">
<summary style="font-size:13px"><span class="chev" aria-hidden="true">›</span>See all 10 years</summary>
</details>
</div>
</div>
</div>
<div class="phonenote">
<h3>Mobile decisions</h3>
<ul>
<li>Distance tile spans both columns — the number a parent came for shouldn't be a half-width cell.</li>
<li>Chart keeps every year but labels only first, middle and last; the endpoint stays labelled.</li>
<li>Year-by-year table collapses into a disclosure rather than forcing a horizontal scroll.</li>
<li>Map rings reuse the existing full-screen hero map sheet, opened from the tile.</li>
</ul>
</div>
</div>
</div>
<script>
// Segmented control on the admissions card swaps the stacked views.
document.querySelectorAll('.seg button[data-view]').forEach(function (btn) {
btn.addEventListener('click', function () {
var group = btn.closest('.seg');
group.querySelectorAll('button').forEach(function (b) { b.setAttribute('aria-pressed', String(b === btn)); });
var target = btn.dataset.view === 'dist' ? 'v-dist' : 'v-year';
['v-year', 'v-dist'].forEach(function (id) {
document.getElementById(id).hidden = id !== target;
});
});
});
</script>
+373
View File
@@ -0,0 +1,373 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Similar Schools Nearby</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600&family=Manrope:wght@500;600;700&display=swap" rel="stylesheet">
<style>
/* Tokens copied verbatim from nextjs-app/app/(frontend)/globals.css so this
mockup cannot drift from the shipped palette. Light values first, dark
under prefers-color-scheme, both overridable by the theme switch. */
:root {
color-scheme: light;
--bg-primary:#FAFAF8; --bg-secondary:#F5EFE6; --bg-card:#FFFFFF;
--text-primary:#1C2731; --text-secondary:#4A5560; --text-muted:#5F6A75;
--border:#E5E7EB; --border-strong:#D3D7DD;
--brand:#0F766E; --brand-strong:#0C5F58; --brand-bg:rgba(15,118,110,.10); --brand-on:#FFFFFF;
--action:#BE3C27; --action-strong:#A33320; --action-on:#FFFFFF;
--sand:#F5EFE6;
--font-display:Manrope,-apple-system,BlinkMacSystemFont,sans-serif;
--font-ui:Inter,-apple-system,BlinkMacSystemFont,sans-serif;
--radius-md:8px; --radius-lg:16px;
--shadow:0 1px 2px rgba(28,39,49,.06),0 1px 3px rgba(28,39,49,.05);
}
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) {
color-scheme: dark;
--bg-primary:#111A20; --bg-secondary:#16222A; --bg-card:#18242C;
--text-primary:#E9EEF0; --text-secondary:#B4C2C7; --text-muted:#8B9AA1;
--border:#26343D; --border-strong:#35454F;
--brand:#5FC7BB; --brand-strong:#7BD6CC; --brand-bg:rgba(95,199,187,.14); --brand-on:#0A1418;
--action:#F08A72; --action-strong:#F5A492; --action-on:#241009;
--sand:#1B2730;
--shadow:0 1px 2px rgba(0,0,0,.3),0 1px 3px rgba(0,0,0,.25);
}
}
:root[data-theme="dark"] {
color-scheme: dark;
--bg-primary:#111A20; --bg-secondary:#16222A; --bg-card:#18242C;
--text-primary:#E9EEF0; --text-secondary:#B4C2C7; --text-muted:#8B9AA1;
--border:#26343D; --border-strong:#35454F;
--brand:#5FC7BB; --brand-strong:#7BD6CC; --brand-bg:rgba(95,199,187,.14); --brand-on:#0A1418;
--action:#F08A72; --action-strong:#F5A492; --action-on:#241009;
--sand:#1B2730;
--shadow:0 1px 2px rgba(0,0,0,.3),0 1px 3px rgba(0,0,0,.25);
}
* { box-sizing:border-box; }
body {
margin:0; padding:32px 16px 80px; background:var(--bg-primary);
color:var(--text-primary); font:15px/1.55 var(--font-ui);
-webkit-font-smoothing:antialiased;
}
.page { max-width:960px; margin:0 auto; }
.page > header { margin-bottom:28px; display:flex; flex-wrap:wrap; gap:16px; align-items:flex-start; justify-content:space-between; }
.page > header h1 { font:700 25px/1.25 var(--font-display); letter-spacing:-.6px; margin:0 0 6px; }
.page > header p { margin:0; color:var(--text-muted); font-size:14px; max-width:60ch; }
.theme-switch { border:1px solid var(--border-strong); background:var(--bg-card); color:var(--text-secondary); border-radius:999px; padding:8px 14px; font:500 13px var(--font-ui); cursor:pointer; min-height:44px; }
.theme-switch:hover { border-color:var(--brand); color:var(--brand); }
/* ── The page context each variant is shown inside ─────────────────── */
.variant { margin-bottom:40px; }
.variant > .context { padding:0 4px 12px; }
.variant .eyebrow { margin:0 0 4px; font-size:12px; letter-spacing:.04em; text-transform:uppercase; color:var(--text-muted); }
.variant .context h2 { font:600 18px/1.35 var(--font-display); margin:0; color:var(--text-secondary); }
.variant .note { margin:10px 4px 0; font-size:12.5px; color:var(--text-muted); }
.variant .note b { color:var(--text-secondary); font-weight:600; }
/* ── The section itself — mirrors components/school/Section ────────── */
.card {
background:var(--bg-card); border:1px solid var(--border);
border-radius:var(--radius-lg); padding:28px; box-shadow:var(--shadow);
}
.top { display:flex; align-items:flex-start; justify-content:space-between; gap:16px; }
.top h2 { font:700 22px/1.25 var(--font-display); letter-spacing:-.4px; margin:0; }
.lede { margin:8px 0 20px; color:var(--text-secondary); font-size:14.5px; max-width:64ch; }
/* ── Carousel ───────────────────────────────────────────────────────
Every card is in the DOM and in the initial HTML — the arrows scroll a
list, they do not swap a view. That keeps all six links crawlable and
keeps the section usable with no JavaScript, where it degrades to a
plain horizontally scrollable row. */
.arrows { display:flex; gap:8px; flex:none; }
.arrow {
width:44px; height:44px; display:grid; place-items:center; cursor:pointer;
border:1px solid var(--border-strong); border-radius:999px;
background:var(--bg-card); color:var(--brand);
}
.arrow:hover:not(:disabled) { border-color:var(--brand); background:var(--brand-bg); }
.arrow:disabled { opacity:.35; cursor:default; }
.arrow:focus-visible { outline:2px solid var(--brand); outline-offset:2px; }
.arrow svg { width:17px; height:17px; }
.scroller {
display:grid; grid-auto-flow:column;
grid-auto-columns:calc((100% - 28px) / 3);
gap:14px; overflow-x:auto; scroll-snap-type:x mandatory;
padding:2px; margin:-2px; /* room for focus rings */
scrollbar-width:none; -ms-overflow-style:none;
list-style:none;
}
.scroller::-webkit-scrollbar { display:none; }
.scroller:focus-visible { outline:2px solid var(--brand); outline-offset:4px; border-radius:var(--radius-md); }
@media (max-width:820px) { .scroller { grid-auto-columns:calc((100% - 14px) / 2); } }
/* Touch widths: the arrows would squeeze the lede into a four-line column for a
control that swiping already provides, so they go and the documented
right-edge fade carries the affordance instead (MOBILE.md). The fade lifts at
the end of the travel, where there is nothing more to hint at. */
@media (max-width:640px) {
.top { display:block; }
.arrows { display:none; }
.scroller { grid-auto-columns:86%; mask-image:linear-gradient(to right, #000 calc(100% - 28px), transparent); }
.scroller[data-at-end=true] { mask-image:none; }
.card { padding:20px; }
}
.school {
position:relative; display:flex; flex-direction:column; scroll-snap-align:start;
border:1px solid var(--border); border-radius:var(--radius-md);
padding:16px; background:var(--bg-card);
}
.school:has(.add[aria-pressed=true]) { border-color:var(--brand); background:var(--brand-bg); }
.distance { display:flex; align-items:center; gap:5px; font-size:12px; color:var(--text-muted); margin:0 0 10px; }
.distance svg { width:13px; height:13px; flex:none; }
.school h3 { font:600 16px/1.35 var(--font-display); margin:0 0 6px; }
/* The whole card is the link target; the button sits above it on z-index so
it stays independently clickable. */
.school h3 a { color:var(--text-primary); text-decoration:none; }
.school h3 a::after { content:""; position:absolute; inset:0; border-radius:var(--radius-md); }
.school:hover { border-color:var(--border-strong); }
.school h3 a:hover { color:var(--brand); text-decoration:underline; }
.school h3 a:focus-visible { outline:none; }
.school:has(h3 a:focus-visible) { outline:2px solid var(--brand); outline-offset:2px; }
.meta { margin:0 0 12px; font-size:12.5px; color:var(--text-muted); }
.shared { display:flex; flex-wrap:wrap; gap:6px; margin:0 0 14px; padding:0; list-style:none; }
.shared li { font-size:11.5px; line-height:1.4; padding:4px 8px; border-radius:999px; background:var(--brand-bg); color:var(--brand); border:1px solid transparent; }
.shared li.loose { background:transparent; color:var(--text-muted); border-color:var(--border); }
.metric { margin-top:auto; padding-top:13px; border-top:1px solid var(--border); }
.value { font:700 26px/1.1 var(--font-display); letter-spacing:-.6px; margin:0; }
.value.absent { font-size:15px; font-weight:600; color:var(--text-muted); letter-spacing:0; }
.metric .label { margin:4px 0 0; font-size:12px; color:var(--text-secondary); }
.metric .ref { margin:2px 0 0; font-size:12px; color:var(--text-muted); }
.add {
position:relative; z-index:1; margin-top:14px; width:100%; min-height:44px;
font:500 13px var(--font-ui); cursor:pointer; border-radius:var(--radius-md);
border:1px solid var(--border-strong); background:var(--bg-card); color:var(--brand);
}
.add:hover { border-color:var(--brand); background:var(--brand-bg); }
.add[aria-pressed=true] { border-color:var(--brand); background:var(--brand-bg); font-weight:600; }
.add:focus-visible { outline:2px solid var(--brand); outline-offset:2px; }
.footer {
display:flex; flex-wrap:wrap; align-items:center; justify-content:space-between;
gap:14px; margin-top:20px; padding-top:18px; border-top:1px solid var(--border);
}
.footer p { margin:0; font-size:12.5px; color:var(--text-muted); }
.footer strong { display:block; font:600 14px var(--font-ui); color:var(--text-primary); }
/* Coral: the one decisive action in this section, and there is only one. */
.compare {
min-height:44px; padding:0 20px; border-radius:var(--radius-md); cursor:pointer;
font:600 14px var(--font-ui); background:var(--action); color:var(--action-on);
border:1px solid var(--action); text-decoration:none; display:inline-flex; align-items:center; gap:8px;
}
.compare:hover { background:var(--action-strong); border-color:var(--action-strong); }
.compare[aria-disabled=true] { opacity:.45; pointer-events:none; }
.caption { margin:16px 0 0; font-size:11.5px; color:var(--text-muted); }
.sr { position:absolute; width:1px; height:1px; padding:0; margin:-1px; overflow:hidden; clip:rect(0 0 0 0); white-space:nowrap; border:0; }
</style>
</head>
<body>
<div class="page">
<header>
<div>
<h1>Similar schools nearby</h1>
<p>A new section on the school detail page. Fictional schools and figures; shipped
colour, type and section shell taken from <code>globals.css</code>.</p>
</div>
<button class="theme-switch" type="button" id="theme">Dark theme</button>
</header>
<div id="variants"></div>
</div>
<script>
const PIN = '<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M20 10c0 6-8 12-8 12s-8-6-8-12a8 8 0 0 1 16 0Z"/><circle cx="12" cy="10" r="3"/></svg>';
const CHEV = (dir) => `<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="${dir === 'prev' ? 'M15 18l-6-6 6-6' : 'M9 18l6-6-6-6'}"/></svg>`;
const variants = [
{
id: 'dense',
eyebrow: 'Variant 1 · Dense urban primary — six matches, carousel active',
context: 'Meadowbrook Primary School — Ages 4–11 · Mixed · No religious character · Community school',
lede: 'Other primary schools near Meadowbrook Primary School, with a similar intake.',
metric: 'Reading, writing & maths',
caption: 'Meeting the expected standard at key stage 2, 2025.',
thisValue: '72%',
note: 'Fourteen schools cleared <b>tier 1</b> within three miles, so the section takes the six nearest and stops there. The arrows scroll a list that is entirely in the HTML — all six links are crawlable, and with JavaScript off the row still scrolls.',
schools: [
{ name:'Willow Lane Primary School', distance:'0.4', meta:'Community school · Ages 4–11', shared:['Mixed','No religious character'], value:'74%' },
{ name:'Oakfield Primary School', distance:'0.6', meta:'Academy converter · Ages 3–11', shared:['Mixed','No religious character'], value:'69%' },
{ name:'Brookside Primary School', distance:'0.9', meta:'Community school · Ages 4–11', shared:['Mixed','No religious character'], value:'Not published' },
{ name:'Hollytree Primary School', distance:'1.3', meta:'Academy converter · Ages 4–11', shared:['Mixed','No religious character'], value:'81%' },
{ name:'Marsh Green Primary School', distance:'1.8', meta:'Community school · Ages 3–11', shared:['Mixed','No religious character'], value:'64%' },
{ name:'Kingsway Primary School', distance:'2.2', meta:'Foundation school · Ages 4–11', shared:['Mixed','No religious character'], value:'77%' },
],
},
{
id: 'secondary',
eyebrow: 'Variant 2 · Secondary — four matches, mixed tiers',
context: 'Meadowbrook High School — Ages 11–18 · Mixed · Non-selective · Academy',
lede: 'Other secondary schools near Meadowbrook High School, with a similar intake.',
metric: 'Attainment 8',
caption: 'Average GCSE attainment score across eight qualifications, 2025.',
thisValue: '51.2',
note: 'Only two schools cleared tier 1, so the search widened to <b>tier 2</b> and found two more. It stops there rather than widening again to reach six — tiers relax to reach a usable set, never to fill the last slots. Selectivity never relaxes, so no grammar school can appear here.',
schools: [
{ name:'Rivermead High School', distance:'0.9', meta:'Academy converter · Ages 11–18', shared:['Mixed','Non-selective','No religious character'], value:'52.8' },
{ name:'Oakfield Academy', distance:'1.7', meta:'Academy sponsor led · Ages 11–16', shared:['Mixed','Non-selective','No religious character'], value:'49.6' },
{ name:'St Aidan’s Catholic High School', distance:'2.4', meta:'Voluntary aided · Ages 11–18', shared:['Mixed','Non-selective'], value:'53.4' },
{ name:'Parkside Community School', distance:'3.8', meta:'Community school · Ages 11–16', shared:['Mixed','Non-selective'], value:'50.9' },
],
},
{
id: 'sparse',
eyebrow: 'Variant 3 · Rural — two matches, no arrows',
context: 'Little Ashby Church of England Primary School — Ages 4–11 · Mixed · Church of England · Voluntary controlled',
lede: 'Other primary schools near Little Ashby Church of England Primary School.',
metric: 'Reading, writing & maths',
caption: 'Meeting the expected standard at key stage 2, 2025.',
thisValue: '66%',
note: 'Nothing matched on religious character within range. At <b>tier 3</b> the lede drops the phrase “with a similar intake” and the chips fall back to the plain phase. Two cards fit the row, so the arrows are not rendered at all. One school fewer and the section would not render either.',
schools: [
{ name:'Great Marden Primary School', distance:'4.2', meta:'Community school · Ages 4–11', shared:['Primary school'], loose:true, value:'71%' },
{ name:'Ashby Vale Academy', distance:'7.8', meta:'Academy converter · Ages 4–11', shared:['Primary school'], loose:true, value:'58%' },
],
},
];
const selected = Object.fromEntries(variants.map((v) => [v.id, new Set()]));
function card(v, s, i) {
const on = selected[v.id].has(i);
const absent = s.value === 'Not published';
return `
<li class="school">
<p class="distance">${PIN}${s.distance} miles away</p>
<h3><a href="#">${s.name}</a></h3>
<p class="meta">${s.meta}</p>
<ul class="shared">${s.shared.map((c) => `<li class="${s.loose ? 'loose' : ''}">${c}</li>`).join('')}</ul>
<div class="metric">
<p class="value ${absent ? 'absent' : ''}">${s.value}</p>
<p class="label">${v.metric}</p>
<p class="ref">${v.thisValue} at this school</p>
</div>
<button class="add" type="button" data-variant="${v.id}" data-index="${i}" aria-pressed="${on}">
${on ? '✓ Added to compare' : '+ Add to compare'}<span class="sr"> — ${s.name}</span>
</button>
</li>`;
}
function render() {
document.getElementById('variants').innerHTML = variants.map((v) => {
const count = selected[v.id].size;
// Three fit the row, so anything more is what the arrows are for.
const scrollable = v.schools.length > 3;
return `
<section class="variant">
<div class="context">
<p class="eyebrow">${v.eyebrow}</p>
<h2>${v.context}</h2>
</div>
<div class="card">
<div class="top">
<div>
<h2 id="h-${v.id}">Similar schools nearby</h2>
<p class="lede">${v.lede}</p>
</div>
${scrollable ? `<div class="arrows">
<button class="arrow" type="button" data-scroll="prev" data-variant="${v.id}" aria-label="Previous schools" aria-controls="sc-${v.id}">${CHEV('prev')}</button>
<button class="arrow" type="button" data-scroll="next" data-variant="${v.id}" aria-label="More schools" aria-controls="sc-${v.id}">${CHEV('next')}</button>
</div>` : ''}
</div>
<ul class="scroller" id="sc-${v.id}" ${scrollable ? `tabindex="0" role="group" aria-labelledby="h-${v.id}"` : ''}>
${v.schools.map((s, i) => card(v, s, i)).join('')}
</ul>
<div class="footer">
<p aria-live="polite">
<strong>${count ? `${count} school${count === 1 ? '' : 's'} selected` : 'Compare side by side'}</strong>
${count ? 'This school is included automatically.' : 'Add a school to compare it with this one.'}
</p>
<a class="compare" href="#" aria-disabled="${count ? 'false' : 'true'}">
${count ? `Compare ${count + 1} schools` : 'Compare'} →
</a>
</div>
<p class="caption">Distances are straight-line from this school, not road distance.
${v.caption} Fictional schools and figures for this mockup.</p>
</div>
<p class="note">${v.note}</p>
</section>`;
}).join('');
variants.forEach((v) => {
const scroller = document.getElementById(`sc-${v.id}`);
if (scroller) syncArrows(v.id, scroller);
});
}
/** An arrow that scrolls nowhere is a dead control, so each end disables its own.
*
* EDGE is not paranoia. The scroller carries 2px of padding so focus rings are
* not clipped, and scroll-snap treats that padding as the first card's snap
* position — so a scroller sitting at its start reports scrollLeft 2, not 0.
* Sub-pixel rounding at other zoom levels moves it again. Testing against an
* exact 0 leaves the back arrow live at the start, pointing nowhere. */
const EDGE = 8;
function syncArrows(id, scroller) {
const max = scroller.scrollWidth - scroller.clientWidth;
const atStart = scroller.scrollLeft <= EDGE;
const atEnd = scroller.scrollLeft >= max - EDGE;
// Drives the mobile scroll-fade, so it is computed even where no arrow is
// rendered to consume it.
scroller.dataset.atEnd = String(atEnd);
const prev = document.querySelector(`.arrow[data-scroll="prev"][data-variant="${id}"]`);
const next = document.querySelector(`.arrow[data-scroll="next"][data-variant="${id}"]`);
if (!prev || !next) return;
prev.disabled = atStart;
next.disabled = atEnd;
}
document.getElementById('variants').addEventListener('click', (event) => {
const arrow = event.target.closest('.arrow');
if (arrow) {
const scroller = document.getElementById(`sc-${arrow.dataset.variant}`);
// A page is what the reader can see, so the viewport is the step.
scroller.scrollBy({ left: (arrow.dataset.scroll === 'next' ? 1 : -1) * scroller.clientWidth, behavior: 'smooth' });
return;
}
const button = event.target.closest('.add');
if (!button) return;
const { variant, index } = button.dataset;
const set = selected[variant];
const i = Number(index);
// Scroll position is DOM state, not React state; keep it across the re-render.
const offset = document.getElementById(`sc-${variant}`).scrollLeft;
set.has(i) ? set.delete(i) : set.add(i);
render();
const scroller = document.getElementById(`sc-${variant}`);
scroller.scrollLeft = offset;
syncArrows(variant, scroller);
document.querySelector(`.add[data-variant="${variant}"][data-index="${index}"]`).focus();
}, true);
document.getElementById('variants').addEventListener('scroll', (event) => {
const scroller = event.target.closest('.scroller');
if (scroller) syncArrows(scroller.id.replace('sc-', ''), scroller);
}, true);
const themeButton = document.getElementById('theme');
themeButton.addEventListener('click', () => {
const dark = document.documentElement.dataset.theme === 'dark';
document.documentElement.dataset.theme = dark ? 'light' : 'dark';
themeButton.textContent = dark ? 'Dark theme' : 'Light theme';
});
render();
</script>
</body>
</html>
+12 -4
View File
@@ -1,8 +1,16 @@
# API Configuration
NEXT_PUBLIC_API_URL=http://localhost:8000/api
# Browser requests use the same-origin Next.js proxy.
NEXT_PUBLIC_API_URL=/api
# Production API URL (for deployment)
# NEXT_PUBLIC_API_URL=https://api.schoolcompare.co.uk/api
# Absolute URL for server-side fetching and the proxy; include /api.
# In the managed container network this is http://backend:80/api (staging differs).
FASTAPI_URL=http://localhost:8000/api
# Payload CMS runtime configuration. Use the managed environment's database;
# Payload owns the payload schema, independently of the school marts.
DATABASE_URL=postgresql://schoolcompare:CHANGE_THIS_PASSWORD@localhost:5432/schoolcompare
# Generate a secret: python -c "import secrets; print(secrets.token_urlsafe(32))"
# Use distinct secrets for staging and production.
PAYLOAD_SECRET=CHANGE_THIS_TO_A_SECURE_RANDOM_SECRET
# Node Environment
NODE_ENV=development
+1
View File
@@ -39,3 +39,4 @@ yarn-error.log*
# typescript
*.tsbuildinfo
next-env.d.ts
+12 -288
View File
@@ -1,291 +1,15 @@
# Deployment Guide
# Frontend deployment
This guide covers deployment options for the SchoolCompare Next.js application.
Next.js and Payload run in the same frontend container. The maintained deployment
procedure is [docs/DEPLOY.md](../docs/DEPLOY.md), with the production and staging
Portainer compose files at the repository root.
## Deployment Options
The frontend Dockerfile builds a standalone Next.js image. Runtime configuration
supplies `FASTAPI_URL`, `DATABASE_URL` and `PAYLOAD_SECRET`; uploaded CMS media is
persisted in a volume. Promote the built image through the repository's Gitea
workflow after human staging approval.
### Option 1: Vercel (Recommended for Next.js)
Vercel is the easiest and most optimized platform for Next.js applications.
#### Steps:
1. **Install Vercel CLI**:
```bash
npm install -g vercel
```
2. **Login to Vercel**:
```bash
vercel login
```
3. **Deploy**:
```bash
vercel --prod
```
4. **Configure Environment Variables** in Vercel dashboard:
- `NEXT_PUBLIC_API_URL`: Your FastAPI endpoint (e.g., `https://api.schoolcompare.co.uk/api`)
- `FASTAPI_URL`: Same as above for server-side requests
#### Benefits:
- Automatic HTTPS
- Global CDN
- Zero-config deployment
- Automatic preview deployments
- Built-in analytics
---
### Option 2: Docker (Self-hosted)
Deploy using Docker containers for full control.
#### Prerequisites:
- Docker 20+
- Docker Compose 2+
#### Steps:
1. **Build Docker Image**:
```bash
docker build -t schoolcompare-nextjs:latest .
```
2. **Run with Docker Compose**:
```bash
# Create .env file with production variables
echo "NEXT_PUBLIC_API_URL=https://api.schoolcompare.co.uk/api" > .env
echo "FASTAPI_URL=http://backend:8000/api" >> .env
# Start services
docker-compose up -d
```
3. **Verify Deployment**:
```bash
curl http://localhost:3000
```
#### Environment Variables:
- `NEXT_PUBLIC_API_URL`: Public API endpoint (client-side)
- `FASTAPI_URL`: Internal API endpoint (server-side)
- `NODE_ENV`: `production`
---
### Option 3: PM2 (Node.js Process Manager)
Deploy directly on a Node.js server using PM2.
#### Prerequisites:
- Node.js 24+
- PM2 (`npm install -g pm2`)
#### Steps:
1. **Build Application**:
```bash
npm run build
```
2. **Create PM2 Ecosystem File** (`ecosystem.config.js`):
```javascript
module.exports = {
apps: [{
name: 'schoolcompare-nextjs',
script: 'npm',
args: 'start',
cwd: '/path/to/nextjs-app',
instances: 'max',
exec_mode: 'cluster',
env: {
NODE_ENV: 'production',
PORT: 3000,
NEXT_PUBLIC_API_URL: 'https://api.schoolcompare.co.uk/api',
FASTAPI_URL: 'http://localhost:8000/api',
},
}],
};
```
3. **Start with PM2**:
```bash
pm2 start ecosystem.config.js
pm2 save
pm2 startup
```
---
### Option 4: Nginx Reverse Proxy
Use Nginx as a reverse proxy in front of Next.js.
#### Nginx Configuration:
```nginx
server {
listen 80;
server_name schoolcompare.co.uk;
# Redirect to HTTPS
return 301 https://$server_name$request_uri;
}
server {
listen 443 ssl http2;
server_name schoolcompare.co.uk;
# SSL Configuration
ssl_certificate /etc/ssl/certs/schoolcompare.crt;
ssl_certificate_key /etc/ssl/private/schoolcompare.key;
# Security Headers
# frame-ancestors replaces X-Frame-Options so the analytics subdomain
# (Umami heatmap/recorder) can embed the site in an iframe.
add_header Content-Security-Policy "frame-ancestors 'self' https://analytics.schoolcompare.co.uk" always;
add_header X-Content-Type-Options "nosniff" always;
add_header X-XSS-Protection "1; mode=block" always;
# Proxy to Next.js
location / {
proxy_pass http://localhost:3000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection 'upgrade';
proxy_set_header Host $host;
proxy_cache_bypass $http_upgrade;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
# Proxy to FastAPI
location /api/ {
proxy_pass http://localhost:8000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
# Cache static files
location /_next/static/ {
proxy_pass http://localhost:3000;
add_header Cache-Control "public, max-age=31536000, immutable";
}
}
```
---
## Pre-Deployment Checklist
- [ ] Run `npm run build` successfully
- [ ] Run `npm test` - all tests pass
- [ ] Environment variables configured
- [ ] FastAPI backend accessible
- [ ] Database migrations applied
- [ ] SSL certificates configured (production)
- [ ] Domain DNS configured
- [ ] Monitoring/logging set up
- [ ] Backup strategy in place
---
## Post-Deployment Verification
1. **Health Check**:
```bash
curl https://schoolcompare.co.uk
```
2. **Test Routes**:
- Home: `https://schoolcompare.co.uk/`
- School Page: `https://schoolcompare.co.uk/school/100001`
- Compare: `https://schoolcompare.co.uk/compare`
- Rankings: `https://schoolcompare.co.uk/rankings`
3. **Check SEO**:
- Sitemap: `https://schoolcompare.co.uk/sitemap.xml`
- Robots: `https://schoolcompare.co.uk/robots.txt`
4. **Performance Audit**:
- Run Lighthouse in Chrome DevTools
- Target scores: 90+ for Performance, Accessibility, Best Practices, SEO
---
## Monitoring
### Recommended Tools:
- **Vercel Analytics** (if using Vercel)
- **Sentry** for error tracking
- **Google Analytics** for user analytics
- **Uptime Robot** for uptime monitoring
### Health Check Endpoint:
The application automatically serves health data at the root route.
---
## Rollback Procedure
### Vercel:
```bash
vercel rollback
```
### Docker:
```bash
docker-compose down
docker-compose up -d --force-recreate
```
### PM2:
```bash
pm2 stop schoolcompare-nextjs
# Restore previous build
pm2 start schoolcompare-nextjs
```
---
## Troubleshooting
### Issue: API requests failing
- **Solution**: Check `NEXT_PUBLIC_API_URL` and `FASTAPI_URL` environment variables
- **Verify**: FastAPI backend is accessible from Next.js container/server
### Issue: Build fails
- **Solution**: Check Node.js version (requires 24+)
- **Clear cache**: `rm -rf .next node_modules && npm install && npm run build`
### Issue: Slow page loads
- **Solution**: Enable caching in API calls
- **Check**: Network latency to FastAPI backend
- **Verify**: CDN is serving static assets
---
## Security Considerations
- ✅ HTTPS enabled
- ✅ Security headers configured (X-Frame-Options, CSP, etc.)
- ✅ API keys in environment variables (never in code)
- ✅ CORS properly configured
- ✅ Rate limiting on API endpoints
- ✅ Regular security updates
- ✅ Dependency vulnerability scanning
---
## Support
For deployment issues, contact the DevOps team or refer to:
- [Next.js Deployment Docs](https://nextjs.org/docs/deployment)
- [Vercel Documentation](https://vercel.com/docs)
- [Docker Documentation](https://docs.docker.com/)
Earlier Vercel and standalone deployment recipes have been retired from this file
because they do not describe the current CMS, persistence and promotion setup.
See [development](../docs/DEVELOPMENT.md) for checks and
[publishing](docs/PUBLISHING.md) for CMS operations.
+17
View File
@@ -28,6 +28,10 @@ ENV NODE_ENV=production
ARG FASTAPI_URL=http://backend:80/api
ENV FASTAPI_URL=${FASTAPI_URL}
ARG BUILD_SHA=development
ARG BUILD_ID=development
RUN node -e 'require("fs").writeFileSync("build-info.json", JSON.stringify({sha:process.argv[1],build_id:process.argv[2]}))' "$BUILD_SHA" "$BUILD_ID"
# Build application
RUN npm run build
@@ -53,6 +57,13 @@ COPY --from=builder /app/.next/static ./.next/static
# a miss here is a silent 500 on /opengraph-image, not a build failure.
COPY --from=builder /app/assets ./assets
# Payload writes uploads here, and the compose file mounts a named volume over
# it. The directory must exist and be owned by the runtime user BEFORE the
# mount: Docker seeds a fresh named volume from the image path, so a missing or
# root-owned directory here makes every upload fail with EACCES at runtime,
# long after the build passed. The chown below covers it.
RUN mkdir -p /app/media
# Set correct permissions
RUN chown -R nextjs:nodejs /app
@@ -63,6 +74,12 @@ USER nextjs
EXPOSE 3000
# Set environment variables
ARG BUILD_SHA=development
ARG BUILD_ID=development
LABEL io.schoolcompare.build-id=$BUILD_ID
LABEL io.schoolcompare.commit=$BUILD_SHA
COPY --from=builder /app/build-info.json ./build-info.json
ENV PORT=3000
ENV HOSTNAME="0.0.0.0"
+43 -141
View File
@@ -1,156 +1,58 @@
# SchoolCompare Next.js Application
# SchoolCompare frontend and CMS
Modern Next.js application for comparing primary school KS2 performance across England.
Next.js App Router with React, TypeScript, CSS Modules, Chart.js, Leaflet and
Payload CMS. It serves school search, comparisons, rankings, school/place detail
pages and editorial content across England.
## Features
Start with the [repository overview](../README.md),
[architecture](../docs/ARCHITECTURE.md) and [development checks](../docs/DEVELOPMENT.md).
- **Server-Side Rendering (SSR)**: Fast initial page loads with pre-rendered content
- **Individual School Pages**: Dedicated pages for each school with full SEO optimization
- **Side-by-Side Comparison**: Compare up to 5 schools simultaneously
- **School Rankings**: Top-performing schools by various metrics
- **Interactive Maps**: Leaflet integration for geographic visualization
- **Performance Charts**: Chart.js visualizations for historical data
- **Responsive Design**: Mobile-first approach with full responsive support
- **SEO Optimized**: Dynamic sitemaps, meta tags, and structured data
## Source map
## Tech Stack
| Path | Purpose |
|---|---|
| `app/(frontend)/` | Public root layout, server pages and FastAPI proxy |
| `app/(payload)/` | Payload root layout, `/admin` and `/cms-api` |
| `app/robots.ts`, `app/opengraph-image.tsx`, root icons | Site-wide metadata endpoints |
| `components/` | Client views and reusable display components |
| `components/school/` | School detail sections |
| `lib/api.ts`, `lib/types.ts` | Fetch wrappers and manual school API types |
| `lib/schoolSections.ts`, `lib/compareLogic.ts` | Presentation decisions and data preparation |
| `context/`, `hooks/` | Comparison state, suggestion state and responsive behaviour |
| `collections/`, `blocks/`, `migrations/` | CMS schema and production migrations |
| `__tests__/` | Jest and React Testing Library tests |
- **Framework**: Next.js 16 (App Router)
- **Language**: TypeScript 5
- **Styling**: CSS Modules + CSS Variables
- **State Management**: React Context API + URL state
- **Data Fetching**: SWR (client-side) + Next.js fetch (server-side)
- **Charts**: Chart.js + react-chartjs-2
- **Maps**: Leaflet + react-leaflet
- **Testing**: Jest + React Testing Library
- **Validation**: Zod
Do not introduce a shared `app/layout.tsx`: public pages and Payload have separate
root layouts. Keep root metadata files outside the route groups.
## Getting Started
## Data and state
### Prerequisites
Server pages fetch initial data directly from `FASTAPI_URL`. Browser fetches use
`/api` by default, forwarded by `app/(frontend)/api/[...path]/route.ts`.
`FASTAPI_URL` must include `/api`. See `.env.example` for CMS and API settings.
- Node.js 24+ (using nvm recommended)
- FastAPI backend running on port 8000
State uses React hooks/context, URL search parameters and localStorage for the
comparison basket. SWR is not installed. Maps use dynamic Leaflet wrappers.
Revalidation intervals are configured in fetch wrappers and pages; they vary by
resource. Backend reloads do not automatically invalidate every Next.js cache.
### Installation
## Commands
```bash
# Install dependencies
npm install
# Copy environment variables
cp .env.example .env.local
# Update .env.local with your configuration
```
### Development
```bash
# Start development server
npm run dev
# Open http://localhost:3000
```
### Building
```bash
# Build for production
```sh
npm ci
npm run typecheck
npm test -- --runInBand
npm run build
# Start production server
npm start
```
### Testing
`test:watch` and `test:coverage` are also available. There is no `lint` script.
A running application needs the backend/data environment described in the
[development guide](../docs/DEVELOPMENT.md).
```bash
# Run tests
npm test
After CMS field or editor changes, run `npm run generate:importmap`. Keep
`payload-types.ts` generated from the CMS schema rather than editing it by hand.
The build must work without a database connection; avoid module-scope CMS queries
and DB-backed `generateStaticParams` functions.
# Run tests in watch mode
npm run test:watch
# Run tests with coverage
npm run test:coverage
```
### Linting
```bash
# Run ESLint
npm run lint
```
## Project Structure
```
nextjs-app/
├── app/ # App Router pages
│ ├── layout.tsx # Root layout
│ ├── page.tsx # Home page
│ ├── compare/ # Compare page
│ ├── rankings/ # Rankings page
│ ├── school/[urn]/ # Individual school pages
│ ├── sitemap.ts # Dynamic sitemap
│ └── robots.ts # Robots.txt
├── components/ # React components
│ ├── SchoolCard.tsx # School card component
│ ├── FilterBar.tsx # Search/filter controls
│ ├── ComparisonView.tsx # Comparison interface
│ ├── RankingsView.tsx # Rankings table
│ └── ...
├── lib/ # Utility libraries
│ ├── api.ts # API client
│ ├── types.ts # TypeScript types
│ └── utils.ts # Helper functions
├── hooks/ # Custom React hooks
├── context/ # React Context providers
├── styles/ # Global styles
├── public/ # Static assets
└── __tests__/ # Test files
```
## Environment Variables
| Variable | Description | Default |
|----------|-------------|---------|
| `NEXT_PUBLIC_API_URL` | Public API endpoint (client-side) | `http://localhost:8000/api` |
| `FASTAPI_URL` | Server-side API endpoint | `http://localhost:8000/api` |
| `NODE_ENV` | Environment mode | `development` |
## Performance Optimizations
- **Server-Side Rendering**: Initial HTML rendered on server
- **Static Generation**: Where possible, pages are pre-generated
- **Image Optimization**: Next.js Image component with AVIF/WebP support
- **Code Splitting**: Automatic route-based code splitting
- **Dynamic Imports**: Heavy components loaded on demand
- **API Caching**: Configurable revalidation for data fetching
- **Bundle Optimization**: Tree shaking and minification
- **Compression**: Gzip compression enabled
## SEO Features
- **Dynamic Meta Tags**: Generated per page with Next.js Metadata API
- **Open Graph**: Social media optimization
- **JSON-LD**: Structured data for search engines
- **Sitemap**: Auto-generated from database
- **Robots.txt**: Search engine crawling rules
- **Canonical URLs**: Duplicate content prevention
## Browser Support
- Chrome (latest)
- Firefox (latest)
- Safari (latest)
- Edge (latest)
## License
Proprietary - SchoolCompare
## Support
For issues and questions, please contact the development team.
See [publishing](docs/PUBLISHING.md) for CMS operations and
[deployment](../docs/DEPLOY.md) for staging and production promotion.
@@ -0,0 +1,31 @@
/**
* The /api/* proxy is public. Anything it forwards is on the internet.
*
* @jest-environment node
*/
// The docblock above is load-bearing. jest.config.js sets jsdom globally, and
// NextRequest/NextResponse need the Web Fetch API globals that only the node
// environment provides — under jsdom this suite fails on import, not on an
// assertion.
import { NextRequest } from 'next/server';
import { GET } from '@/app/(frontend)/api/[...path]/route';
function request(path: string) {
return new NextRequest(`http://localhost:3000/api/${path}`);
}
describe('public API proxy', () => {
it('refuses to forward internal-only paths', async () => {
// /api/flags names every unreleased feature and its state. Forwarding it
// publishes the thing shipping dark exists to keep quiet.
const res = await GET(request('flags'), { params: Promise.resolve({ path: ['flags'] }) });
expect(res.status).toBe(404);
});
it('does not deny a path that merely starts with the same letters', async () => {
// A prefix match would take /api/flagship down with /api/flags.
const res = await GET(
request('flagship'), { params: Promise.resolve({ path: ['flagship'] }) });
expect(res.status).not.toBe(404);
});
});
@@ -0,0 +1,34 @@
import { metadata } from '@/app/(frontend)/about/page';
import { personJsonLd, organizationJsonLd } from '@/lib/jsonld';
describe('/about metadata', () => {
it('canonicalises to the bare path', () => {
expect(metadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/about');
});
});
describe('author structured data', () => {
it('describes a Person with a first name and a photo', () => {
const person = personJsonLd();
expect(person['@type']).toBe('Person');
expect(person.name).toBe('Tudor');
expect(person.image).toBe('https://www.schoolcompare.co.uk/brand/tudor.jpg');
expect(person.url).toBe('https://www.schoolcompare.co.uk/about');
});
it('never publishes a surname or an employer', () => {
// Author identity constraint: first name only. A surname here would be
// the one place it leaks, since JSON-LD is machine-read and archived.
const serialised = JSON.stringify(personJsonLd());
expect(serialised).not.toMatch(/familyName|Sitaru/i);
expect(serialised).not.toMatch(/worksFor|affiliation/i);
});
it('describes the site as an Organization the Person authors for', () => {
const org = organizationJsonLd();
expect(org['@type']).toBe('Organization');
expect(org.name).toBe('schoolcompare');
expect(org.url).toBe('https://www.schoolcompare.co.uk');
});
});
@@ -0,0 +1,66 @@
/**
* The blog index imports getCachedPayload, which pulls in Payload — ESM-only,
* and next/jest will not transform node_modules. Mocking that one module keeps
* the page's metadata testable without loading the CMS; the mock is never
* called, because `metadata` is a static export evaluated at import time.
*/
jest.mock('@/lib/payload', () => ({ getCachedPayload: jest.fn() }));
import { metadata } from '@/app/(frontend)/blog/page';
import { blogPostingJsonLd, breadcrumbJsonLd } from '@/lib/jsonld';
const post = {
title: 'What the data cannot tell you',
slug: 'what-the-data-cannot-tell-you',
excerpt: 'Results describe one year group on a handful of days.',
publishedAt: '2026-09-15T00:00:00.000Z',
};
describe('/blog metadata', () => {
it('canonicalises to the bare path', () => {
expect(metadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/blog');
});
});
describe('BlogPosting structured data', () => {
it('names the same Person entity the about page declares', () => {
// By @id, not by repeating the person: search engines must resolve every
// post and the about page to one author entity, or the site has several.
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld['@type']).toBe('BlogPosting');
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk/about#tudor');
expect(ld.publisher['@id']).toBe('https://www.schoolcompare.co.uk#organization');
});
it('attributes to the organization when the about page is dark', () => {
/*
* The two flags are independent, so blog-on-about-off is a reachable
* state. The Person entity lives at /about#tudor and that URL 404s while
* the flag is dark, so claiming it would declare an author that resolves
* to nothing — worse for the blog's credibility than having no named
* author at all. Attribute to the publisher instead.
*/
const ld = blogPostingJsonLd(post, { namedAuthor: false });
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk#organization');
expect(JSON.stringify(ld)).not.toContain('/about');
});
it('carries a self-referencing canonical url and the publish date', () => {
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld.url).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
expect(ld.datePublished).toBe('2026-09-15T00:00:00.000Z');
});
});
describe('breadcrumbs', () => {
it('places the post under the blog index', () => {
const ld = breadcrumbJsonLd(post);
expect(ld.itemListElement[0].item).toBe('https://www.schoolcompare.co.uk/blog');
expect(ld.itemListElement[1].item).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
});
});
@@ -0,0 +1,51 @@
import { fireEvent, render, screen } from '@testing-library/react';
import HomePage from '@/app/(frontend)/page';
import SchoolPage from '@/app/(frontend)/school/[slug]/page';
import ErrorPage from '@/app/(frontend)/error';
import { APIFetchError, fetchSchools, fetchFilters, fetchSchoolDetails } from '@/lib/api';
import { fetchPlace, fetchPlaces } from '@/lib/places';
jest.mock('@/lib/api', () => ({
...jest.requireActual('@/lib/api'),
fetchSchools: jest.fn(),
fetchSchoolDetails: jest.fn(),
fetchFilters: jest.fn(async () => ({})),
fetchDataInfo: jest.fn(async () => null),
fetchNationalAverages: jest.fn(async () => null),
}));
jest.mock('@/lib/flags', () => ({ getFlags: jest.fn(async () => ({})) }));
jest.mock('next/navigation', () => ({
notFound: () => { throw new Error('NEXT_NOT_FOUND'); },
redirect: jest.fn(),
}));
const realFetch = global.fetch;
afterEach(() => { global.fetch = realFetch; jest.clearAllMocks(); });
test('school outages propagate; only a real 404 becomes not found', async () => {
const request = { params: Promise.resolve({ slug: '100001-school' }) };
const outage = new APIFetchError('unavailable', 503);
jest.mocked(fetchSchoolDetails).mockRejectedValueOnce(outage);
await expect(SchoolPage(request)).rejects.toBe(outage);
jest.mocked(fetchSchoolDetails).mockRejectedValueOnce(new APIFetchError('missing', 404));
await expect(SchoolPage(request)).rejects.toThrow('NEXT_NOT_FOUND');
});
test('homepage search failure is not returned as an empty successful page', async () => {
jest.mocked(fetchSchools).mockRejectedValueOnce(new APIFetchError('unavailable', 503));
await expect(HomePage({ searchParams: Promise.resolve({ search: 'school' }) })).rejects.toThrow('unavailable');
});
test('a place is absent only on 404; other failures propagate', async () => {
global.fetch = jest.fn().mockResolvedValue({ ok: false, status: 404 });
await expect(fetchPlace('town', 'example')).resolves.toBeNull();
jest.mocked(global.fetch).mockResolvedValue({ ok: false, status: 503 } as Response);
await expect(fetchPlace('town', 'example')).rejects.toMatchObject({ status: 503 });
await expect(fetchPlaces()).rejects.toMatchObject({ status: 503 });
});
test('the error boundary offers a retry without showing an empty search', () => {
const reset = jest.fn();
render(<ErrorPage error={new Error('offline')} reset={reset} />);
fireEvent.click(screen.getByRole('button', { name: 'Try again' }));
expect(reset).toHaveBeenCalledTimes(1);
});
+169
View File
@@ -0,0 +1,169 @@
import { metadata as homeMetadata } from '@/app/(frontend)/page';
import { metadata as rankingsMetadata } from '@/app/(frontend)/rankings/page';
import { metadata as admissionsMetadata } from '@/app/(frontend)/admissions/page';
import { generateMetadata as compareMetadata } from '@/app/(frontend)/compare/page';
import { metadata as rootMetadata } from '@/app/(frontend)/layout';
describe('canonical URLs', () => {
it('the homepage canonicalises to the bare root', () => {
// page.tsx reads eleven search params. Without this, every filter
// combination is a crawlable near-duplicate of the one page we want to
// rank for "compare schools".
expect(homeMetadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/');
});
it('rankings canonicalises to the bare path', () => {
expect(rankingsMetadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/rankings');
});
it('admissions canonicalises to the bare path', () => {
expect(admissionsMetadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/admissions');
});
});
describe('/compare indexability', () => {
it('the bare compare page is indexable and canonical to itself', async () => {
// This is the landing page for the "compare schools" head term.
const meta = await compareMetadata({ searchParams: Promise.resolve({}) });
expect(meta.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/compare');
expect(meta.robots).toBeUndefined();
});
it('a comparison of specific schools is noindex, follow', async () => {
// ~317 million pairs before triples. Indexing the parameter space would
// swamp everything else in the corpus.
const meta = await compareMetadata({
searchParams: Promise.resolve({ urns: '100001,100002' }),
});
expect(meta.robots).toEqual({ index: false, follow: true });
});
it('a parameterised comparison still canonicalises to the bare path', async () => {
// follow:true plus a canonical means the outbound links to each school
// page still pass value even though this URL is not indexed.
const meta = await compareMetadata({
searchParams: Promise.resolve({ urns: '100001,100002' }),
});
expect(meta.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/compare');
});
});
/*
* W8 — snippet copy for the C1 cluster.
*
* The baseline (GSC, 16 months to 2026-08-20) showed these pages ranking on
* page one and converting at a tenth of the normal rate: "compare school
* performance" at position 6.1 with 0.43% CTR, against 9.16% for the brand
* query from the same neighbourhood. The SERP is dominated by the DfE's own
* "Compare school performance" service, so the job of this copy is to say
* what that service does not offer, without losing intent match on the title.
*
* These tests guard the mechanics that make a snippet work — length, intent
* keyword, differentiator, no brand-first — not the exact wording, which
* should stay free to iterate.
*/
// Google truncates titles near 60 characters and descriptions near 155.
const TITLE_MAX = 60;
const DESC_MIN = 110;
const DESC_MAX = 155;
type Meta = { title?: unknown; description?: unknown };
const titleOf = (m: Meta): string => {
const t = m.title as string | { absolute?: string } | undefined;
return typeof t === 'string' ? t : (t?.absolute ?? '');
};
describe('C1 snippet copy', () => {
const pages: Array<[string, Meta, RegExp]> = [
['home', homeMetadata as Meta, /compare schools/i],
['rankings', rankingsMetadata as Meta, /league table/i],
['admissions', admissionsMetadata as Meta, /admission/i],
];
for (const [name, meta, intent] of pages) {
it(`${name}: title carries the search intent and fits the SERP`, () => {
const t = titleOf(meta);
expect(t).toMatch(intent);
expect(t.length).toBeLessThanOrEqual(TITLE_MAX);
});
it(`${name}: title does not open with the brand`, () => {
// The measured 0.43% CTR came from a brand-first title. The most
// valuable pixels go to the thing the searcher typed.
expect(titleOf(meta).toLowerCase().startsWith('schoolcompare')).toBe(false);
});
it(`${name}: description is long enough to be worth reading, short enough to survive`, () => {
const d = meta.description as string;
expect(d.length).toBeGreaterThanOrEqual(DESC_MIN);
expect(d.length).toBeLessThanOrEqual(DESC_MAX);
});
}
it('the homepage description names what gov.uk does not publish', () => {
// Admissions distance is the one fact the DfE service has no equivalent
// for. If it ever leaves this description, the snippet is competing with
// gov.uk on gov.uk's own ground.
expect(homeMetadata.description).toMatch(/close you had to live|distance/i);
});
it('/compare targets the tool phrasing rather than repeating the homepage', () => {
// Two pages chasing one phrase is how a site competes with itself.
return compareMetadata({ searchParams: Promise.resolve({}) }).then((m) => {
expect(m.title).toMatch(/comparison tool/i);
expect(m.title).not.toBe(titleOf(homeMetadata as Meta));
});
});
it('no C1 page claims a school count that will drift', () => {
// The corpus moves with every data refresh; this repo has already shipped
// one copy bug of that kind ("three schools" against MAX_SCHOOLS = 5).
for (const [, meta] of pages) {
expect(meta.description as string).not.toMatch(/\b\d{2},\d{3}\b|\b\d{2},000\b/);
}
});
});
/**
* The share card must be declared, not inherited.
*
* `app/opengraph-image.tsx` is a metadata file convention, and it does attach
* to routes in the app root segment — `_not-found` gets an og:image from it.
* It does NOT attach to the site's pages, which live in the `(frontend)`
* route group whose own layout is a root layout. Staging served og:title,
* og:description, og:url, og:site_name and og:type and no og:image at all,
* so every link pasted into a chat rendered bare.
*
* The file stays at the app root, because /robots.txt and /icon.png depend on
* it being there. The site's root layout points at the route it generates.
*/
describe('the share card', () => {
it('declares an opengraph image on the site root layout', () => {
// No og:image means every link pasted into a chat renders bare.
const images = rootMetadata.openGraph?.images;
expect(images).toBeTruthy();
expect(JSON.stringify(images)).toContain('/opengraph-image');
});
it('declares a twitter image too', () => {
// twitter.card is summary_large_image. Claiming a large-image card and
// supplying no image is worse than claiming a summary card.
// Metadata['twitter'] is a union and `card` is not on every member, so
// this reads the serialised shape rather than narrowing the type.
const twitter = JSON.stringify(rootMetadata.twitter);
expect(twitter).toContain('summary_large_image');
expect(twitter).toContain('/opengraph-image');
});
it('resolves the card to an absolute url via metadataBase', () => {
// The e2e journey does `new URL(ogUrl)`, which throws on a relative path.
expect(rootMetadata.metadataBase?.toString())
.toBe('https://www.schoolcompare.co.uk/');
});
});
@@ -0,0 +1,66 @@
/**
* next.config.mjs carries the staging noindex rule. Breaking it turns
* stx.schoolcompare.co.uk into a fully crawlable duplicate of production,
* and nothing else in the suite would notice.
*
* The non-null assertions are deliberate: every key asserted here is optional
* on NextConfig, and a missing one is precisely the regression under test, so
* the assertion below should fail the test rather than the compile.
*/
import nextConfig from '@/next.config.mjs';
async function headerRules() {
return nextConfig.headers!();
}
describe('next.config.mjs', () => {
it('keeps the staging host out of the index', async () => {
const headers = await headerRules();
const stagingRule = headers.find((rule) =>
rule.has?.some(
(cond) => cond.type === 'host' && cond.value === 'stx.schoolcompare.co.uk',
),
);
expect(stagingRule).toBeDefined();
expect(stagingRule!.headers).toContainEqual({
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
});
});
it('still emits standalone output for the Docker runner', () => {
expect(nextConfig.output).toBe('standalone');
});
it('still traces the share-card fonts into the standalone bundle', () => {
expect(nextConfig.outputFileTracingIncludes!['/opengraph-image']).toEqual([
'./assets/**',
]);
});
it('still allows the analytics subdomain to frame the site', async () => {
const headers = await headerRules();
const csp = headers
.flatMap((rule) => rule.headers)
.find((header) => header.key === 'Content-Security-Policy');
expect(csp).toBeDefined();
expect(csp!.value).toContain('https://analytics.schoolcompare.co.uk');
});
});
describe('admin surface', () => {
it('serves noindex on the admin panel and the CMS API', async () => {
// robots.txt disallows these too, but a Disallow only blocks crawling — a
// URL found from an external link can still be indexed without ever being
// fetched. This header is what actually keeps them out.
const headers = await headerRules();
for (const source of ['/admin/:path*', '/cms-api/:path*']) {
const rule = headers.find((entry) => entry.source === source);
expect(rule).toBeDefined();
expect(rule!.headers).toContainEqual({
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
});
}
});
});
@@ -0,0 +1,42 @@
import { generateMetadata as placeMeta } from '@/app/(frontend)/schools/[place]/page';
jest.mock('@/lib/places', () => ({
...jest.requireActual('@/lib/places'),
fetchPlace: jest.fn(async (kind: string, slug: string) =>
slug === 'atlantis' ? null : ({
place: { kind, slug, name: 'Brentwood', count: 29,
parent_authority: 'Essex' },
schools: [], averages: { rwm_expected_pct: 63, attainment_8_score: null },
})),
fetchPlaces: jest.fn(async () => []),
}));
describe('place page metadata', () => {
it('titles the page the way the place is searched', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'brentwood' }) });
expect((m.title as { absolute: string }).absolute).toMatch(/schools in brentwood/i);
});
it('canonicalises to its own path on the www host', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'brentwood' }) });
expect(m.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/schools/brentwood');
});
it('opts out of the layout template, which would double the brand', () => {
// The root layout appends '| schoolcompare' to a plain string title, and
// these titles already carry it — every place page shipped reading
// '... | schoolcompare | schoolcompare' until this was made absolute.
return placeMeta({ params: Promise.resolve({ place: 'brentwood' }) })
.then((m) => {
expect(typeof m.title).toBe('object');
expect((m.title as { absolute: string }).absolute)
.not.toMatch(/schoolcompare.*schoolcompare/);
});
});
it('an unknown place gets a not-found title rather than inventing one', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'atlantis' }) });
expect(m.title).toMatch(/not found/i);
});
});
@@ -0,0 +1,30 @@
/** @jest-environment node */
import { GET } from '@/app/(frontend)/release.json/route';
import { readFile } from 'node:fs/promises';
jest.mock('node:fs/promises', () => ({ readFile: jest.fn() }));
const realFetch = global.fetch;
const identity = { sha: 'a'.repeat(40), build_id: 'b'.repeat(32) };
beforeEach(() => {
jest.mocked(readFile).mockResolvedValue(JSON.stringify(identity));
global.fetch = jest.fn(async () => Response.json(identity));
});
afterEach(() => { global.fetch = realFetch; jest.resetAllMocks(); });
test('reports immutable file identity and backend identity without caching', async () => {
const response = await GET();
expect(response.status).toBe(200);
expect(response.headers.get('Cache-Control')).toBe('no-store');
expect(await response.json()).toEqual({ frontend: identity, backend: identity });
expect(fetch).toHaveBeenCalledWith(expect.stringMatching(/\/api\/release$/), expect.objectContaining({ cache: 'no-store', signal: expect.anything() }));
});
test('missing build metadata cannot pass the release gate', async () => {
jest.mocked(readFile).mockRejectedValueOnce(new Error('missing file'));
expect((await GET()).status).toBe(503);
});
test('backend failure cannot pass the release gate', async () => {
jest.mocked(fetch).mockResolvedValueOnce(new Response('', { status: 503 }));
expect((await GET()).status).toBe(503);
});
+20
View File
@@ -0,0 +1,20 @@
import robots from '@/app/robots';
describe('robots.txt', () => {
it('disallows the admin panel and the CMS API', () => {
const rules = robots().rules;
const rule = Array.isArray(rules) ? rules[0] : rules;
expect(rule.disallow).toEqual(
expect.arrayContaining(['/api/', '/_next/', '/admin/', '/cms-api/']),
);
});
});
describe('sitemap discovery', () => {
it('lists both the proxied school sitemap and the Next-owned content sitemap', () => {
expect(robots().sitemap).toEqual([
'https://www.schoolcompare.co.uk/sitemap.xml',
'https://www.schoolcompare.co.uk/content-sitemap.xml',
]);
});
});
@@ -104,7 +104,7 @@ describe('CompareAdmissions', () => {
render(<CompareAdmissions schools={[grammar]} data={data} isSecondary={true} />);
expect(
screen.getByText(/Entry is by entrance test — the school is selective/),
screen.getByText(/Entry is by entrance test\. The school is selective/),
).toBeInTheDocument();
expect(screen.queryByText(/non-faith primaries/)).toBeNull();
});
@@ -0,0 +1,152 @@
/**
* The postcode check.
*
* This is the one place on the site that answers a question about a specific
* family rather than about a school, so the tests here are mostly about what it
* refuses to say — and that matters more now than it did, because there is only
* one year to answer with. A run of years used to soften a single close call;
* nothing does now, so the "too close to call" band is the whole safety margin.
*/
import { render, screen, fireEvent, waitFor } from '@testing-library/react';
import { CutoffMapPanel } from '@/components/school/CutoffMapPanel';
import { CUTOFF_UNCERTAINTY_M } from '@/components/school/lastDistanceOffered';
import type { School, SchoolAdmissionDistance } from '@/lib/types';
// Leaflet needs a real layout box and network tiles; neither exists in jsdom.
jest.mock('@/components/LeafletCutoffMapInner', () => ({
__esModule: true,
default: () => <div data-testid="cutoff-map" />,
}));
const mockGeocode = jest.fn();
jest.mock('@/lib/api', () => ({
...jest.requireActual('@/lib/api'),
geocodePostcode: (pc: string) => mockGeocode(pc),
}));
const SCHOOL = { urn: 100010, school_name: 'Test Primary', latitude: 51.5, longitude: -0.12 } as School;
const cutoff = (distance_m: number | null, year = 2026): SchoolAdmissionDistance => ({
year, distance_m, route_count: 1, la_name: 'Camden', distance_unit_raw: 'miles',
});
/** A point due north of the school, `metres` away. 1° latitude ≈ 111,320 m. */
function northOf(metres: number) {
return { latitude: SCHOOL.latitude! + metres / 111_320, longitude: SCHOOL.longitude! };
}
const renderPanel = (c = cutoff(800)) =>
render(<CutoffMapPanel schoolInfo={SCHOOL} cutoff={c} />);
async function check(postcode: string) {
fireEvent.change(screen.getByLabelText('Your postcode'), { target: { value: postcode } });
fireEvent.click(screen.getByRole('button', { name: 'Check' }));
}
beforeEach(() => mockGeocode.mockReset());
describe('CutoffMapPanel', () => {
it('renders nothing without a figure to compare against', () => {
const { container } = renderPanel(cutoff(null));
expect(container).toBeEmptyDOMElement();
});
it('renders nothing when the school has no coordinates', () => {
const { container } = render(
<CutoffMapPanel
schoolInfo={{ ...SCHOOL, latitude: null, longitude: null } as School}
cutoff={cutoff(800)}
/>,
);
expect(container).toBeEmptyDOMElement();
});
it('rejects a malformed postcode without calling the geocoder', async () => {
renderPanel();
await check('not a postcode');
expect(await screen.findByRole('alert')).toHaveTextContent(/does not look like a UK postcode/);
expect(mockGeocode).not.toHaveBeenCalled();
});
it('names the year in the verdict, so the figure is never free-floating', async () => {
mockGeocode.mockResolvedValue(northOf(200));
renderPanel(cutoff(800, 2026));
await check('SE23 3NA');
const result = await screen.findByRole('status');
expect(result).toHaveTextContent(/inside the/);
expect(result).toHaveTextContent(/September 2026/);
});
it('reports a home clearly beyond the cut-off', async () => {
mockGeocode.mockResolvedValue(northOf(5000));
renderPanel(cutoff(800));
await check('SE23 3NA');
expect(await screen.findByRole('status')).toHaveTextContent(/beyond the/);
});
it('declines to call a result that sits inside the measurement error', async () => {
// Nominally inside the 800 m cut-off, but by half the uncertainty band —
// which a postcode centroid cannot resolve. With only one year published
// there is nothing else to fall back on, so this must not read as a pass.
mockGeocode.mockResolvedValue(northOf(800 - CUTOFF_UNCERTAINTY_M / 2));
renderPanel(cutoff(800));
await check('SE23 3NA');
const result = await screen.findByRole('status');
expect(result).toHaveTextContent(/too close/);
expect(result).toHaveTextContent(/measurement error/);
// Explanation is supporting text, not part of the bold verdict line.
expect(result.querySelector('[class*="cutoffCheckHeadline"]')!.textContent)
.not.toMatch(/measurement error/);
expect(result).not.toHaveTextContent(/^\S+ away, inside/);
});
it('surfaces a postcode the geocoder cannot find', async () => {
mockGeocode.mockResolvedValue(null);
renderPanel();
await check('ZZ99 9ZZ');
expect(await screen.findByRole('alert')).toHaveTextContent(/could not find that postcode/);
});
it('recovers from a geocoder failure instead of leaving a stale verdict', async () => {
mockGeocode.mockResolvedValue(northOf(200));
renderPanel();
await check('SE23 3NA');
await screen.findByRole('status');
mockGeocode.mockRejectedValue(new Error('network'));
await check('SE23 3NB');
await waitFor(() => expect(screen.queryByRole('status')).not.toBeInTheDocument());
expect(screen.getByRole('alert')).toHaveTextContent(/Something went wrong/);
});
it('keeps the map behind a request until there is a reason to show it', async () => {
renderPanel();
expect(screen.queryByTestId('cutoff-map')).not.toBeInTheDocument();
mockGeocode.mockResolvedValue(northOf(200));
await check('SE23 3NA');
await screen.findByRole('status');
expect(screen.getByTestId('cutoff-map')).toBeInTheDocument();
});
it('can also show the map without a postcode, on request', () => {
renderPanel();
fireEvent.click(screen.getByRole('button', { name: /Show this distance on a map/ }));
expect(screen.getByTestId('cutoff-map')).toBeInTheDocument();
});
it('states its limits before it is used, not with the answer', () => {
renderPanel();
const caveat = screen.getByText(/Distance is the last criterion applied/);
expect(caveat).toBeInTheDocument();
expect(caveat).toHaveTextContent(/not a catchment boundary/);
expect(caveat).toHaveTextContent(/walking route/);
});
});
@@ -0,0 +1,149 @@
import { render, screen } from '@testing-library/react';
import { DestinationsSection } from '@/components/school/DestinationsSection';
import type { DestinationPhase } from '@/lib/types';
import type { DestinationCategory, DestinationStatus } from '@/lib/destinations';
const cell = (
category: DestinationCategory,
pupils: number | null,
status: DestinationStatus = 'published',
) => ({
category, pupils,
percentage: pupils === null ? null : (pupils / 180) * 100,
status,
});
const ALL_PUBLISHED = [
cell('school_sixth_form', 75), cell('sixth_form_college', 21),
cell('further_education', 55), cell('other_education', 6),
cell('apprenticeship', 8), cell('employment', 6),
cell('not_sustained', 5), cell('not_captured', 4),
];
const fullPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: { all: { cohort: 180, categories: ALL_PUBLISHED } },
};
const suppressedPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: [
cell('school_sixth_form', 75), cell('sixth_form_college', null, 'suppressed'),
cell('further_education', 55), cell('other_education', 6),
cell('apprenticeship', 8), cell('employment', 6),
cell('not_sustained', 5), cell('not_captured', 4),
],
},
},
};
describe('DestinationsSection', () => {
it('dates its own cohort so it is not read as stale next to the GCSE section', () => {
render(<DestinationsSection destinations={fullPhase} />);
expect(screen.getByText(/2022\/23/)).toBeInTheDocument();
});
it('renders one bar segment per published category', () => {
const { container } = render(<DestinationsSection destinations={fullPhase} />);
expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(8);
});
it('renders NO bar at all when a category is withheld', () => {
const { container } = render(<DestinationsSection destinations={suppressedPhase} />);
// R1: a bar with a gap in it publishes the withheld figure by its width.
expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(0);
expect(screen.getAllByText(/withheld/i).length).toBeGreaterThan(0);
});
it('never states the remainder for a partially suppressed group', () => {
const { container } = render(<DestinationsSection destinations={suppressedPhase} />);
// 180 cohort - 159 published = 21, the withheld figure. It must appear nowhere.
expect(container.textContent).not.toMatch(/\b21\b/);
});
it('shows a card value for a group whose components are all published', () => {
render(<DestinationsSection destinations={fullPhase} />);
// academic route = 75 + 21 = 96 of 180 = 53%
expect(screen.getByText('53%')).toBeInTheDocument();
});
it('refuses a card value when one of its components is withheld', () => {
render(<DestinationsSection destinations={suppressedPhase} />);
// academic route needs sixth_form_college, which is suppressed.
expect(screen.getByText(/not published/i)).toBeInTheDocument();
expect(screen.queryByText('53%')).not.toBeInTheDocument();
});
it('never claims a pupil stayed at this school', () => {
const { container } = render(<DestinationsSection destinations={fullPhase} />);
// The published file reports destination TYPE, never destination institution.
expect(container.textContent).not.toMatch(/stayed on (here|at this school)/i);
});
it('renders nothing when no group carries categories', () => {
const empty: DestinationPhase = { cohort_year: '2022/23', groups: {} };
const { container } = render(<DestinationsSection destinations={empty} />);
expect(container.firstChild).toBeNull();
});
});
describe('the detail table keeps the three statuses apart', () => {
// 'suppressed' and 'not_applicable' are different claims, and the mart, the
// SQLAlchemy model and the serialiser all preserve the difference. The table
// used to key its Share column off `percentage === null`, which is true for
// both, so a category that simply does not apply was labelled "withheld" —
// while the Pupils column beside it rendered blank.
const mixedPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: [
cell('school_sixth_form', 75),
cell('sixth_form_college', null, 'suppressed'),
cell('further_education', null, 'suppressed'),
cell('apprenticeship', null, 'not_applicable'),
],
},
},
};
const rowFor = (container: HTMLElement, category: string) =>
Array.from(container.querySelectorAll('tbody tr'))
.find(tr => tr.textContent?.includes(category));
it('never labels a not-applicable category as withheld', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'Apprenticeship');
expect(row).toBeTruthy();
expect(row!.textContent).not.toMatch(/withheld/i);
});
it('labels a genuinely suppressed category as withheld in both columns', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'Sixth-form college');
expect(row).toBeTruthy();
expect(row!.querySelectorAll('td')).toHaveLength(2);
Array.from(row!.querySelectorAll('td')).forEach(td =>
expect(td.textContent).toMatch(/withheld/i));
});
it('the two columns of a row never disagree about what the row is', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
Array.from(container.querySelectorAll('tbody tr')).forEach(tr => {
const cells = Array.from(tr.querySelectorAll('td'))
.map(td => /withheld/i.test(td.textContent ?? ''));
expect(new Set(cells).size).toBe(1);
});
});
it('shows a published category its real figures', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'State-funded school sixth form');
expect(row!.textContent).toMatch(/75/);
expect(row!.textContent).toMatch(/42%/);
});
});
@@ -0,0 +1,110 @@
import { render, screen, fireEvent, waitFor } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { FilterBar } from '@/components/FilterBar';
const push = jest.fn();
let searchParams = new URLSearchParams();
jest.mock('next/navigation', () => ({
useRouter: () => ({ push, replace: jest.fn(), prefetch: jest.fn() }),
usePathname: () => '/',
useSearchParams: () => searchParams,
}));
const FILTERS = {
local_authorities: [], school_types: [], years: [], phases: [],
genders: [], admissions_policies: [],
};
const realFetch = global.fetch;
beforeEach(() => {
global.fetch = jest.fn(async () => ({
ok: true,
json: async () => ({ suggestions: [{
urn: 100010, school_name: 'Brecknock Primary School',
local_authority: 'Camden', postcode: 'NW1 1AA',
phase: 'Primary', school_type: 'Community school' }] }),
})) as unknown as typeof fetch;
push.mockClear();
searchParams = new URLSearchParams();
});
afterEach(() => { global.fetch = realFetch; });
describe('FilterBar autosuggest', () => {
it('is a combobox only when the flag is on', () => {
const { rerender } = render(<FilterBar filters={FILTERS} autosuggest={false} />);
expect(screen.queryByRole('combobox')).not.toBeInTheDocument();
rerender(<FilterBar filters={FILTERS} autosuggest />);
expect(screen.getByRole('combobox')).toBeInTheDocument();
});
it('makes no request while the flag is off', async () => {
// Off means off: no listener, no fetch, no markup.
render(<FilterBar filters={FILTERS} autosuggest={false} />);
await userEvent.type(screen.getByPlaceholderText(/School name or postcode/i),
'brecknock');
expect(global.fetch).not.toHaveBeenCalled();
});
it('shows suggestions and navigates when one is chosen', async () => {
render(<FilterBar filters={FILTERS} autosuggest />);
await userEvent.type(screen.getByRole('combobox'), 'brecknock');
const option = await screen.findByRole('option', { name: /Brecknock/ });
await userEvent.click(option);
expect(push).toHaveBeenCalledWith(
expect.stringContaining('/school/100010'));
});
it('suppresses suggestions once the value is a postcode', async () => {
// The box takes a name OR a postcode; suggestions must get out of the way.
//
// fireEvent.change, not userEvent.type: typing sets "N", "NW", "NW1"... and
// "NW1" is not a postcode, so a request for it is correct behaviour. Only
// the settled value is the assertion, so set it in one go.
render(<FilterBar filters={FILTERS} autosuggest />);
fireEvent.change(screen.getByRole('combobox'), { target: { value: 'NW1 1AA' } });
await new Promise((r) => setTimeout(r, 300)); // past the 200ms debounce
expect(global.fetch).not.toHaveBeenCalled();
});
it('Enter with no active option still submits the free-text search', async () => {
// The existing behaviour is preserved, not replaced.
render(<FilterBar filters={FILTERS} autosuggest />);
const input = screen.getByRole('combobox');
await userEvent.type(input, 'brecknock{Enter}');
// updateURL pushes inside startTransition, so the call is not synchronous.
await waitFor(() => expect(push).toHaveBeenCalledWith(
expect.stringContaining('search=brecknock')));
});
});
describe('FilterBar autosuggest does not reopen over results', () => {
it('stays shut when the input arrives pre-filled from the URL', async () => {
/*
* The results-page bar renders with the search term already in the input.
* Opening on that would drop the dropdown on top of the results the search
* just produced — which is exactly what happened: the first result became
* unclickable, because the list sat over it and swallowed the pointer.
*
* Suggestions answer typing, not the presence of a value.
*/
searchParams = new URLSearchParams('search=brecknock');
render(<FilterBar filters={FILTERS} autosuggest />);
expect(screen.getByRole('combobox')).toHaveValue('brecknock');
await new Promise((r) => setTimeout(r, 300)); // past the 200ms debounce
expect(global.fetch).not.toHaveBeenCalled();
expect(screen.queryByRole('listbox')).not.toBeInTheDocument();
});
it('closes the dropdown when the search is submitted', async () => {
render(<FilterBar filters={FILTERS} autosuggest />);
const input = screen.getByRole('combobox');
await userEvent.type(input, 'brecknock');
expect(await screen.findByRole('listbox')).toBeInTheDocument();
await userEvent.type(input, '{Enter}');
await waitFor(() =>
expect(screen.queryByRole('listbox')).not.toBeInTheDocument());
});
});
@@ -0,0 +1,47 @@
/**
* The footer is the only navigational route to /about and /blog, so it is
* where a dark flag would otherwise leave a link into a 404.
*
* Both props default to false. A caller that forgets to pass them hides the
* links, which is the direction that cannot break a page — the same reasoning
* as backend/flags.py's "every flag defaults to False".
*/
import { render, screen } from '@testing-library/react';
import { Footer } from '@/components/Footer';
describe('footer feature links', () => {
it('links to both when both flags are on', () => {
render(<Footer aboutEnabled blogEnabled />);
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toHaveAttribute('href', '/about');
expect(screen.getByRole('link', { name: /^blog$/i }))
.toHaveAttribute('href', '/blog');
});
it('omits the about link when that flag is dark', () => {
render(<Footer blogEnabled />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.getByRole('link', { name: /^blog$/i })).toBeInTheDocument();
});
it('omits the blog link when that flag is dark', () => {
render(<Footer aboutEnabled />);
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toBeInTheDocument();
});
it('drops the whole section when both are dark, not an empty heading', () => {
// Shipping dark means the footer renders as it did before the feature
// existed, not as a section with its contents removed.
render(<Footer />);
expect(screen.queryByRole('heading', { name: /^about$/i })).toBeNull();
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
});
it('defaults to dark when a caller passes nothing', () => {
render(<Footer />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
});
});
@@ -0,0 +1,81 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import { HomeView } from '@/components/HomeView';
import { fetchSchools } from '@/lib/api';
import { primaryFixture } from '../support/schoolFixtures';
import type { SchoolsResponse, School } from '@/lib/types';
let params = new URLSearchParams('postcode=SW1A+1AA');
jest.mock('next/navigation', () => ({
useSearchParams: () => params,
usePathname: () => '/',
useRouter: () => ({ push: jest.fn(), replace: jest.fn() }),
}));
jest.mock('@/context/ComparisonContext', () => ({
useComparisonContext: () => ({ addSchool: jest.fn(), removeSchool: jest.fn(), selectedSchools: [] }),
}));
jest.mock('@/lib/api', () => ({
fetchSchools: jest.fn(),
fetchNationalAverages: jest.fn(async () => ({})),
fetchLAaverages: jest.fn(async () => ({ secondary: { attainment_8_by_la: {} } })),
}));
jest.mock('@/components/FilterBar', () => ({ FilterBar: () => null }));
jest.mock('@/components/SchoolRow', () => ({ SchoolRow: ({ school }: {school: School}) => <div>{school.school_name}</div> }));
jest.mock('@/components/SchoolMap', () => ({ SchoolMap: ({ schools }: {schools: School[]}) => <div data-testid="map">{schools.map(s => s.school_name).join(',')}</div> }));
const filters = { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
function response(name: string): SchoolsResponse {
return { schools: [{ ...primaryFixture.schoolInfo, school_name: name }],
total: 2, page: 1, page_size: 1, total_pages: 2 };
}
function deferred() {
let resolve!: (value: SchoolsResponse) => void;
let reject!: (error: Error) => void;
const promise = new Promise<SchoolsResponse>((yes, no) => { resolve = yes; reject = no; });
return { promise, resolve, reject };
}
beforeEach(() => {
params = new URLSearchParams('postcode=SW1A+1AA');
jest.mocked(fetchSchools).mockReset();
});
test('load-more results from an old search are discarded, even after returning to it', async () => {
const pending = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(pending.promise);
const view = render(<HomeView initialSchools={response('Initial A')} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Load more schools' }));
const signal = jest.mocked(fetchSchools).mock.calls[0][1]?.signal;
params = new URLSearchParams('postcode=SW2+1AA');
view.rerender(<HomeView initialSchools={response('Initial B')} filters={filters} />);
expect(signal?.aborted).toBe(true);
params = new URLSearchParams('postcode=SW1A+1AA');
view.rerender(<HomeView initialSchools={response('Fresh A')} filters={filters} />);
await act(async () => pending.resolve(response('Stale append')));
expect(screen.queryByText('Stale append')).not.toBeInTheDocument();
expect(screen.getByRole('button', { name: 'Load more schools' })).toBeEnabled();
});
test('an older map response cannot overwrite the current search', async () => {
const first = deferred(), second = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(first.promise).mockReturnValueOnce(second.promise);
const initial = response('Initial A');
const view = render(<HomeView initialSchools={initial} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Map' }));
params = new URLSearchParams('postcode=SW2+1AA');
view.rerender(<HomeView initialSchools={response('Initial B')} filters={filters} />);
await act(async () => second.resolve(response('Current map')));
await act(async () => first.resolve(response('Stale map')));
expect(screen.getByTestId('map')).toHaveTextContent('Current map');
expect(screen.getByTestId('map')).not.toHaveTextContent('Stale map');
});
test('failed map requests can be retried by reopening the map', async () => {
const pending = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(pending.promise).mockResolvedValue(response('Retry result'));
render(<HomeView initialSchools={response('Initial')} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Map' }));
await act(async () => pending.reject(new Error('offline')));
fireEvent.click(screen.getByRole('button', { name: 'List' }));
await act(async () => fireEvent.click(screen.getByRole('button', { name: 'Map' })));
expect(fetchSchools).toHaveBeenCalledTimes(2);
expect(screen.getByTestId('map')).toHaveTextContent('Retry result');
});
@@ -0,0 +1,92 @@
/**
* The module that ends the stranding: before it, a school page's only anchor
* pointed at the school's own website, so ~27k pages sent authority off-site
* and none of it reached the location layer.
*/
import { render, screen } from '@testing-library/react';
import { NearbyPlaces } from '@/components/school/NearbyPlaces';
const essex = { kind: 'authority', slug: 'essex', name: 'Essex', count: 480, url: '/schools/authority/essex', phases: [] };
const brentwood = { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 37, url: '/schools/brentwood', phases: [] };
const cm15 = { kind: 'outcode', slug: 'cm15', name: 'CM15', count: 12, url: '/schools/near/cm15', phases: [] };
describe('NearbyPlaces', () => {
it('links to every place the school belongs to', () => {
render(<NearbyPlaces places={[essex, brentwood, cm15]} />);
expect(screen.getByRole('link', { name: /Brentwood/ }))
.toHaveAttribute('href', '/schools/brentwood');
expect(screen.getByRole('link', { name: /Essex/ }))
.toHaveAttribute('href', '/schools/authority/essex');
expect(screen.getByRole('link', { name: /CM15/ }))
.toHaveAttribute('href', '/schools/near/cm15');
});
it('says how many schools each link leads to', () => {
// An anchor that states its destination's size is worth more to a reader
// and to a crawler than "see more".
render(<NearbyPlaces places={[brentwood]} />);
expect(screen.getByRole('link', { name: /37 schools in Brentwood/ }))
.toBeInTheDocument();
});
it('renders nothing at all when the school has no published places', () => {
// Not an empty heading. A school whose town and authority both fall below
// the threshold has nowhere to point, and the page should look as it did
// before the module existed.
const { container } = render(<NearbyPlaces places={[]} />);
expect(container).toBeEmptyDOMElement();
});
it('puts the narrowest place first, which is the most useful link', () => {
// The API orders widest-first for the breadcrumb; a reader on a school
// page wants its town before its county.
render(<NearbyPlaces places={[essex, brentwood, cm15]} />);
const hrefs = screen.getAllByRole('link').map((a) => a.getAttribute('href'));
expect(hrefs.indexOf('/schools/brentwood'))
.toBeLessThan(hrefs.indexOf('/schools/authority/essex'));
});
it('handles a singular count without saying "1 schools"', () => {
render(<NearbyPlaces places={[{ ...brentwood, count: 1 }]} />);
expect(screen.getByRole('link', { name: /1 school in Brentwood/ }))
.toBeInTheDocument();
});
it('links the phase page the school appears on', () => {
// "primary schools in brentwood" is the query these pages exist for.
render(<NearbyPlaces places={[{
...brentwood,
phases: [{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' }],
}]} />);
expect(screen.getByRole('link', { name: /22 primary schools in Brentwood/ }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('links both phase pages for an all-through school', () => {
render(<NearbyPlaces places={[{
...brentwood,
phases: [
{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' },
{ phase: 'secondary', count: 9, url: '/schools/brentwood/secondary' },
],
}]} />);
expect(screen.getByRole('link', { name: /22 primary schools/ })).toBeInTheDocument();
expect(screen.getByRole('link', { name: /9 secondary schools/ })).toBeInTheDocument();
});
it('keeps a phase link next to the place it belongs to', () => {
// Grouping matters: "22 primary schools in Brentwood" directly after
// "37 schools in Brentwood" reads as one place, not two unrelated links.
render(<NearbyPlaces places={[essex, {
...brentwood,
phases: [{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' }],
}]} />);
const hrefs = screen.getAllByRole('link').map((a) => a.getAttribute('href'));
expect(hrefs.indexOf('/schools/brentwood/primary'))
.toBe(hrefs.indexOf('/schools/brentwood') + 1);
});
});
@@ -0,0 +1,192 @@
/**
* The section's job is to be honest about what it is showing. These tests pin
* the ways it could mislead: rendering below the minimum, claiming a likeness
* it does not rank on, showing a missing figure as a number, or hiding a card
* behind an arrow where a crawler cannot reach it.
*/
import { render, screen } from '@testing-library/react';
import {
nearbyNoun,
NearbySchoolsSection,
shouldRenderNearby,
} from '@/components/school/NearbySchoolsSection';
import type { NearbySchool } from '@/lib/types';
jest.mock('@/components/school/AddToCompareButton', () => ({
AddToCompareButton: ({ school }: { school: NearbySchool }) => (
<button type="button">Add {school.school_name} to compare</button>
),
}));
jest.mock('@/components/school/NearbySchoolsCompareBar', () => ({
NearbySchoolsCompareBar: ({ thisUrn }: { thisUrn: number }) => (
<div data-testid="compare-bar">bar for {thisUrn}</div>
),
}));
function school(overrides: Partial<NearbySchool> = {}): NearbySchool {
return {
urn: 100002,
school_name: 'Willow Lane Primary School',
distance_miles: 0.6,
school_type: 'Community school',
age_range: '4-11',
shared: ['Mixed', 'No religious character'],
metric_value: 74,
metric_key: 'rwm_expected_pct',
metric_year: 202425,
...overrides,
};
}
function renderSection(nearby: NearbySchool[]) {
return render(
<NearbySchoolsSection
urn={100001}
schoolName="Meadowbrook Primary School"
phase="Primary"
thisMetricValue={72}
nearby={nearby}
/>,
);
}
describe('render gates', () => {
it.each([
['undefined', undefined],
['null', null],
['empty', []],
['a single school', [school()]],
])('renders nothing for %s', (_label, value) => {
expect(shouldRenderNearby(value as NearbySchool[] | null | undefined)).toBe(false);
});
it('renders for two or more schools', () => {
expect(shouldRenderNearby([school(), school({ urn: 100003 })])).toBe(true);
});
it('returns null rather than an empty shell below the minimum', () => {
const { container } = renderSection([school()]);
expect(container).toBeEmptyDOMElement();
});
});
describe('what the section claims', () => {
it('never claims a similar intake, because it does not rank on one', () => {
renderSection([school(), school({ urn: 100003, shared: [] })]);
expect(screen.queryByText(/similar intake/i)).not.toBeInTheDocument();
});
it('is headed "Other schools nearby", not "similar"', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByRole('heading', { name: 'Other schools nearby' })).toBeInTheDocument();
});
it('shows chips for what is shared', () => {
renderSection([school({ shared: ['Mixed', 'Roman Catholic'] }), school({ urn: 100003 })]);
expect(screen.getAllByText('Roman Catholic').length).toBe(1);
});
it('shows no chips at all when nothing is shared, rather than inventing one', () => {
const { container } = render(
<NearbySchoolsSection
urn={100001}
schoolName="Meadowbrook Primary School"
phase="Primary"
thisMetricValue={72}
nearby={[school({ shared: [] }), school({ urn: 100003, shared: [] })]}
/>,
);
// The card still carries its distance, name, type and figure — just no
// claim of likeness.
expect(container.querySelectorAll('li ul').length).toBe(0);
expect(screen.getAllByText(/miles away/).length).toBe(2);
});
});
describe('what the lede calls the set', () => {
it.each([
['Primary', 'primary schools'],
['Middle deemed primary', 'primary schools'],
['Secondary', 'secondary schools'],
['Middle deemed secondary', 'secondary schools'],
['All-through', 'all-through schools'],
// GIAS phase 6. Its candidates span the whole secondary group, so no
// single noun fits and it takes the honest general one.
['16 plus', 'schools and colleges'],
['', 'schools'],
[null, 'schools'],
])('calls a %s school\'s neighbours "%s"', (phase, expected) => {
expect(nearbyNoun(phase)).toBe(expected);
});
it('never calls a sixth form college\'s neighbours primary schools', () => {
render(
<NearbySchoolsSection
urn={100001}
schoolName="Barnet Sixth Form College"
phase="16 plus"
thisMetricValue={null}
nearby={[school(), school({ urn: 100003 })]}
/>,
);
expect(screen.getByText(/Other schools and colleges near Barnet Sixth Form College/)).toBeInTheDocument();
expect(screen.queryByText(/primary schools/)).not.toBeInTheDocument();
});
});
describe('cards', () => {
it('links each school to its canonical slug', () => {
renderSection([school(), school({ urn: 100003, school_name: 'Oakfield Primary School' })]);
const link = screen.getByRole('link', { name: /Willow Lane Primary School/ });
expect(link).toHaveAttribute('href', '/school/100002-willow-lane-primary-school');
});
it('shows the distance and the shared characteristics', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getAllByText('0.6 miles away').length).toBeGreaterThan(0);
expect(screen.getAllByText('Mixed').length).toBeGreaterThan(0);
});
it('renders a missing figure as "Not published", never as a number', () => {
renderSection([school({ metric_value: null }), school({ urn: 100003 })]);
expect(screen.getByText('Not published')).toBeInTheDocument();
});
it('anchors each figure against this school', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getAllByText('72% at this school').length).toBe(2);
});
it('offers the compare bar once, for this school', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByTestId('compare-bar')).toHaveTextContent('bar for 100001');
});
it('keeps every card in the DOM, including the ones scrolled out of view', () => {
const six = Array.from({ length: 6 }, (_, n) =>
school({ urn: 100002 + n, school_name: `Peer ${n} School` }),
);
renderSection(six);
expect(screen.getAllByRole('link', { name: /Peer \d School/ })).toHaveLength(6);
});
it('offers no arrows when three cards fit the row', () => {
renderSection([school(), school({ urn: 100003 }), school({ urn: 100004 })]);
expect(screen.queryByRole('button', { name: /More schools/ })).not.toBeInTheDocument();
});
it('offers arrows once there is a fourth school', () => {
renderSection(Array.from({ length: 4 }, (_, n) => school({ urn: 100002 + n })));
expect(screen.getByRole('button', { name: /More schools/ })).toBeInTheDocument();
expect(screen.getByRole('button', { name: /Previous schools/ })).toBeInTheDocument();
});
it('says distances are straight-line, and offers no method panel', () => {
const { container } = renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByText(/straight-line from this school/i)).toBeInTheDocument();
expect(container.querySelector('details')).toBeNull();
});
});
@@ -0,0 +1,479 @@
import { render, screen } from '@testing-library/react';
import { PlaceView } from '@/components/places/PlaceView';
import type { PlaceDetail } from '@/lib/places';
const detail: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 29,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', rwm_expected_pct: 82,
ofsted_grade: 1, phase: 'Primary' } as never,
{ urn: 2, school_name: 'Beta Primary', rwm_expected_pct: 44,
ofsted_grade: 3, phase: 'Primary' } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
describe('PlaceView', () => {
it('leads with an H1 that matches how the place is searched', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByRole('heading', { level: 1 }))
.toHaveTextContent(/primary schools in brentwood/i);
});
it('states the count so the page says something before the table', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByText(/29 schools/i)).toBeInTheDocument();
});
it('compares the local average against England, which a list cannot', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByTestId('local-vs-england')).toHaveTextContent('63');
expect(screen.getByTestId('local-vs-england')).toHaveTextContent('61');
});
it('links every school in scope, which is what de-orphans them', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getAllByRole('link', { name: /Primary$/ })).toHaveLength(2);
});
it('links to the parent authority so the place sits in a hierarchy', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByRole('link', { name: /Essex/i }))
.toHaveAttribute('href', '/schools/authority/essex');
});
it('shows the Ofsted distribution, not just a count of Outstanding', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByTestId('ofsted-distribution')).toBeInTheDocument();
});
it('links to neighbouring places so the page is not a dead end', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[{ kind: 'town', slug: 'romford', name: 'Romford', count: 40 }]} />);
expect(screen.getByRole('link', { name: /Romford/ }))
.toHaveAttribute('href', '/schools/romford');
});
it('says nothing about an average it does not have', () => {
render(<PlaceView detail={{ ...detail, averages:
{ rwm_expected_pct: null, attainment_8_score: null } }}
phase="primary" englandAverage={61} neighbours={[]} />);
expect(screen.queryByTestId('local-vs-england')).not.toBeInTheDocument();
});
});
describe('PlaceView structured data', () => {
function jsonLd() {
const { container } = render(<PlaceView detail={detail} phase="primary"
englandAverage={61} neighbours={[]} />);
const el = container.querySelector('script[type="application/ld+json"]');
return JSON.parse(el!.textContent!);
}
it('declares the page as a ranked list, not prose', () => {
const types = jsonLd()['@graph'].map((n: { '@type': string }) => n['@type']);
expect(types).toContain('ItemList');
expect(types).toContain('BreadcrumbList');
});
it('gives every listed school an absolute URL on the canonical host', () => {
const list = jsonLd()['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListElement).toHaveLength(2);
for (const item of list.itemListElement) {
expect(item.url).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\/school\//);
}
});
});
describe('PlaceView phase variants', () => {
it('links the phase variants that exist', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /Primary schools in Brentwood/i }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('links no variant for a phase below its own threshold', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.queryByRole('link', { name: /Secondary schools in Brentwood/i }))
.not.toBeInTheDocument();
});
it('does not link sideways from a variant page to itself', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.queryByRole('link', { name: /Primary schools in Brentwood/i }))
.not.toBeInTheDocument();
});
});
describe('PlaceView presentation', () => {
// /schools/brentwood shipped with 8 of 27 rows blank: an unphased page shows
// one primary-only measure for a list that also holds secondaries.
const mixed: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 4,
parent_authority: 'Essex', phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta High', phase: 'Secondary',
rwm_expected_pct: null, attainment_8_score: 47 } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: 45 },
};
it('gives each phase its own table rather than one column of blanks', () => {
render(<PlaceView detail={mixed} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('heading', { name: /^Primary schools/ })).toBeInTheDocument();
expect(screen.getByRole('heading', { name: /^Secondary schools/ })).toBeInTheDocument();
expect(screen.getByText('82%')).toBeInTheDocument();
expect(screen.getByText('47')).toBeInTheDocument();
});
it('names the measure in plain words, not jargon', () => {
// The first cut said "RWM expected", which appears nowhere else on the site.
render(<PlaceView detail={mixed} englandAverage={61} neighbours={[]} />);
expect(screen.getByText('Reading, writing & maths')).toBeInTheDocument();
expect(screen.getByText('Attainment 8')).toBeInTheDocument();
expect(screen.queryByText(/RWM expected/i)).not.toBeInTheDocument();
});
it('says a missing result is unpublished rather than showing a bare dash', () => {
const noResult: PlaceDetail = {
...mixed,
schools: [{ urn: 3, school_name: 'New Primary', phase: 'Primary',
rwm_expected_pct: null, attainment_8_score: null } as never],
};
render(<PlaceView detail={noResult} englandAverage={61} neighbours={[]} />);
expect(screen.getByText('Not published')).toBeInTheDocument();
});
it('styles school links to the site convention rather than browser default', () => {
const { container } = render(<PlaceView detail={mixed} englandAverage={61}
neighbours={[]} />);
const link = container.querySelector('a[href^="/school/"]');
expect(link?.className).toBeTruthy();
});
it('a phased page shows one table and no phase headings', () => {
render(<PlaceView detail={mixed} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.queryByRole('heading', { name: /^Secondary schools/ }))
.not.toBeInTheDocument();
});
});
describe('PlaceView table alignment', () => {
const aligned: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('aligns the measure heading and its values with the same class', () => {
// They were aligned by two different selectors whose specificity did not
// match: `.table th:last-child` (0,2,1) won and went right, while `.num`
// (0,1,0) lost to `.table td` (0,1,1) and stayed left. Sharing one class
// is what makes them impossible to drift apart.
const { container } = render(<PlaceView detail={aligned} englandAverage={61}
neighbours={[]} />);
const th = container.querySelectorAll('th')[1];
const td = container.querySelectorAll('tbody td')[1];
expect(th.className).toBeTruthy();
expect(td.className).toBe(th.className);
});
it('leaves the school-name column unclassed so it takes the spare width', () => {
const { container } = render(<PlaceView detail={aligned} englandAverage={61}
neighbours={[]} />);
expect(container.querySelectorAll('th')[0].className).toBe('');
});
});
describe('PlaceView authorities', () => {
const straddling: PlaceDetail = {
place: { kind: 'outcode', slug: 'sw19', name: 'SW19', count: 33,
parent_authority: 'Merton', phases: ['primary'],
authorities: [
{ name: 'Merton', slug: 'merton', count: 26 },
{ name: 'Wandsworth', slug: 'wandsworth', count: 7 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names every authority the place straddles, not just the largest', () => {
// SW19 is mostly Merton but partly Wandsworth. Naming one asserts
// something false about a quarter of outcodes.
render(<PlaceView detail={straddling} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' }))
.toHaveAttribute('href', '/schools/authority/merton');
expect(screen.getByRole('link', { name: 'Wandsworth' }))
.toHaveAttribute('href', '/schools/authority/wandsworth');
});
it('joins them readably rather than as a bare list', () => {
// Asserted on the summary line's whole text: a loose /and/ matcher also
// hits "Wandsworth".
const { container } = render(<PlaceView detail={straddling}
englandAverage={61} neighbours={[]} />);
const summary = container.querySelector('header p');
expect(summary?.textContent).toContain('Merton and Wandsworth');
});
it('falls back to the single parent when the field is absent', () => {
// A cached API response predating the authorities field must not blank
// the line entirely.
const legacy = { ...straddling,
place: { ...straddling.place, authorities: undefined } };
render(<PlaceView detail={legacy} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' })).toBeInTheDocument();
});
});
describe('PlaceView list ordering', () => {
const detail3: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 40, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta Primary', phase: 'Primary',
rwm_expected_pct: 90, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 65, attainment_8_score: null },
};
it('renders schools in the order the API sent them, not by score', () => {
// The API sorts alphabetically now; the component must not re-sort.
render(<PlaceView detail={detail3} englandAverage={61} neighbours={[]} />);
const links = screen.getAllByRole('link', { name: /Primary$/ });
expect(links.map((l) => l.textContent))
.toEqual(['Alpha Primary', 'Beta Primary']);
});
it('declares the list as ascending rather than implying a ranking', () => {
// An ItemList carrying `position` reads as a ranking unless it says
// otherwise, and the table is A-Z.
const { container } = render(<PlaceView detail={detail3} englandAverage={61}
neighbours={[]} />);
const ld = JSON.parse(
container.querySelector('script[type="application/ld+json"]')!.textContent!);
const list = ld['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListOrder).toBe('https://schema.org/ItemListOrderAscending');
});
});
describe('PlaceView phase links', () => {
const authority: PlaceDetail = {
place: { kind: 'authority', slug: 'barnet', name: 'Barnet', count: 156,
parent_authority: null, phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('keeps an authority phase link in the authority namespace', () => {
// The link was built as `/schools/${slug}/${phase}` for every kind, so an
// authority page pointed into the town namespace. For 87 of 151
// authorities that 404'd; for the other 64 it silently landed on the town
// page of the same name — a different set of schools, and exactly the
// duplicate the two namespaces exist to prevent. Barnet is one of the 64.
render(<PlaceView detail={authority} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /^Primary schools in Barnet$/ }))
.toHaveAttribute('href', '/schools/authority/barnet/primary');
expect(screen.getByRole('link', { name: /^Secondary schools in Barnet$/ }))
.toHaveAttribute('href', '/schools/authority/barnet/secondary');
});
it('still uses the bare namespace for a town', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /^Primary schools in Brentwood$/ }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('offers no phase link when the place publishes none', () => {
// Outcodes are the case: no phase route exists for them, so the registry
// reports no phases and the nav does not render.
const outcode = { ...detail,
place: { ...detail.place, kind: 'outcode', slug: 'cm13', name: 'CM13',
phases: [] } };
render(<PlaceView detail={outcode} englandAverage={61} neighbours={[]} />);
expect(screen.queryByRole('navigation', { name: 'By phase' }))
.not.toBeInTheDocument();
});
});
describe('PlaceView unlinkable authorities', () => {
const withUnpublished: PlaceDetail = {
place: { kind: 'outcode', slug: 'tr21', name: 'TR21', count: 8,
parent_authority: 'Cornwall', phases: [],
authorities: [
{ name: 'Cornwall', slug: 'cornwall', count: 6 },
{ name: 'Isles Of Scilly', slug: null, count: 2 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names an authority with no page without linking it', () => {
// City of London and the Isles of Scilly hold fewer schools than a page
// needs. Saying where the place is stays right; linking there would 404.
const { container } = render(<PlaceView detail={withUnpublished}
englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Cornwall' })).toBeInTheDocument();
expect(screen.queryByRole('link', { name: 'Isles Of Scilly' }))
.not.toBeInTheDocument();
expect(container.querySelector('header p')?.textContent)
.toContain('Isles Of Scilly');
});
});
describe('PlaceView school attributes', () => {
/*
* The table shipped with one column of scores, which answers "how did they
* do" and nothing about whether the school is one a family could use. Age
* range, faith, nursery and constituency are the four facts a parent
* filters on before they look at a number at all.
*/
const withAttributes: PlaceDetail = {
place: { kind: 'town', slug: 'chelmsford', name: 'Chelmsford', count: 3,
parent_authority: 'Essex', phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null,
age_range: '4-11', religious_denomination: 'Church of England',
nursery_provision: true,
parliamentary_constituency: 'Chelmsford' } as never,
{ urn: 2, school_name: 'Beta High', phase: 'Secondary',
rwm_expected_pct: null, attainment_8_score: 47,
age_range: '11-16', religious_denomination: 'Does not apply',
nursery_provision: false,
parliamentary_constituency: 'Witham' } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: 45 },
};
function headings(container: HTMLElement, table = 0): string[] {
return Array.from(container.querySelectorAll('table')[table]
.querySelectorAll('thead th')).map((th) => th.textContent ?? '');
}
it('heads a primary table with all four attributes', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
expect(headings(container)).toEqual([
'School', 'Reading, writing & maths',
'Ages', 'Religious character', 'Nursery', 'Constituency',
]);
});
it('omits nursery from a secondary table, where it does not apply', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
expect(headings(container, 1)).toEqual([
'School', 'Attainment 8', 'Ages', 'Religious character', 'Constituency',
]);
});
it('keeps the measure beside the school name, where a phone can see it', () => {
// Six columns overflow a phone and .tableWrap turns that into a swipe.
// With the measure last, the one number the page exists for is the one
// scrolled off the screen.
const { container } = render(<PlaceView detail={withAttributes}
phase="primary" englandAverage={61} neighbours={[]} />);
expect(headings(container)[1]).toBe('Reading, writing & maths');
});
it('shows the age range without repeating the column heading', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('4–11')).toBeInTheDocument();
expect(screen.queryByText('Ages 4–11')).not.toBeInTheDocument();
});
it('names the faith of a faith school', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('Church of England')).toBeInTheDocument();
});
it('reads "Does not apply" as no religious character, not as a value', () => {
// GIAS spells the absence of a faith as "Does not apply", which is a
// database answer rather than an English one. The school page already
// suppresses it; the two must not disagree about the same school.
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
const secondary = container.querySelectorAll('table')[1]
.querySelectorAll('tbody td');
expect(secondary[3].textContent).toBe('—');
expect(screen.queryByText(/Does not apply/)).not.toBeInTheDocument();
});
it('marks a nursery as such and a school without one as not', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
const cells = container.querySelectorAll('table')[0]
.querySelectorAll('tbody td');
expect(cells[4].textContent).toBe('Yes');
});
it('names the constituency of each school', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('Chelmsford', { selector: 'td' })).toBeInTheDocument();
expect(screen.getByText('Witham', { selector: 'td' })).toBeInTheDocument();
});
it('dashes an attribute the data does not carry', () => {
// nursery_provision and parliamentary_constituency are absent from marts
// the pipeline has not rebuilt, and the API degrades them to null rather
// than failing. A row must survive that.
const bare: PlaceDetail = {
...withAttributes,
schools: [{ urn: 3, school_name: 'Gamma Primary', phase: 'Primary',
rwm_expected_pct: 70 } as never],
};
const { container } = render(<PlaceView detail={bare} phase="primary"
englandAverage={61} neighbours={[]} />);
const cells = Array.from(container.querySelectorAll('tbody td'))
.map((td) => td.textContent);
expect(cells.slice(2)).toEqual(['—', '—', '—', '—']);
});
it('gives an all-through school its nursery under primary only', () => {
// All-through schools render in both groups. Nursery belongs to the
// primary reading of the same school, not the secondary one.
const allThrough: PlaceDetail = {
...withAttributes,
schools: [{ urn: 4, school_name: 'Delta Academy', phase: 'All-through',
rwm_expected_pct: 66, attainment_8_score: 51,
age_range: '4-18', religious_denomination: 'None',
nursery_provision: true,
parliamentary_constituency: 'Chelmsford' } as never],
};
const { container } = render(<PlaceView detail={allThrough}
englandAverage={61} neighbours={[]} />);
const tables = container.querySelectorAll('table');
expect(tables[0].textContent).toContain('Yes');
expect(tables[1].textContent).not.toContain('Yes');
});
});
@@ -0,0 +1,44 @@
import { render, screen } from '@testing-library/react';
import { Post16DestinationsSection } from '@/components/school/Post16DestinationsSection';
import type { DestinationPhase } from '@/lib/types';
const phase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 96,
categories: [
{ category: 'higher_education', pupils: 56, percentage: 58.3, status: 'published' },
{ category: 'further_education', pupils: 12, percentage: 12.5, status: 'published' },
{ category: 'apprenticeship', pupils: 9, percentage: 9.4, status: 'published' },
{ category: 'employment', pupils: 13, percentage: 13.5, status: 'published' },
{ category: 'not_sustained', pupils: 6, percentage: 6.3, status: 'published' },
],
},
},
};
describe('Post16DestinationsSection', () => {
it('names the Year 13 cohort, not Year 11', () => {
const { container } = render(<Post16DestinationsSection destinations={phase} />);
expect(container.textContent).toMatch(/Year 13/);
expect(container.textContent).not.toMatch(/Year 11/);
});
it('reports higher education destinations', () => {
render(<Post16DestinationsSection destinations={phase} />);
expect(screen.getByText(/UK higher education/i)).toBeInTheDocument();
});
it('uses its own anchor so the nav does not collide with After Year 11', () => {
const { container } = render(<Post16DestinationsSection destinations={phase} />);
expect(container.querySelector('#post16-destinations')).toBeTruthy();
expect(container.querySelector('#destinations')).toBeNull();
});
it('renders nothing when no group carries categories', () => {
const empty: DestinationPhase = { cohort_year: '2022/23', groups: {} };
const { container } = render(<Post16DestinationsSection destinations={empty} />);
expect(container.firstChild).toBeNull();
});
});
@@ -0,0 +1,38 @@
/**
* The trail has to be written by something, and it has to be written on every
* route — not only the ones that happen to track an event.
*/
import { render } from '@testing-library/react';
const recordVisitedPath = jest.fn();
let pathname = '/schools/brentwood';
jest.mock('next/navigation', () => ({ usePathname: () => pathname }));
jest.mock('@/lib/analytics', () => ({
recordVisitedPath: (p: string) => recordVisitedPath(p),
}));
// eslint-disable-next-line @typescript-eslint/no-var-requires
const { RouteTrail } = require('@/components/RouteTrail');
describe('RouteTrail', () => {
beforeEach(() => recordVisitedPath.mockClear());
it('records the page it is mounted on', () => {
render(<RouteTrail />);
expect(recordVisitedPath).toHaveBeenCalledWith('/schools/brentwood');
});
it('records each new route as the user moves through the app', () => {
const { rerender } = render(<RouteTrail />);
pathname = '/school/115429-brentwood-school';
rerender(<RouteTrail />);
expect(recordVisitedPath).toHaveBeenLastCalledWith(
'/school/115429-brentwood-school');
});
it('renders nothing, so it can sit anywhere in the layout', () => {
const { container } = render(<RouteTrail />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -0,0 +1,50 @@
import { render, screen } from '@testing-library/react';
import { SuggestList, suggestOptionId } from '@/components/SuggestList';
const ROWS = [
{ urn: 1, school_name: "St Mary's Primary", local_authority: 'Camden',
postcode: 'NW1 1AA', phase: 'Primary', school_type: 'Voluntary aided school' },
{ urn: 2, school_name: "St Mary's Primary", local_authority: 'Barnet',
postcode: 'EN5 2AA', phase: 'Primary', school_type: 'Community school' },
];
describe('SuggestList', () => {
it('is a listbox of options', () => {
render(<SuggestList id="s" suggestions={ROWS} activeIndex={-1}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getByRole('listbox')).toBeInTheDocument();
expect(screen.getAllByRole('option')).toHaveLength(2);
});
it('shows the local authority, which is what tells two schools apart', () => {
// Both rows are "St Mary's Primary". Without the authority the list is
// unusable for exactly the query autosuggest exists to serve.
render(<SuggestList id="s" suggestions={ROWS} activeIndex={-1}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getByText('Camden')).toBeInTheDocument();
expect(screen.getByText('Barnet')).toBeInTheDocument();
});
it('marks only the active option selected', () => {
render(<SuggestList id="s" suggestions={ROWS} activeIndex={1}
onPick={() => {}} onHover={() => {}} />);
const options = screen.getAllByRole('option');
expect(options[0]).toHaveAttribute('aria-selected', 'false');
expect(options[1]).toHaveAttribute('aria-selected', 'true');
});
it('gives each option the id the input will point at', () => {
// aria-activedescendant on the input has to name a real element id, or
// a screen reader announces nothing as the user arrows through.
render(<SuggestList id="s" suggestions={ROWS} activeIndex={0}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getAllByRole('option')[0]).toHaveAttribute(
'id', suggestOptionId('s', 0));
});
it('renders nothing when there is nothing to suggest', () => {
const { container } = render(<SuggestList id="s" suggestions={[]}
activeIndex={-1} onPick={() => {}} onHover={() => {}} />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -0,0 +1,45 @@
import { render } from '@testing-library/react';
import { TrackPlaceView } from '@/components/places/TrackPlaceView';
const trackMock = jest.fn();
jest.mock('@/lib/analytics', () => ({
track: (...args: unknown[]) => trackMock(...args),
getNavigationSource: () => 'search',
}));
describe('TrackPlaceView', () => {
beforeEach(() => trackMock.mockClear());
it('reports which kind of location page was viewed', () => {
/*
* `kind` is the reason this event exists. Whether to keep investing in the
* location layer turns on which *sort* of page earns engagement — towns,
* authorities or postcode districts — and a bare pageview cannot say,
* because all four families share the /schools/ prefix.
*/
render(<TrackPlaceView kind="authority" slug="kent" count={412} />);
expect(trackMock).toHaveBeenCalledWith('place_viewed', {
kind: 'authority', slug: 'kent', phase: 'all',
school_count: 412, from: 'search',
});
});
it('names the phase when the page is a phase variant', () => {
render(<TrackPlaceView kind="town" slug="brentwood" count={29} phase="primary" />);
expect(trackMock).toHaveBeenCalledWith('place_viewed',
expect.objectContaining({ phase: 'primary' }));
});
it('fires once, not once per render', () => {
const { rerender } = render(
<TrackPlaceView kind="town" slug="brentwood" count={29} />);
rerender(<TrackPlaceView kind="town" slug="brentwood" count={29} />);
expect(trackMock).toHaveBeenCalledTimes(1);
});
it('renders nothing', () => {
const { container } = render(
<TrackPlaceView kind="town" slug="brentwood" count={29} />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -0,0 +1,192 @@
import fs from 'fs';
import path from 'path';
/**
* Guards against light-theme-only CSS.
*
* The site themes entirely through tokens redefined under
* `@media (prefers-color-scheme: dark)`. A hardcoded colour therefore does not
* fail loudly — it renders perfectly in the theme it was written for and
* quietly wrongly in the other, which nobody sees unless they happen to be in
* dark mode when they look.
*
* Both rules below are drawn from real defects in SchoolHeroMap.module.css,
* found by eye rather than by any test:
*
* - the map's fade to the header ramped through hardcoded white and landed on
* `var(--bg-card)`. Invisible in light; a bright band across the full width
* of a near-black card in dark.
* - the controls floating over the map paired a hardcoded white background
* with `color: var(--text-primary)`, which resolves to #E9EEF0 in dark —
* near-white text on a near-white button.
*/
const COMPONENTS = path.join(__dirname, '..', '..', 'components');
function stylesheets(dir: string): string[] {
return fs.readdirSync(dir, { withFileTypes: true }).flatMap((entry) => {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) return stylesheets(full);
return entry.name.endsWith('.module.css') ? [full] : [];
});
}
/** Innermost `selector { body }` pairs. Nested at-rules never match as rules,
* because their body contains braces.
*
* Comments are stripped before matching rather than after, so that the whole
* selector survives. Taking only its last line — which is what stripping a
* leading comment used to require — silently discarded every selector in a
* grouped rule but the final one, and a safety guard that cannot see half its
* input fails open. */
function rules(css: string): Array<{ selector: string; body: string }> {
const bare = css.replace(/\/\*[\s\S]*?\*\//g, '');
return Array.from(bare.matchAll(/([^{}]+)\{([^{}]*)\}/g), (m) => ({
selector: m[1].trim().replace(/\s*\n\s*/g, ' '),
body: m[2],
}));
}
const HARDCODED_WHITE_BG = /background[^;]*(?:255,\s*255,\s*255|#fff\b|#ffffff\b)/i;
const THEMED_COLOR = /(?:^|[^-])color:\s*var\(--/;
const files = stylesheets(COMPONENTS);
/** Component sources, for the third-party-surface rule below. */
function sources(dir: string): string[] {
return fs.readdirSync(dir, { withFileTypes: true }).flatMap((entry) => {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) return sources(full);
return entry.name.endsWith('.tsx') ? [full] : [];
});
}
describe('dark-theme safety', () => {
it('finds stylesheets to check', () => {
expect(files.length).toBeGreaterThan(0);
});
it('never pairs a hardcoded white background with a themed text colour', () => {
const offenders = files.flatMap((file) =>
rules(fs.readFileSync(file, 'utf8'))
.filter((r) => HARDCODED_WHITE_BG.test(r.body) && THEMED_COLOR.test(r.body))
.map((r) => `${path.relative(COMPONENTS, file)} ${r.selector}`));
// Either the surface follows the theme and so should the text, or it does
// not and the text must be literal too. Mixing them is how near-white text
// ends up on a near-white button.
expect(offenders).toEqual([]);
});
it('never fades to a themed colour through a hardcoded one', () => {
const offenders = files.flatMap((file) =>
rules(fs.readFileSync(file, 'utf8'))
.filter((r) => /linear-gradient/.test(r.body)
&& /var\(--bg-(card|primary|secondary)\)/.test(r.body)
&& /255,\s*255,\s*255|#fff\b/i.test(r.body))
.map((r) => `${path.relative(COMPONENTS, file)} ${r.selector}`));
// A gradient that lands on a token has to be made of that token, or the
// ramp and its destination disagree in one theme. Use the matching
// `--*-rgb` token for the transparent stops.
expect(offenders).toEqual([]);
});
});
/**
* The same defect one stylesheet further out.
*
* The rules above scan our own CSS modules. They cannot see a surface painted
* by a third-party sheet: leaflet.css hardcodes `background: white` on
* `.leaflet-popup-content-wrapper` and `.leaflet-popup-tip`, and
* LeafletMapInner builds its popup as an HTML string with inline
* `color: var(--text-primary)`. Neither half lives in a .module.css, so the
* module scan passed while dark mode rendered #E9EEF0 on #FFFFFF — 1.17:1,
* with the school name and the headline figure effectively invisible.
*
* globals.css already pulls the rest of Leaflet's chrome onto the tokens (the
* attribution bar, the zoom controls) for exactly this reason. The popup was
* simply missed.
*/
describe('third-party surfaces under themed text', () => {
const GLOBALS = path.join(__dirname, '..', '..', 'app', '(frontend)', 'globals.css');
/** Leaflet surfaces our own code writes token-coloured text onto. */
const LEAFLET_POPUP_SURFACES = [
'.leaflet-popup-content-wrapper',
'.leaflet-popup-tip',
];
it('still finds a component painting themed text into a Leaflet popup', () => {
// Guards the rule below against passing vacuously if the popups are ever
// rewritten as React components rather than HTML strings.
const themed = sources(COMPONENTS).filter((file) => {
const src = fs.readFileSync(file, 'utf8');
return /bindPopup\(/.test(src) && /color:var\(--|color: var\(--/.test(src);
});
expect(themed.length).toBeGreaterThan(0);
});
it('themes the Leaflet popup surface, because the text on it is themed', () => {
const globals = rules(fs.readFileSync(GLOBALS, 'utf8'));
const unthemed = LEAFLET_POPUP_SURFACES.filter((surface) => {
const rule = globals.find((r) => r.selector.includes(surface));
return !rule || !/background[^;]*var\(--/.test(rule.body);
});
// Leaflet's white is not a colour this site owns. Either the surface
// follows the theme or the text on it must be literal — and the text is
// already themed.
expect(unthemed).toEqual([]);
});
it('never puts a literal white label on a themed fill', () => {
/*
* The mirror image of the module-CSS rule above, and the half of the popup
* that theming the card does not reach. "View Details" is
* `background:var(--status-above);color:white`; --status-above is #36743F
* in light but #7FCB8A in dark, so the label went from 5.63:1 to 1.94:1.
*
* --text-inverse is the token for ink on a saturated fill — #FFFFFF in
* light, #111A20 in dark — and the popup's Ofsted badge already uses it.
*/
const offenders = sources(COMPONENTS).flatMap((file) => {
const src = fs.readFileSync(file, 'utf8');
return Array.from(
src.matchAll(/background:\s*var\(--[^;"']*;[^"']*?color:\s*(white|#fff\b|#ffffff\b)/gi),
() => path.relative(COMPONENTS, file));
});
expect(offenders).toEqual([]);
});
});
/**
* Destination measures add the first new colour family since the palette was
* set. The tokens have to exist in both blocks or the section renders one
* theme's fills on the other theme's ground — the exact failure the suite
* above exists to catch, but for tokens rather than literals.
*/
describe('destination tokens', () => {
const css = fs.readFileSync(
path.join(__dirname, '..', '..', 'app', '(frontend)', 'globals.css'), 'utf8');
const TOKENS = [
'--dest-sixthform', '--dest-sfcollege', '--dest-fecollege',
'--dest-apprentice', '--dest-employment', '--dest-none', '--dest-none-hatch',
];
const DARK_AT = css.indexOf('@media (prefers-color-scheme: dark)');
it('defines every destination token in the light palette', () => {
const light = css.slice(0, DARK_AT);
expect(TOKENS.filter((t) => !light.includes(`${t}:`))).toEqual([]);
});
it('redefines every destination token for dark', () => {
const dark = css.slice(DARK_AT);
expect(TOKENS.filter((t) => !dark.includes(`${t}:`))).toEqual([]);
});
});
@@ -0,0 +1,108 @@
import fs from 'fs';
import path from 'path';
/**
* The hero search and the results filter bar are the same component in two
* costumes. `.filterBar` is the card — background, border, shadow, padding —
* and `.heroMode` strips all of it so the search sits directly on the hero
* panel.
*
* Both selectors have specificity (0,1,0), so **source order decides**, and
* `.heroMode` only wins because it is declared immediately after. Any later
* bare `.filterBar` rule — which in practice means one inside a media query —
* silently wins instead, and the hero grows a card's padding back.
*
* That is exactly what happened: `@media (max-width: 768px) { .filterBar {
* padding: 0.875rem } }` re-added 14px in hero mode, indenting the search box,
* the hint and the location link 14px past the headline above them and costing
* the search field 28px of width on a 390px screen. The two rules directly
* below it in the same block were correctly written as
* `.filterBar:not(.heroMode)`; this one was missed, and nothing caught it
* because the result is a plausible-looking layout rather than a broken one.
*/
const CSS = path.join(__dirname, '..', '..', 'components', 'FilterBar.module.css');
/** Properties `.heroMode` resets. A later bare `.filterBar` rule setting any
* of these puts the card back on the hero. */
const RESET_BY_HERO_MODE = [
'background', 'border', 'border-radius', 'box-shadow', 'padding',
];
/**
* Comments are stripped before anything is parsed.
*
* A `{` or `}` inside a comment would otherwise desynchronise the brace walk
* below and the rule regex alike, and the selector text captured for each rule
* would carry the preceding comment along with it.
*/
function withoutComments(css: string): string {
return css.replace(/\/\*[\s\S]*?\*\//g, '');
}
/**
* The individual selectors in a rule's prelude.
*
* Split on commas, because a selector list is a list: `.filterBar, .other { }`
* applies to `.filterBar` just as surely as `.filterBar { }` does, and an
* earlier version of this guard compared the whole prelude against the literal
* string '.filterBar' — so writing the regression as a comma list, or across
* two lines, would have walked straight past it.
*/
function selectorsOf(prelude: string): string[] {
return prelude.split(',').map((sel) => sel.trim().replace(/\s+/g, ' '))
.filter(Boolean);
}
function mediaQueryBodies(css: string): string[] {
const bodies: string[] = [];
const re = /@media[^{]*\{/g;
let m: RegExpExecArray | null;
while ((m = re.exec(css)) !== null) {
// Walk braces from the opening one to find this at-rule's whole body.
let depth = 1;
let i = m.index + m[0].length;
const start = i;
while (i < css.length && depth > 0) {
if (css[i] === '{') depth++;
else if (css[i] === '}') depth--;
i++;
}
bodies.push(css.slice(start, i - 1));
}
return bodies;
}
describe('FilterBar hero-mode scoping', () => {
const css = withoutComments(fs.readFileSync(CSS, 'utf8'));
it('confirms heroMode still resets the card, which is what makes this matter', () => {
const hero = css.match(/\.heroMode\s*\{([^}]*)\}/);
expect(hero).not.toBeNull();
expect(hero![1]).toMatch(/padding:\s*0/);
});
it('never re-applies card styling to the hero from inside a media query', () => {
const offenders: string[] = [];
for (const body of mediaQueryBodies(css)) {
for (const rule of body.matchAll(/([^{}]+)\{([^{}]*)\}/g)) {
// Only a *bare* .filterBar is dangerous, and it is dangerous wherever
// it appears in a selector list. Scoped variants
// (`.filterBar:not(.heroMode)`) and descendants are fine.
const selectors = selectorsOf(rule[1]);
if (!selectors.includes('.filterBar')) continue;
for (const prop of RESET_BY_HERO_MODE) {
if (new RegExp(`(^|[;\\s])${prop}\\s*:`).test(rule[2])) {
offenders.push(`${rule[1].trim()} sets ${prop}`);
}
}
}
}
// Fix by scoping the rule as `.filterBar:not(.heroMode)`, the way the
// neighbouring rules in the same block already are.
expect(offenders).toEqual([]);
});
});
@@ -0,0 +1,202 @@
/**
* Last distance offered, rendered on both detail templates.
*
* The figure is the one number on these pages that a parent may act on — it is
* easy to read as "we live inside the catchment, we will get a place". These
* tests pin the things that stop it being read that way: the year is always
* present, the caveat is always present, and the blanket "not available"
* sentence appears only when it is actually true.
*/
import { screen, fireEvent } from '@testing-library/react';
import { renderSchoolDetail, renderSecondarySchoolDetail } from '../support/renderSchoolDetail';
import { primaryFixture, secondaryFixture } from '../support/schoolFixtures';
import type { SchoolAdmissionDistance } from '@/lib/types';
const cutoff = (over: Partial<SchoolAdmissionDistance> = {}): SchoolAdmissionDistance => ({
year: 2025,
distance_m: 500,
route_count: 1,
la_name: 'Camden',
distance_unit_raw: 'miles',
...over,
});
describe('primary detail page', () => {
it('shows the figure with the year it belongs to', () => {
renderSchoolDetail({ ...primaryFixture, admissionDistance: cutoff({ distance_m: 772.49, year: 2024 }) });
expect(screen.getByText('0.48 miles')).toBeInTheDocument();
expect(screen.getByText(/Last distance offered/)).toHaveTextContent('September 2024');
});
it('keeps the figure and its metric support readable as two numbers', () => {
// They are flex children with a CSS gap and nothing between them in the
// text layer, which read as "0.48 miles770 m" to a screen reader and to any
// text matcher. Cheap to lose again, so pinned.
renderSchoolDetail({ ...primaryFixture, admissionDistance: cutoff({ distance_m: 772.49 }) });
const tile = document.querySelector('[class*="admissionsTileDistance"]')!;
expect(tile.textContent).toMatch(/0\.48 miles\s+770 m/);
expect(tile.textContent).not.toMatch(/miles\d/);
});
it('never shows the figure without saying it is not a catchment', () => {
renderSchoolDetail({ ...primaryFixture, admissionDistance: cutoff() });
expect(screen.getByText(/not a fixed catchment/)).toBeInTheDocument();
expect(screen.getByText(/moves every year/)).toBeInTheDocument();
});
it('flags that a banded school\'s figure is the widest of several routes', () => {
renderSchoolDetail({ ...primaryFixture, admissionDistance: cutoff({ route_count: 4 }) });
expect(screen.getByText(/4 admission routes/)).toBeInTheDocument();
});
it('renders nothing distance-related when the LA publishes none', () => {
renderSchoolDetail({ ...primaryFixture, admissionDistance: null });
expect(screen.queryByText(/Last distance offered/)).not.toBeInTheDocument();
expect(screen.queryByText(/not a fixed catchment/)).not.toBeInTheDocument();
});
it('carries the figure even with no EES admissions row', () => {
// The two sources are independent; this school has a cut-off and no
// admissions figures. Before this feature the section did not render at all.
renderSchoolDetail({
...primaryFixture,
admissions: null,
admissionsHistory: [],
admissionDistance: cutoff({ distance_m: 1421.05 }),
});
expect(screen.getByText('0.88 miles')).toBeInTheDocument();
});
});
describe('secondary detail page', () => {
it('shows the figure with the year it belongs to', () => {
renderSecondarySchoolDetail({ ...secondaryFixture, admissionDistance: cutoff({ distance_m: 3472.96 }) });
expect(screen.getByText('2.16 miles')).toBeInTheDocument();
expect(screen.getByText(/Last distance offered/)).toHaveTextContent('September 2025');
});
it('drops the blanket "not available" line once a distance exists', () => {
renderSecondarySchoolDetail({ ...secondaryFixture, admissionDistance: cutoff() });
expect(screen.queryByText(/has not published a cut-off distance/)).not.toBeInTheDocument();
expect(screen.getByText(/not a fixed catchment/)).toBeInTheDocument();
});
it('names the authority that would hold the data when there is none', () => {
// The old copy asserted "Historical distance cut-off data is not available
// for this school" on every secondary page, including the ones whose
// council does publish it.
renderSecondarySchoolDetail({ ...secondaryFixture, admissionDistance: null });
expect(screen.getByText(/has not published a cut-off distance/)).toBeInTheDocument();
});
it('makes no claim about publication when the feature is switched off', () => {
// Absent, not null. The API omits the key entirely while the
// admission_distance flag is off, and "Islington has not published a
// cut-off distance" is then a statement about us, not about Islington —
// false wherever the authority does publish one.
renderSecondarySchoolDetail({ ...secondaryFixture, admissionDistance: undefined });
expect(screen.queryByText(/has not published a cut-off distance/)).not.toBeInTheDocument();
expect(screen.queryByText(/Contact the admissions authority/)).not.toBeInTheDocument();
});
});
// ── The Distance section ───────────────────────────────────────────────
describe('Distance section', () => {
it('appears for a school with a figure and coordinates', () => {
const { container } = renderSchoolDetail({
...primaryFixture,
schoolInfo: { ...primaryFixture.schoolInfo, latitude: 51.5, longitude: -0.12 },
admissionDistance: cutoff({ distance_m: 700, year: 2026 }),
});
expect(container.querySelector('#distance')).toBeInTheDocument();
expect(screen.getByText('How far away are you?')).toBeInTheDocument();
});
it('stays away when the school has no coordinates to measure from', () => {
const { container } = renderSchoolDetail({
...primaryFixture,
schoolInfo: { ...primaryFixture.schoolInfo, latitude: null, longitude: null },
admissionDistance: cutoff({ distance_m: 700, year: 2026 }),
});
expect(container.querySelector('#distance')).not.toBeInTheDocument();
});
it('stays away when no figure has been published', () => {
const { container } = renderSchoolDetail({
...primaryFixture,
schoolInfo: { ...primaryFixture.schoolInfo, latitude: 51.5, longitude: -0.12 },
admissionDistance: null,
});
expect(container.querySelector('#distance')).not.toBeInTheDocument();
});
it('shows no year-by-year record — that is held back as a paid feature', () => {
// The page must not leak the history through a table, a chart or a strip of
// per-year verdicts. The API no longer sends it either; this guards the
// render side so a future component cannot quietly put it back.
renderSchoolDetail({
...primaryFixture,
schoolInfo: { ...primaryFixture.schoolInfo, latitude: 51.5, longitude: -0.12 },
admissionDistance: cutoff({ distance_m: 700, year: 2026 }),
});
// Scoped to the section: the page has other tables (the history section's).
const section = document.querySelector('#distance')!;
expect(section.querySelector('table')).toBeNull();
expect(screen.queryByText(/Last distance offered, by year/)).not.toBeInTheDocument();
expect(screen.queryByText(/Not published/)).not.toBeInTheDocument();
expect(screen.queryByText(/too few to read as a trend/)).not.toBeInTheDocument();
});
});
describe('secondary Distance section', () => {
it('renders on the secondary template too', () => {
const { container } = renderSecondarySchoolDetail({
...secondaryFixture,
schoolInfo: { ...secondaryFixture.schoolInfo, latitude: 51.5, longitude: -0.12 },
admissionDistance: cutoff({ distance_m: 3472.96, year: 2026 }),
});
expect(container.querySelector('#distance')).toBeInTheDocument();
});
it('explains a selective school by how it admits rather than as missing data', () => {
renderSecondarySchoolDetail({
...secondaryFixture,
schoolInfo: { ...secondaryFixture.schoolInfo, admissions_policy: 'Selective' },
admissionDistance: null,
});
expect(screen.getByText(/ranked by the entrance test/)).toBeInTheDocument();
expect(screen.queryByText(/has not published a cut-off distance/)).not.toBeInTheDocument();
});
it('reads a consistently undersubscribed school as good news', () => {
renderSecondarySchoolDetail({
...secondaryFixture,
admissionDistance: null,
admissionsHistory: [
{ year: 2022, oversubscribed: false },
{ year: 2023, oversubscribed: false },
{ year: 2024, oversubscribed: false },
],
});
expect(screen.getByText(/has not needed a distance cut-off/)).toBeInTheDocument();
});
});
@@ -0,0 +1,72 @@
import fs from 'fs';
import path from 'path';
import ts from 'typescript';
/**
* Keeps the em dash out of public copy.
*
* Every visitor-facing sentence was rewritten to use a colon, comma, full stop
* or parentheses instead. A new one slips back in unnoticed, because it reads
* fine to whoever wrote it, so this walks the source with the TypeScript parser
* and checks every string literal, template chunk and JSX text node. Comments
* are not nodes the walk visits, so they stay free to use it.
*
* Two uses are allowed:
* - a lone dash standing in for a missing value in a table cell or stat slot,
* which is a data convention rather than prose;
* - the message of a thrown Error, which only a developer reads.
*/
const ROOT = path.join(__dirname, '..', '..');
const SCANNED = ['app', 'components', 'lib'];
const SKIPPED = [path.join(ROOT, 'app', '(payload)')];
function sources(dir: string): string[] {
if (SKIPPED.includes(dir)) return [];
return fs.readdirSync(dir, { withFileTypes: true }).flatMap((entry) => {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) return sources(full);
return /\.tsx?$/.test(entry.name) ? [full] : [];
});
}
function isErrorMessage(node: ts.Node): boolean {
for (let p = node.parent; p; p = p.parent) {
if (ts.isNewExpression(p) && /Error$/.test(p.expression.getText())) return true;
if (ts.isStatement(p)) return false;
}
return false;
}
const EMPTY_VALUE = /^(—|&mdash;)$/;
function dashedCopy(file: string): string[] {
const text = fs.readFileSync(file, 'utf8');
const kind = file.endsWith('.tsx') ? ts.ScriptKind.TSX : ts.ScriptKind.TS;
const sf = ts.createSourceFile(file, text, ts.ScriptTarget.Latest, true, kind);
const found: string[] = [];
const visit = (node: ts.Node) => {
if (
ts.isStringLiteral(node) || ts.isNoSubstitutionTemplateLiteral(node) || ts.isJsxText(node)
|| ts.isTemplateHead(node) || ts.isTemplateMiddle(node) || ts.isTemplateTail(node)
) {
const copy = node.getText().replace(/\s+/g, ' ').trim();
const bare = node.text.trim();
if (/—|&mdash;/.test(copy) && !EMPTY_VALUE.test(bare) && !isErrorMessage(node)) {
const { line } = sf.getLineAndCharacterOfPosition(node.getStart());
found.push(`${path.relative(ROOT, file)}:${line + 1} ${copy}`);
}
}
ts.forEachChild(node, visit);
};
visit(sf);
return found;
}
describe('public copy', () => {
it('uses no em dashes', () => {
const files = SCANNED.flatMap((dir) => sources(path.join(ROOT, dir)));
expect(files.length).toBeGreaterThan(50);
expect(files.flatMap(dashedCopy)).toEqual([]);
});
});
@@ -108,7 +108,7 @@ describe('all-through schools', () => {
it('labels the KS2 block explicitly', () => {
renderSchoolDetail(allThroughFixture);
expect(screen.getAllByText(/Primary — KS2 SATs/).length).toBeGreaterThan(0);
expect(screen.getAllByText(/Primary: KS2 SATs/).length).toBeGreaterThan(0);
});
});
@@ -0,0 +1,73 @@
import { renderHook, act, waitFor } from '@testing-library/react';
import { useSchoolSuggest } from '@/hooks/useSchoolSuggest';
const realFetch = global.fetch;
function mockFetch(rows: unknown[], delayMs = 0) {
global.fetch = jest.fn(async (_url: unknown, init?: { signal?: AbortSignal }) => {
if (delayMs) {
await new Promise((resolve, reject) => {
const t = setTimeout(resolve, delayMs);
init?.signal?.addEventListener('abort', () => {
clearTimeout(t);
reject(Object.assign(new Error('aborted'), { name: 'AbortError' }));
});
});
}
return { ok: true, json: async () => ({ suggestions: rows }) };
}) as unknown as typeof fetch;
}
const ROW = {
urn: 1, school_name: 'Brecknock Primary School', local_authority: 'Camden',
postcode: 'NW1 1AA', phase: 'Primary', school_type: 'Community school',
};
describe('useSchoolSuggest', () => {
beforeEach(() => { jest.useFakeTimers(); });
afterEach(() => { jest.useRealTimers(); global.fetch = realFetch; });
it('does not fetch below the minimum query length', () => {
mockFetch([ROW]);
renderHook(() => useSchoolSuggest('b', true));
act(() => { jest.advanceTimersByTime(500); });
expect(global.fetch).not.toHaveBeenCalled();
});
it('does not fetch at all when disabled', () => {
// The flag being off must mean no request, not a hidden dropdown.
mockFetch([ROW]);
renderHook(() => useSchoolSuggest('brecknock', false));
act(() => { jest.advanceTimersByTime(500); });
expect(global.fetch).not.toHaveBeenCalled();
});
it('debounces rather than firing per keystroke', () => {
mockFetch([ROW]);
const { rerender } = renderHook(
({ q }) => useSchoolSuggest(q, true), { initialProps: { q: 'br' } });
rerender({ q: 'bre' });
rerender({ q: 'brec' });
act(() => { jest.advanceTimersByTime(199); });
expect(global.fetch).not.toHaveBeenCalled();
act(() => { jest.advanceTimersByTime(2); });
expect(global.fetch).toHaveBeenCalledTimes(1);
});
it('opens with results once they arrive', async () => {
mockFetch([ROW]);
const { result } = renderHook(() => useSchoolSuggest('brecknock', true));
act(() => { jest.advanceTimersByTime(200); });
await waitFor(() => expect(result.current.suggestions).toHaveLength(1));
expect(result.current.open).toBe(true);
});
it('close() hides the list without clearing the query', async () => {
mockFetch([ROW]);
const { result } = renderHook(() => useSchoolSuggest('brecknock', true));
act(() => { jest.advanceTimersByTime(200); });
await waitFor(() => expect(result.current.open).toBe(true));
act(() => { result.current.close(); });
expect(result.current.open).toBe(false);
});
});
+158
View File
@@ -0,0 +1,158 @@
import { getNavigationSource } from '@/lib/analytics';
/** jsdom's document.referrer is read-only; redefining it is the way in. */
function referrer(url: string) {
Object.defineProperty(document, 'referrer', { value: url, configurable: true });
}
const ORIGIN = 'http://localhost';
describe('getNavigationSource', () => {
afterEach(() => referrer(''));
it('attributes a visit from a location page to the place layer', () => {
/*
* The one this was added for.
*
* W2 published ~3,900 location pages whose entire purpose is to funnel
* search traffic onto school pages. Before this case existed they fell
* through to 'direct' — so the location layer's contribution was not
* merely missing from the funnel, it was being counted in the bucket you
* read as "typed the URL". The measurement that decides whether W2 worked
* was confidently reporting the wrong answer.
*/
referrer(`${ORIGIN}/schools/barnet`);
expect(getNavigationSource()).toBe('place');
});
it.each([
['/schools/authority/kent', 'authority'],
['/schools/near/sw11', 'outcode'],
['/schools/brentwood/primary', 'phase variant'],
])('covers %s (%s)', (path) => {
referrer(`${ORIGIN}${path}`);
expect(getNavigationSource()).toBe('place');
});
it('still calls a school page "detail", one character away', () => {
// /school/ and /schools/ differ by one letter and mean different things.
// A prefix test written in the wrong order silently merges them.
referrer(`${ORIGIN}/school/100010-brecknock-primary-school`);
expect(getNavigationSource()).toBe('detail');
});
it.each([
['/', 'search'],
['/rankings', 'rankings'],
['/compare?urns=1,2', 'compare'],
])('leaves %s attributed as %s', (path, expected) => {
referrer(`${ORIGIN}${path}`);
expect(getNavigationSource()).toBe(expected);
});
it('treats an external referrer as direct', () => {
// Umami records the real referrer on the pageview; this field is only
// about internal navigation.
referrer('https://www.google.com/search?q=schools+in+barnet');
expect(getNavigationSource()).toBe('direct');
});
it('treats no referrer as direct', () => {
referrer('');
expect(getNavigationSource()).toBe('direct');
});
});
/*
* The defect the existing suite could not see.
*
* Every test above sets document.referrer, which the browser writes only when
* a *document* loads. Every internal navigation in this app is an App Router
* soft navigation — history.pushState, no new document — so document.referrer
* keeps naming whatever opened the tab for the whole session. Verified on
* staging: /schools/brentwood → click a school → URL changes to /school/…
* and document.referrer is still "".
*
* So `from` reported 'direct' for essentially every in-app journey, and the
* suite passed because it only ever exercised the full-page-load path.
*/
function freshAnalytics() {
let mod!: typeof import('@/lib/analytics');
jest.isolateModules(() => {
mod = require('@/lib/analytics');
});
return mod;
}
function at(path: string) {
window.history.pushState({}, '', path);
}
describe('getNavigationSource across a soft navigation', () => {
afterEach(() => {
referrer('');
at('/');
});
it('attributes a school view to the place page the user actually came from', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/schools/brentwood');
recordVisitedPath('/schools/brentwood');
at('/school/115429-brentwood-school');
recordVisitedPath('/school/115429-brentwood-school');
expect(source()).toBe('place');
});
it('does not depend on whether the new path was recorded first', () => {
// The trail is written by a layout-level effect and read by a page-level
// one. React orders those by tree position, which is not a contract worth
// resting a measurement on, so the answer must be the same either way.
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
recordVisitedPath('/rankings');
at('/school/115429-brentwood-school');
expect(source()).toBe('rankings');
});
it('names the previous page, not the current one, when both are schools', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/school/100010-brecknock-primary-school');
recordVisitedPath('/school/100010-brecknock-primary-school');
at('/school/115429-brentwood-school');
recordVisitedPath('/school/115429-brentwood-school');
expect(source()).toBe('detail');
});
it('looks past a return visit to the page the user came back from', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
for (const p of ['/schools/brentwood', '/school/115429-brentwood-school',
'/schools/brentwood']) {
at(p);
recordVisitedPath(p);
}
expect(source()).toBe('detail');
});
it('falls back to the referrer on a real document load, where it is true', () => {
// A fresh module is a fresh document: nothing has been recorded, and
// document.referrer is meaningful again.
const { getNavigationSource: source } = freshAnalytics();
at('/school/115429-brentwood-school');
referrer(`${ORIGIN}/schools/barnet`);
expect(source()).toBe('place');
});
it('still reads an arrival from outside as direct', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/schools/brentwood');
recordVisitedPath('/schools/brentwood');
referrer('https://www.google.com/search?q=schools+in+brentwood');
expect(source()).toBe('direct');
});
});
@@ -0,0 +1,87 @@
import {
canAggregate, aggregateCells,
canRenderBar, toBarSegments, CARD_GROUPS,
type DestinationCell, type DestinationGroup, type DestinationCategory,
} from '@/lib/destinations';
const pub = (category: DestinationCategory, pupils: number, cohort: number): DestinationCell => ({
category, pupils, percentage: (pupils / cohort) * 100, status: 'published',
});
const sup = (category: DestinationCategory): DestinationCell => ({
category, pupils: null, percentage: null, status: 'suppressed',
});
const fullGroup = (): DestinationGroup => ({
cohort: 180,
cells: [
pub('school_sixth_form', 75, 180), pub('sixth_form_college', 21, 180),
pub('further_education', 55, 180), pub('other_education', 6, 180),
pub('apprenticeship', 8, 180), pub('employment', 6, 180),
pub('not_sustained', 5, 180), pub('not_captured', 4, 180),
],
});
describe('canAggregate — R2, computing from components', () => {
it('allows a sum when every component is published', () => {
expect(canAggregate([pub('apprenticeship', 8, 180), pub('employment', 6, 180)])).toBe(true);
});
it('refuses a sum when any component is suppressed', () => {
expect(canAggregate([pub('apprenticeship', 8, 180), sup('employment')])).toBe(false);
});
it('refuses a sum when every component is suppressed', () => {
expect(canAggregate([sup('apprenticeship'), sup('employment')])).toBe(false);
});
});
describe('aggregateCells', () => {
it('sums published cells and derives a percentage from the cohort', () => {
expect(aggregateCells([pub('apprenticeship', 8, 180), pub('employment', 6, 180)], 180))
.toEqual({ pupils: 14, percentage: (14 / 180) * 100 });
});
it('returns null rather than a partial sum when a component is suppressed', () => {
expect(aggregateCells([pub('apprenticeship', 8, 180), sup('employment')], 180)).toBeNull();
});
});
describe('canRenderBar — R1', () => {
it('allows a bar when the whole group is published', () => {
expect(canRenderBar(fullGroup())).toBe(true);
});
it('refuses a bar when a single category is suppressed', () => {
const g = fullGroup();
g.cells[1] = sup('sixth_form_college');
expect(canRenderBar(g)).toBe(false);
});
});
describe('toBarSegments', () => {
it('derives widths from counts, not from rounded percentages', () => {
const segs = toBarSegments(fullGroup());
expect(segs).toHaveLength(8);
expect(segs[0].widthPct).toBeCloseTo((75 / 180) * 100, 10);
expect(segs.reduce((a, s) => a + s.widthPct, 0)).toBeCloseTo(100, 6);
});
it('throws rather than silently leaving a gap when the group is suppressed', () => {
const g = fullGroup();
g.cells[1] = sup('sixth_form_college');
expect(() => toBarSegments(g)).toThrow(/suppressed/i);
});
});
describe('CARD_GROUPS', () => {
it('partitions every destination category exactly once, plus the absence', () => {
const grouped = Object.values(CARD_GROUPS).flat();
expect(new Set(grouped).size).toBe(grouped.length);
expect(grouped).toEqual(expect.arrayContaining([
'school_sixth_form', 'sixth_form_college', 'further_education',
'other_education', 'apprenticeship', 'employment',
]));
expect(grouped).not.toContain('not_sustained');
expect(grouped).not.toContain('not_captured');
});
});
+58
View File
@@ -0,0 +1,58 @@
import { getFlags, FLAGS_REVALIDATE } from '@/lib/flags';
// jsdom provides no global fetch, so there is nothing for jest.spyOn to attach
// to — assign it and restore the original afterwards. This is the first test
// here to mock fetch; later ones should follow this shape.
const realFetch = global.fetch;
function mockFetch(impl: () => Promise<unknown>) {
global.fetch = jest.fn(impl) as unknown as typeof fetch;
}
describe('getFlags', () => {
afterEach(() => { global.fetch = realFetch; });
it('returns the flags the API reports', async () => {
mockFetch(async () => ({
ok: true,
json: async () => ({ admission_distance: true }),
}));
await expect(getFlags()).resolves.toEqual({ admission_distance: true });
});
it('returns no flags rather than throwing when the API is down', async () => {
// A page that cannot read flags must render everything dark, not 500.
// Fail-closed is the same direction as the backend's default.
mockFetch(async () => { throw new Error('ECONNREFUSED'); });
await expect(getFlags()).resolves.toEqual({});
});
it('returns no flags rather than throwing on a non-200', async () => {
mockFetch(async () => ({ ok: false, status: 503 }));
await expect(getFlags()).resolves.toEqual({});
});
/*
* Reading a flag pins the calling route's ISR floor: Next uses the LOWEST
* revalidate among a route's fetches for the whole route. That is why the
* revalidate is an argument rather than the constant.
*
* Every SEO route here declares `revalidate = 604800`. A gate that read
* flags at the 300s default would drop the whole school and place corpus
* from a weekly cache to a 5-minute one, which is a large origin-load
* regression to pay for a feature flag.
*/
it('reads at the 300s floor by default', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags();
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: FLAGS_REVALIDATE } });
});
it('lets a caller pass its own route floor instead', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags(604800);
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: 604800 } });
});
});
@@ -0,0 +1,182 @@
/**
* Last distance offered — formatting and the caveats attached to the figure.
*
* The assertions about the route note and the year are not cosmetic. A cut-off
* shown without its year, or a banded school's widest cut-off shown as if it
* were the only one, tells a parent something false about their chances of a
* place — so both are pinned here rather than left to the component.
*/
import { formatCutoffDistance, formatMiles, formatEntryYear } from '@/lib/utils';
import {
describeCutoff, describeCutoffAbsence, compareToCutoff, CUTOFF_UNCERTAINTY_M,
} from '@/components/school/lastDistanceOffered';
import type { SchoolAdmissionDistance } from '@/lib/types';
const distance = (over: Partial<SchoolAdmissionDistance> = {}): SchoolAdmissionDistance => ({
year: 2025,
distance_m: 500,
route_count: 1,
la_name: 'Camden',
distance_unit_raw: 'miles',
...over,
});
describe('formatCutoffDistance', () => {
it('leads with miles, the unit councils publish in', () => {
expect(formatCutoffDistance(500)).toEqual({ primary: '0.31 miles', secondary: '500 m' });
expect(formatCutoffDistance(1609.344)).toEqual({ primary: '1.00 miles', secondary: '1.6 km' });
});
it('switches to kilometres for the support figure above a kilometre', () => {
expect(formatCutoffDistance(3472.96)!.secondary).toBe('3.5 km');
});
it('stays in miles at short range, where it used to swap to metres', () => {
// The swap made a single number easier to read and a comparison harder:
// "69 m away — inside the cut-off of 0.17 miles" asked the reader to
// convert between units to check a claim we had already made for them.
expect(formatCutoffDistance(27)).toEqual({ primary: '0.02 miles', secondary: '30 m' });
expect(formatCutoffDistance(69)!.primary).toMatch(/miles$/);
});
it('describes a distance too short for two decimal places', () => {
// Rather than a flat "0.00 miles", which reads as no distance at all.
expect(formatMiles(5)).toBe('under 0.01 miles');
expect(formatMiles(0)).toBe('under 0.01 miles');
expect(formatMiles(20)).toBe('0.01 miles');
});
it('returns null rather than a zero cut-off', () => {
// 0.0 miles appears in the source where a school filled on a higher
// criterion. Rendered as "0.00 miles" it would read as the opposite.
expect(formatCutoffDistance(0)).toBeNull();
expect(formatCutoffDistance(null)).toBeNull();
expect(formatCutoffDistance(undefined)).toBeNull();
expect(formatCutoffDistance(Number.NaN)).toBeNull();
});
});
describe('formatEntryYear', () => {
it('names the intake, not the academic year', () => {
// formatAcademicYear would render 2025 as "2025/26", which reads as a
// school year rather than the September a child started.
expect(formatEntryYear(2025)).toBe('September 2025');
expect(formatEntryYear(null)).toBe('');
});
});
describe('describeCutoff', () => {
it('always carries the entry year alongside the figure', () => {
const d = describeCutoff(distance({ distance_m: 772.49, year: 2024 }));
expect(d).not.toBeNull();
expect(d!.primary).toBe('0.48 miles');
expect(d!.entryYear).toBe('September 2024');
});
it('says nothing about routes for a school with one', () => {
expect(describeCutoff(distance({ route_count: 1 }))!.routeNote).toBeNull();
expect(describeCutoff(distance({ route_count: null }))!.routeNote).toBeNull();
});
it('warns that a banded school\'s figure is the widest of several', () => {
const note = describeCutoff(distance({ route_count: 4 }))!.routeNote;
expect(note).toContain('4 admission routes');
expect(note).toContain('shorter cut-off');
});
it('is null when there is nothing publishable', () => {
expect(describeCutoff(null)).toBeNull();
expect(describeCutoff(undefined)).toBeNull();
expect(describeCutoff(distance({ distance_m: null }))).toBeNull();
});
});
// ── "Would we have got in?" ────────────────────────────────────────────
describe('compareToCutoff', () => {
it('never states the two figures in different units', () => {
/*
* The reported defect: "69 m away — inside the September 2026 cut-off of
* 0.17 miles". Both numbers are correct and the sentence is still useless,
* because checking it means converting one of them.
*
* Swept across the range where the old formatter switched units, so a
* future readability tweak to one figure cannot reintroduce the mismatch
* in the other.
*/
const mixed: string[] = [];
for (const homeM of [0, 5, 27, 69, 99, 100, 260, 800, 1609, 5000]) {
for (const cutoffM of [30, 69, 100, 270, 1000, 3500]) {
const { headline } = compareToCutoff(homeM, cutoffM, 2026);
const hasMetres = /\d\s?m\b/.test(headline);
const milesCount = (headline.match(/miles/g) ?? []).length;
// Two figures, both in miles, and no metric reading anywhere near them.
if (hasMetres || milesCount !== 2) {
mixed.push(`home=${homeM}m cutoff=${cutoffM}m -> ${headline}`);
}
}
}
expect(mixed).toEqual([]);
});
it('reads back the reported case in one unit', () => {
expect(compareToCutoff(69, 270, 2026).headline)
.toBe('0.04 miles away, inside the September 2026 cut-off of 0.17 miles.');
});
it('calls a clearly nearer home inside, and names the year', () => {
const r = compareToCutoff(300, 800, 2026);
expect(r.verdict).toBe('inside');
expect(r.headline).toContain('September 2026');
});
it('calls a clearly further home beyond', () => {
expect(compareToCutoff(4000, 800, 2026).verdict).toBe('outside');
});
it('refuses to call a result inside the measurement error, either way', () => {
// A postcode centroid covers several addresses, so a margin this fine is
// noise. With one published year there is no other year to fall back on,
// which makes this band the only thing standing between a parent and a
// place they do not have.
expect(compareToCutoff(800 - CUTOFF_UNCERTAINTY_M / 2, 800, 2026).verdict).toBe('too-close');
expect(compareToCutoff(800 + CUTOFF_UNCERTAINTY_M / 2, 800, 2026).verdict).toBe('too-close');
expect(compareToCutoff(800, 800, 2026).detail).toContain('measurement error');
// The explanation is not welded to the headline, so it does not run at
// headline weight in the result block.
expect(compareToCutoff(800, 800, 2026).headline).not.toContain('measurement error');
expect(compareToCutoff(300, 800, 2026).detail).toBeNull();
});
it('treats the band as exclusive at its edge', () => {
// Exactly on the boundary is still too close; one metre past it is not.
expect(compareToCutoff(800 - CUTOFF_UNCERTAINTY_M, 800, 2026).verdict).toBe('too-close');
expect(compareToCutoff(800 - CUTOFF_UNCERTAINTY_M - 1, 800, 2026).verdict).toBe('inside');
});
});
describe('describeCutoffAbsence', () => {
it('explains a selective school by how it admits, not as missing data', () => {
const s = describeCutoffAbsence({ localAuthority: 'Kent', admissionsPolicy: 'Selective' });
expect(s).toContain('entrance test');
expect(s).not.toContain('has not published');
});
it('reads a consistently undersubscribed school as good news', () => {
const s = describeCutoffAbsence({
localAuthority: 'Camden',
admissionsHistory: [
{ year: 2022, oversubscribed: false },
{ year: 2023, oversubscribed: false },
{ year: 2024, oversubscribed: false },
],
});
expect(s).toContain('has not needed a distance cut-off');
});
it('otherwise names the authority that would hold the figure', () => {
expect(describeCutoffAbsence({ localAuthority: 'Camden' }))
.toContain('Camden has not published');
});
});
@@ -0,0 +1,68 @@
/**
* School pages had no BreadcrumbList and no links into the location layer.
* Both are fixed by the same data — the `places` array the API now returns —
* so they are tested together.
*/
import { schoolBreadcrumbJsonLd } from '@/lib/jsonld';
const essex = { kind: 'authority', slug: 'essex', name: 'Essex', count: 480, url: '/schools/authority/essex', phases: [] };
const brentwood = { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 37, url: '/schools/brentwood', phases: [] };
const outcode = { kind: 'outcode', slug: 'cm15', name: 'CM15', count: 12, url: '/schools/near/cm15', phases: [] };
describe('school breadcrumbs', () => {
it('reads home to authority to town to school', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood],
});
expect(ld['@type']).toBe('BreadcrumbList');
expect(ld.itemListElement.map((i) => i.name))
.toEqual(['schoolcompare', 'Essex', 'Brentwood', 'Brentwood School']);
expect(ld.itemListElement.map((i) => i.position)).toEqual([1, 2, 3, 4]);
});
it('skips a level the school has no published place for', () => {
// A school whose town falls below the publish threshold has no town page.
// The trail closes over the gap rather than linking to a 404.
const ld = schoolBreadcrumbJsonLd({
name: 'Lone School', url: '/school/1-lone-school', places: [essex],
});
expect(ld.itemListElement.map((i) => i.name))
.toEqual(['schoolcompare', 'Essex', 'Lone School']);
expect(ld.itemListElement.map((i) => i.position)).toEqual([1, 2, 3]);
});
it('omits outcodes, which are not a place a breadcrumb reads through', () => {
// CM15 is a useful link in the module but nonsense in a trail: nobody
// navigates Essex → CM15 → school.
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood, outcode],
});
expect(JSON.stringify(ld)).not.toContain('cm15');
});
it('still produces a valid trail when the school has no places at all', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Orphan School', url: '/school/2-orphan-school', places: [],
});
expect(ld.itemListElement.map((i) => i.name)).toEqual(['schoolcompare', 'Orphan School']);
});
it('uses absolute urls, as every other entity on the site does', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood],
});
for (const item of ld.itemListElement) {
expect(item.item).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\//);
}
// The root is the homepage: there is no /schools index page to link to.
expect(ld.itemListElement[0].item).toBe('https://www.schoolcompare.co.uk/');
});
});
@@ -0,0 +1,83 @@
import { computeSecondaryFlags, buildSecondaryNavItems } from '@/lib/schoolSections';
import type { School, SchoolDestinations } from '@/lib/types';
const schoolInfo = {
urn: 137083, school_name: 'Northbrook Academy', phase: 'Secondary',
has_sixth_form: true,
} as unknown as School;
const base = { schoolInfo, yearlyData: [], deprivation: null, finance: null };
const phase = (categories = 1) => ({
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: Array.from({ length: categories }, () => ({
category: 'school_sixth_form' as const,
pupils: 75, percentage: 41.7, status: 'published' as const,
})),
},
},
});
const ks4Only: SchoolDestinations = { ks4: phase(), ks5: null };
const both: SchoolDestinations = { ks4: phase(), ks5: phase() };
describe('computeSecondaryFlags — destinations', () => {
it('flags KS4 destinations when the block carries categories', () => {
const flags = computeSecondaryFlags({ ...base, destinations: ks4Only });
expect(flags.hasKs4Destinations).toBe(true);
expect(flags.hasKs5Destinations).toBe(false);
});
it('flags both phases when both are present', () => {
const flags = computeSecondaryFlags({ ...base, destinations: both });
expect(flags.hasKs4Destinations).toBe(true);
expect(flags.hasKs5Destinations).toBe(true);
});
it('flags neither when the block is absent', () => {
const flags = computeSecondaryFlags({ ...base, destinations: null });
expect(flags.hasKs4Destinations).toBe(false);
expect(flags.hasKs5Destinations).toBe(false);
});
it('does not flag a phase whose groups carry no categories', () => {
const empty: SchoolDestinations = {
ks4: { cohort_year: '2022/23', groups: {} }, ks5: null,
};
expect(computeSecondaryFlags({ ...base, destinations: empty }).hasKs4Destinations)
.toBe(false);
});
it('does not flag a phase whose only group has an empty category list', () => {
const empty: SchoolDestinations = { ks4: phase(0), ks5: null };
expect(computeSecondaryFlags({ ...base, destinations: empty }).hasKs4Destinations)
.toBe(false);
});
});
describe('buildSecondaryNavItems — destinations', () => {
const navInput = {
ofsted: null, admissions: null, admissionDistance: null,
hasLocation: false, yearlyDataLength: 0,
};
it('adds both entries, after GCSEs', () => {
const flags = computeSecondaryFlags({ ...base, destinations: both });
const ids = buildSecondaryNavItems({ ...flags, hasResults: true }, navInput)
.map(i => i.id);
expect(ids).toContain('destinations');
expect(ids).toContain('post16-destinations');
expect(ids.indexOf('destinations')).toBeGreaterThan(ids.indexOf('gcse'));
expect(ids.indexOf('post16-destinations')).toBe(ids.indexOf('destinations') + 1);
});
it('adds no entry for a phase that will not render — the nav must not link to a missing anchor', () => {
const flags = computeSecondaryFlags({ ...base, destinations: null });
const ids = buildSecondaryNavItems(flags, navInput).map(i => i.id);
expect(ids).not.toContain('destinations');
expect(ids).not.toContain('post16-destinations');
});
});
Loaded 100 of 299 files, more files were not shown because too many files have changed in this diff. Show more