Compare commits

...
Author SHA1 Message Date
TudorandClaude Opus 5 0571d1c0ff fix(web): stop the sheet-open rule stealing .sectionNav's layout
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
Review catch, and a bad one: the previous commit anchored its insertion on
`padding: 0.5rem 0.75rem;` and closed .sectionNav there. Everything that
followed in the rule — margin-bottom, box-shadow, display: flex, align-items,
gap — was orphaned into .sectionNavSheetOpen, which is only applied while the
mobile jump sheet is open.

So the sticky nav lost its flex layout, spacing and shadow in the closed
state, which is virtually every page view on every school detail page. A
site-wide regression introduced by a fix for one mobile menu.

Redone by anchoring on the complete rule, closing brace included, so nothing
can be orphaned. .sectionNav is now byte-identical to main and the diff is
purely additive; .sectionNavSheetOpen carries the z-index and nothing else.

The staging experiment that validated this fix set nav.style.zIndex = '1100'
with every other declaration intact, so it was always testing this version
rather than the broken one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 20:43:58 +01:00
TudorandClaude Opus 5 180d6e9b3e fix(web): lift the jump sheet above the bottom tab bar
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 37s
On mobile the last item in "Jump to section" was painted over by the fixed
bottom tab bar and could not be tapped. Reported as nearby schools missing
from the menu; it was there, underneath the bar.

The sticky nav sets `position: sticky` with `z-index: 10`, which makes it a
stacking context. The sheet's own `z-index: 1600` therefore orders it only
inside that context — against the tab bar (z-index 1000) the nav's 10 is what
counts, so the bar wins. Verified on staging: every menu item returns itself
from elementFromPoint except the last, which returns the tab bar.

Latent rather than new. With five sections the list stopped just above the bar;
"Nearby schools" made six, and the sixth is the first to reach it. Any section
added later would have done the same.

Lifted only while the sheet is open, and only to 1100 — above the bar, below
the comparison toast (2000), the fullscreen map (5000) and the info popover
(9999). The backdrop rises with it, so tapping over the bar now dismisses the
sheet instead of navigating away.

The journey asks what a thumb asks: for each item, whether it is the topmost
element at its own centre. A bounding-box check cannot see this — the item is
in the viewport and the right size, just underneath something. Confirmed to
fail against current staging, naming "Nearby schools", before the fix.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 20:35:25 +01:00
tudor 80f405123e Merge pull request 'test(e2e): wait for the carousel's smooth scroll to settle before measuring it' (#153) from fix/nearby-journey-waits-for-smooth-scroll into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m25s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 28s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m53s
Reviewed-on: #153
2026-09-22 14:55:16 +00:00
TudorandClaude Opus 5 077aca6008 test(e2e): wait for the carousel's smooth scroll to settle before measuring it
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 32s
Second failure of the same assertion, and the first fix addressed a real but
different problem. This is the one that was actually producing the number.

The arrows scroll with `behavior: 'smooth'`. The test polled for "has it moved
at all" — satisfied 50ms in, at 13px of a 1300px journey — then recorded the
offset, clicked, and recorded again. The row was still travelling throughout,
so the delta it measured was the tail of the arrow's animation, not the effect
of the selection. Hence a deterministic 659, roughly half of the 1317 this
school's row scrolls.

Measured against staging rather than reasoned about: the animation runs about
700ms, and the samples are in the helper's comment.

settledScrollLeft waits for two identical readings before trusting one. The app
was never at fault — driving staging by hand, the offset holds at 1317 across
the selection, exactly as intended.

Verified against staging both ways: main's version of this test fails there,
this version passes, along with all three mobile widths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 15:13:32 +01:00
tudor 9dba5ff1ff Merge pull request 'test(e2e): fix the nearby-schools journey clicking a card it scrolled past' (#152) from fix/nearby-scroll-journey-clicks-offscreen-card into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m23s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 25s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m49s
Reviewed-on: #152
2026-09-22 13:50:25 +00:00
TudorandClaude Opus 5 f530a912bc test(e2e): stop the nearby-schools journey clicking a card it scrolled past
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 21s
The assertion that adding to compare does not reset the carousel failed on
staging: 537 → 2. The app was not at fault. Playwright scrolls a target into
view before clicking, and the test clicked the FIRST card's button after
paging the row to the end — so Playwright scrolled the container back to the
start, and the assertion measured that.

Reproduced on a static page with no React on it: a snap scroller at 615,
Playwright clicks the off-screen first card, scrollLeft becomes 2. Scroll-snap
was ruled out first — mandatory, proximity and no-snap all behave identically
when the button mutates in place.

Now clicks the last card's button, which is visible at the end of the travel,
and allows a few pixels for snap and sub-pixel adjustment while still failing
on a reset to the start. The property was never actually under test before;
it is now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 14:19:00 +01:00
tudor 029fe8d8a6 Merge pull request 'fix: order nearby schools by distance, not by how alike they are' (#151) from fix/nearby-schools-order-by-distance into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m24s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 30s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m51s
Reviewed-on: #151
2026-09-22 13:09:38 +00:00
TudorandClaude Opus 5 cd1c5d1e1a docs: revise the spec to the design that survived staging
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 57s
The spec described the tier system as the design of record. It is gone, so
the document was describing something the code deliberately does not do.

The revision note and the "why not, having built it the other way first"
passage are kept rather than overwritten. The mistake is the instructive part:
treating a preference as a constraint inverted the ranking, and the stopping
rule added to prevent weak distant matches is what guaranteed six Catholic
schools and no community school down the road. A spec that quietly presents the
second design as the plan teaches nobody why the first one failed.

The mockup link is annotated as one revision behind rather than silently left
to look current.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
TudorandClaude Opus 5 cd6a45bf7d refactor: rename similar → nearby, so the code says what the section does
The section ranks on distance and is headed "Other schools nearby", but every
identifier still called it "similar" — the exact drift that leaves a later
reader trusting a name over the behaviour.

Mechanical: files, the module, the payload key, the type, the components, the
prop. No behaviour change; the suites are unchanged in count and still green.
Free to do now because #150 has not merged, so the payload key rename needs no
lockstep deploy. Uses of "similar" that are ordinary English — progress
measures compared to similar pupils, and unrelated comments — are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
TudorandClaude Opus 5 5c0ccc693d fix: order nearby schools by distance, not by how alike they are
Reported from staging: a Catholic primary showed six Catholic primaries, none
of them close enough to be a real option, and omitted the community school
down the road.

Three causes, compounding. Ranking put tier before distance, so a faith match
at 2.9 miles outranked a community school at 0.3. The ENOUGH=3 stopping rule —
added so a cap of six would not drag in weak distant matches — filled the row
from the best tier before it ever widened, which is what made every card
Catholic. And a 3-mile tier-1 radius is sane for a secondary and most of a city
for a primary, whose catchments are routinely under a mile.

The premise was backwards. For a parent, distance is a constraint and intake is
a preference; a school beyond a primary catchment is not a weaker option, it is
not an option. So distance now decides the order and nothing else does. The
hard filters are untouched — they were always where the defensibility lived.
Similarity survives as chips on the card: reported, so a reader applies their
own weighting, rather than ranked, so we apply ours for them.

Reach is capped per phase (primary 2, secondary 6, post-16 10) as a sanity
bound, not a target: ordering already handles density, so the cap only decides
what happens where an area is sparse. A primary with nothing inside two miles
now renders no section, which is the honest answer.

Deleted: the tier system, the stopping rule, the tier-dependent lede, the
`tier` field, the tier-3 fallback chip and its style. select_similar also stops
taking is_secondary — it reads the phase from the subject's own row, so no
caller can hand it one that disagrees with the data.

The heading is now "Other schools nearby". The hard filters still guarantee a
comparable set, but nothing ranks on likeness, so the heading no longer says it
does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
tudor 151cf4bc80 Merge pull request 'feat: similar schools nearby on the detail page' (#150) from feat/similar-schools-nearby into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 44s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m25s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m14s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 6s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 42s
Reviewed-on: #150
2026-09-22 05:53:06 +00:00
TudorandClaude Opus 5 83dc5ae5dc docs: correct the metric rule for "16 plus"
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m11s
The spec said the metric follows the template the page renders. That is no
longer true for GIAS phase 6: a sixth-form college renders the primary
template but is matched, correctly, against secondaries. The rule is phase
group membership, decided once in is_secondary_phase — and the section's lede
noun comes from the school's phase rather than its template for the same
reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 06:41:59 +01:00
TudorandClaude Opus 5 2175dccb7c fix(api): treat "16 plus" as secondary, the way PHASE_GROUPS already does
PR Checks / Frontend Typecheck + Tests (pull_request) Canceled after 9s
PR Checks / Backend Smoke (pull_request) Canceled after 0s
PR Checks / Build Backend (no push) (pull_request) Canceled after 0s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 0s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
GIAS phase 6 is "16 plus", and PHASE_GROUPS deliberately files it in the
secondary group. The payload helper decided the same question with
`"secondary" in phase_text`, which that value does not satisfy — so a
sixth-form college was handed the primary bucket and offered infant schools
as its peers, with the KS2 metric key to label them. No crash; just a page
confidently showing the wrong schools.

The decision now lives in similar_schools.is_secondary_phase, beside the
PHASE_GROUPS bucket it selects from, so the two cannot drift again. A test
pins them together.

The same binary assumption had a second output. computeSchoolFlags tests for
the substring too, so a 16-plus school renders the primary template, and the
composer was labelling the section from the template: "Other primary schools
near <sixth form college>" above a row of secondaries. The section now
derives its noun from the school's own phase, which also removes the
duplicated wording from both composers. A 16-plus school's candidates span
the whole secondary group, so no single noun fits and it gets the honest
general one.

Reported in review on #150.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 06:41:49 +01:00
TudorandClaude Opus 5 bd2a6c385b test(e2e): cover the similar-schools section, compare hand-off and mobile widths
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m13s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 37s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 3m1s
Two things here cannot be covered anywhere else. jsdom has no layout, so
scrollWidth and clientWidth are both 0 and the arrows' disabled state can only
be measured by a real engine. And the scroll position surviving a selection is
DOM state rather than React state, so only a real browser can prove the row
does not jump back when the footer re-renders.

MOBILE.md asks for a Playwright width check and records that it was not written
because Playwright was not in the project. It is — this suite — so the check
exists now, scoped to the page this feature touches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:43:38 +01:00
TudorandClaude Opus 5 4e0d8bcf87 feat(web): render similar schools on both detail templates
Inside SchoolDetailShell rather than after it, because the sticky nav's
scroll-spy finds sections with getElementById and can only reach one that
lives in the shell. Last in the order, and last in the nav, because the two
must agree or the nav links to an anchor that was never rendered.

hasSimilarSchools is optional on NavItemsInput, matching hasLocation beside
it: absent has to mean "no section", and making it required would have
churned ten unrelated call sites for no added safety.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:42:59 +01:00
TudorandClaude Opus 5 91314a80b5 feat(web): similar schools section, honest about what it matched
Server-rendered cards inside a client carousel that scrolls rather than
paginates, so all six links stay in the initial HTML and the row still works
with JavaScript off. Three client islands, split by what each needs: a school,
the whole selection, a DOM ref.

The lede claims a similar intake only when no card came from tier 3, chips
list what a school actually shares, a missing figure reads "Not published",
and the neighbour's number carries no valence colour — green and terracotta
mean "against England" everywhere else, and colouring it here would read as
ranking the neighbours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:41:40 +01:00
TudorandClaude Opus 5 52b00ac752 feat(api): serve similar schools on the detail endpoint
Rides in the existing payload rather than a new endpoint: the page already
makes one server fetch for its data, and /school/[slug] regenerates weekly,
so the per-request cost is paid once per school per week.

Wrapped so a failure in selection never 500s a page that is otherwise
complete — the posture get_supplementary_data already takes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:40:35 +01:00
TudorandClaude Opus 5 b571d9c549 feat(api): decide which nearby schools a page may offer
Hard filters encode claims the section may not make — a selective school is
not an alternative to a non-selective one, a special school is not comparable
to a mainstream one, a Girls school is not an option for a Boys school's
reader — so they never relax. Soft preferences describe closeness of fit, so
they relax across three tiers, and only far enough to reach three; the
remaining slots up to six fill from the tiers already opened.

PHASE_GROUPS moves to schemas.py so this module can share it without
importing app, which would be a cycle.

_mask() exists because Series.apply on an empty Series returns a DataFrame,
and using that as a mask drops every column — so the next lookup raises
KeyError instead of yielding no rows. A special school with no special school
near it empties the frame at the provision filter, which is the ordinary case
for most special schools, so this was a crash on a common path rather than an
edge case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:40:10 +01:00
TudorandClaude Opus 5 8a23e3657d docs: meet the mobile baseline, which the carousel was not doing
Checked the section against MOBILE.md at its three reference widths instead
of assuming the breakpoints were enough. Two real failures at 360px.

The arrows sat in the heading's flex row, taking 96px from a 328px card and
crushing the lede into a four-line column — for a control that swiping
already provides. Below 640px they are now gone, the header is a single
column, one card shows at 86% so the next one peeks, and the affordance is
the right-edge scroll-fade MOBILE.md already documents for horizontal
scrollers. The fade lifts at the end of the travel, so the at-end state is
computed whether or not an arrow exists to consume it.

The arrow and add-to-compare buttons were 40px against a 44px floor. Both
are 44 now. A card title's own box is shorter, but its hit area is the whole
card through the ::after overlay, so it passes on the target that actually
receives the tap.

360, 390 and 430 now all report zero overflow, no failing tap targets and no
text under 11px. MOBILE.md wanted a Playwright width check and recorded that
Playwright was not in the project; it is, so the journey now carries one for
this page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:37:34 +01:00
TudorandClaude Opus 5 e4e8f02599 docs: six schools behind a carousel, and a rule for when to stop widening
Three schools is not a neighbourhood in inner London, so the cap is six with
three visible and arrows for the rest.

Raising the cap exposes something the old cap hid. Tiers exist to reach a
usable set, and with six slots a naive loop would keep widening to fill them
— dragging in tier-3 schools ten miles away to sit beside three good matches
that had already earned the row. So tiers now stop relaxing once three are
found, and the remaining slots are filled only from the tiers already used.
Four tier-1 matches never open tier 2.

The carousel scrolls a list rather than swapping a view: all six cards are in
the initial HTML, so every link stays crawlable and the row still scrolls with
JavaScript off. The arrows' edge test carries an 8px tolerance because the
scroller's focus-ring padding is the first snap position — a row at rest
reports scrollLeft 2, and an exact test for 0 left the back arrow live and
pointing nowhere. Caught in the mockup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:35:11 +01:00
TudorandClaude Opus 5 b62dc17532 docs: drop the method disclosure, keep the one caveat that earns its place
The "how these schools are chosen" panel restated what the section already
shows — the phase in the lede, the shared characteristics on each card, the
distance above each name — so it cost space to say nothing new.

One line survives, and it is not a method note. A reader who sees "0.6 miles
away" and takes it for the walk has been misled by us, and no other element
on the card corrects that. The rest were claims the selection rules keep
true without narrating them.

Also records what happens past the third school: surplus matches are dropped
silently, because NearbyPlaces below already leads to the full lists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:29:07 +01:00
TudorandClaude Opus 5 4d7762d796 docs: implementation plan for similar schools nearby
Five tasks, each ending in a green test run and a commit: the pure
selection module, the endpoint key, the section and its two client
islands, the wiring into both templates, and the journey.

The selection logic gets its own module rather than another 200 lines in
app.py, which means the tier rules are testable against a synthetic frame
with no TestClient, no database and no monkeypatch. Moving PHASE_GROUPS
to schemas.py is what keeps that import acyclic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:21:14 +01:00
TudorandClaude Opus 5 3650f7d8b7 docs: design for similar schools nearby on the detail page
A school page links outward to its places and never to another school.
This section adds that edge: three nearby schools of the same phase and a
comparable intake, each a crawlable link and each addable to the basket.

The design separates hard filters from soft preferences and never confuses
them. Selectivity, provision and opposite-sex intake are claims the section
cannot make, so they never relax, even where that means no section renders.
Gender and religious character describe closeness of fit, so they relax in
tiers — and the card states what actually survived rather than padding with
a match it did not earn.

Includes the mockup the design is drawn against, with all three tier states
live in both themes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:18:38 +01:00
tudor dfce308f1f Merge pull request 'fix(ci): make a failed release check say what it actually saw' (#149) from fix/release-check-diagnostics into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m24s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 28s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m9s
Reviewed-on: #149
2026-09-15 15:09:46 +00:00
TudorandClaude Opus 5 64ae71d7ab fix(ci): make a failed release check say what it actually saw
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m4s
The staging poller swallowed every failure identically, so a run that
timed out told us only that the expected release never appeared — not
whether the proxy refused us, the endpoint was down, or the containers
were still serving an older build. The public staging proxy also answers
403 to urllib's default user agent while the release endpoint is healthy,
which looked exactly like a deployment that never arrived.

Identify the poller, and report each distinct observation once: HTTP
status, connection failure type, invalid JSON, or the release identities
actually reported. The timeout error carries the last observation and the
identity it wanted. Responses and the base URL stay out of the logs —
only validated sha/build_id fields are echoed back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 16:02:14 +01:00
tudor 7ab084dd3a Merge pull request 'feat: verify what we actually deployed, and publish data all-or-nothing' (#148) from feat/release-identity-and-reliability-gate into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 0s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m31s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m27s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Failing after 5m3s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Skipped
Reviewed-on: #148
2026-09-15 11:13:08 +00:00
TudorandClaude Opus 5 5ad1cbfb53 fix(api): bound the search candidate set instead of draining Typesense
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 39s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 6m44s
Fetching every match kept scoped searches correct but left the number of
round trips in the caller's hands: a one-letter query, or a deliberately
broad one, walked the whole collection a page at a time.

Cap the candidate set at 1,000 URNs — four pages — and return the
relevance-ordered prefix when the ceiling is hit. That is still far more
than one page, so the API's own authority, phase and postcode filters
keep the matches they need, while latency and upstream load stay bounded.
A capped query is logged so a genuinely truncated search is visible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 11:21:45 +01:00
TudorandClaude Opus 5 0c901cd0d1 feat(ci): gate promotion on the image set that actually passed E2E
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 6m3s
Staging health polling asked only whether something answered HTTP 200 at
the base URL. It could not tell the new deployment from the old one, so
journeys could pass against the previous release, and concurrent merges
could move the staging tags underneath a run in flight.

Each staging run now mints a build ID and stamps all three images with
the commit and that ID, as labels and — for frontend and backend — as a
build-time JSON file that environment overrides cannot rewrite.
/release.json reports both identities uncached, and scripts/ci/release.py
polls for the expected pair before and after the journeys. Only then are
the captured build digests tagged verified-<sha>.

Promotion resolves those verified tags to immutable digests, revalidates
their labels, and refuses a mixed or incomplete set before any :prod tag
moves. The whole staging workflow shares one concurrency group with
cancellation disabled, so releases serialise.

The scripts are stdlib-only and unit-tested against mocked registry and
HTTP behaviour; PR checks now run the pipeline and CI suites too. The
runbook records what this cannot prove locally, and that the first
rollout needs a commit built by this workflow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:17:50 +01:00
TudorandClaude Opus 5 7b41218e6e fix(web): show an outage as an outage, and drop superseded fetches
The home page caught every fetch failure and rendered its empty state,
so a backend outage looked like a site with no schools in it. School
pages turned any error into notFound(), which told visitors — and
crawlers — that a real school had ceased to exist. Place fetches did the
same by returning [] and null. Failures now reach a retryable error
boundary; only a genuine 404 still calls notFound().

"Load more" and the map fetch resolved against whatever state existed
when they returned, so results from an abandoned search appended
themselves to the new ones. Each fetch now carries an AbortController
and checks that its search scope is still current before touching state.
The map only records its cache key on success, so a failed load retries
instead of pinning the stale marker set.

Jest ignored .next/, whose build output otherwise shadowed real suites.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:17:18 +01:00
TudorandClaude Opus 5 9b75f54206 fix(api): publish a validated dataset, and stop truncating search
Reload cleared the caches first and rebuilt afterwards, so any failure
left the API serving nothing, and requests arriving mid-reload saw a
half-swapped state. It now builds and validates the replacement frames,
place registry, reverse index and sitemaps off the request loop, then
publishes them in one synchronous step under a lock. A failed reload
returns 503 and keeps the previous data. Sitemap regeneration takes the
same path rather than clearing the live registry up front.

Typesense search returned at most one page of hits and used an empty
list for both "no matches" and "search is down", so a genuine empty
result silently fell back to substring matching. It now pages through
every candidate and returns None only on failure; the fallback matches
literally, since a query containing regex metacharacters used to throw.

Empty datasets answer 503 rather than 200-with-nothing or a misleading
404, so callers can tell an outage from an absent school. Adds
/api/release, which reports the build identity baked into the image.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:16:58 +01:00
TudorandClaude Opus 5 38bc17cab3 feat(search): validate the index before the alias points at it
The old sync created a collection, imported batches without reading a
single import response, and swapped the alias regardless. A partial
import published a half-empty index, and two overlapping DAG runs could
prune each other's collections.

Publication now checks every import response and the final document
count before upserting the alias, and holds a session-scoped advisory
lock across the read and the publish so concurrent runs serialise.
Cleanup keeps the previous collection as a rollback pointer and is
best-effort: an uncertain alias response must never delete what might
still be live.

Also parses the Typesense URL properly instead of splitting on colons,
which mangled any host carrying a scheme and a default port.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:16:36 +01:00
tudor dc156058fe Merge pull request 'docs: describe the system that exists, and remove the importer's remains' (#147) from docs/current-truth-and-legacy-inventory into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 22s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m22s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m57s
Reviewed-on: #147
2026-09-15 08:58:54 +00:00
TudorandClaude Opus 5 1d8858fbda chore: remove the code the legacy CSV importer left behind
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m2s
`backend/migration.py` and `scripts/migrate_csv_to_db.py` import `School`,
`SchoolResult`, `init_db` and `set_db_schema_version` — names that no longer
exist. `scripts/geocode_schools.py` imports the same removed ORM model. None of
the three can be imported against the current backend, so they were not dormant
utilities anyone could fall back on; they were files that would fail on the
first line. `backend/version.py` existed only to hand `SCHEMA_VERSION` to that
importer, and the FastAPI lifespan performs no version-triggered import.

Three symbols go with them, each confirmed to have no caller: the unvectorised
`haversine_distance`, superseded by the inline NumPy calculation in search;
`fetcher`, an SWR helper for a dependency this project does not install; and
`kmToMiles`. `calculateDistance` stays — CutoffMapPanel uses it.

Two comments pointed at `migrate_csv_to_db.py --drop` to explain why Payload
owns its own schema. The reason survives the script: blog content must stay
clear of the school marts and Airflow's metadata. Reworded rather than deleted,
so the constraint keeps its justification.

docs/LEGACY_CODE.md records what was removed and where to find it in history. It
also records what was deliberately *not* removed, which is the more useful half:
unused UI components awaiting a design decision, manual data utilities whose
operators a repository search cannot see, and fallbacks that look obsolete but
are load-bearing — `data_loader.py`'s older-mart branches, the generated GIAS
dictionary copies, and the `legacy`-named dbt models that annual DAG selectors
explicitly include. A zero-import count is evidence, not a verdict.

The scripts that fetch DfE CSVs are marked historical and kept, pending
confirmation that nobody runs them by hand.

Checked: 190 backend tests, 429 frontend tests, `tsc --noEmit` clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan
2026-09-14 23:01:22 +01:00
TudorandClaude Opus 5 eaf5e5d180 docs: describe the system that exists, not the one we started with
The README still opened on "Primary School Compass", a KS2 tool for Wandsworth
and Merton served by FastAPI and vanilla JavaScript with Chart.js. Every layer
of that sentence is now wrong: coverage is England-wide across KS2, KS4,
all-through and post-16, Next.js owns the public UI, and school data comes from
dbt-built `marts.*` rather than CSVs loaded at startup. The setup instructions
walked a reader into a virtualenv and a CSV import that cannot build the current
schema, so following the docs produced an empty database and a wrong mental
model at the same time.

Replace the narrative docs with two reference documents that were checked
against the code: docs/ARCHITECTURE.md for request flow, data ownership, the
backend/frontend module boundaries and the real publication sequence, and
docs/DEVELOPMENT.md for the checks that actually run, including the container
and CI version skew that makes "just run pytest" misleading.

The env examples drifted the same way. ALLOWED_ORIGINS is a JSON array, not a
comma-separated list; the frontend needs FASTAPI_URL, DATABASE_URL and
PAYLOAD_SECRET, none of which were documented; and RATE_LIMIT_BURST,
DEFAULT_PAGE_SIZE and MAX_PAGE_SIZE were presented as tuning controls the routes
do not consult. Each is now stated as it behaves.

MIGRATION_SUMMARY.md keeps its content but gains a banner, because it reads like
setup instructions and is not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan
2026-09-14 23:01:15 +01:00
tudor be780ebe13 Merge pull request 'fix(seo): declare the share card, which the route group stopped inheriting' (#146) from fix/og-image-route-group into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m21s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m58s
Reviewed-on: #146
2026-09-14 21:03:14 +00:00
67 changed files with 6073 additions and 2069 deletions

No files matched your search

+13 -5
View File
@@ -20,7 +20,7 @@ PORT=80
# =============================================================================
# CORS
# =============================================================================
# Comma-separated list of allowed origins
# JSON array of allowed origins (pydantic-settings format)
# In production, only include your actual domain
ALLOWED_ORIGINS=["https://schoolcompare.co.uk"]
@@ -33,13 +33,21 @@ ADMIN_API_KEY=CHANGE_THIS_TO_A_SECURE_RANDOM_KEY
# Rate limiting (requests per minute per IP)
RATE_LIMIT_PER_MINUTE=60
RATE_LIMIT_BURST=10
GLOBAL_RATE_LIMIT_PER_MINUTE=3000
# Maximum request body size in bytes (default 1MB)
MAX_REQUEST_SIZE=1048576
# =============================================================================
# API
# SEARCH AND OPTIONAL FEATURE FLAGS
# =============================================================================
DEFAULT_PAGE_SIZE=50
MAX_PAGE_SIZE=100
TYPESENSE_URL=http://localhost:8108
TYPESENSE_API_KEY=CHANGE_THIS_TO_YOUR_TYPESENSE_KEY
# Empty URL disables Unleash-backed flags. Match the managed environment when used.
UNLEASH_URL=
UNLEASH_API_TOKEN=
# Page-size limits are currently declared by route Query parameters.
# DEFAULT_PAGE_SIZE, MAX_PAGE_SIZE and RATE_LIMIT_BURST are not reliable tuning
# controls in the current routes; see docs/LEGACY_CODE.md.
+73 -17
View File
@@ -5,6 +5,11 @@ on:
branches:
- main
# Serialise the entire build/deploy/test cycle: no other run can move staging tags.
concurrency:
group: staging-release
cancel-in-progress: false
env:
REGISTRY: privaterepo.sitaru.org
BACKEND_IMAGE_NAME: ${{ gitea.repository }}-backend
@@ -12,7 +17,18 @@ env:
PIPELINE_IMAGE_NAME: ${{ gitea.repository }}-pipeline
jobs:
prepare:
runs-on: ubuntu-latest
outputs:
build_id: ${{ steps.identity.outputs.build_id }}
steps:
- id: identity
run: python3 -c 'import uuid; print("build_id=" + uuid.uuid4().hex)' >> "$GITHUB_OUTPUT"
build-backend:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Backend (FastAPI)
runs-on: ubuntu-latest
steps:
@@ -46,17 +62,24 @@ jobs:
type=raw,value=staging
- name: Build and push Backend Docker image
id: build
uses: docker/build-push-action@v5
with:
context: .
file: ./Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-backend.outputs.tags }}
labels: ${{ steps.meta-backend.outputs.labels }}
cache-from: type=registry,ref=${{ env.REGISTRY }}/${{ env.BACKEND_IMAGE_NAME }}:buildcache
cache-to: type=registry,ref=${{ env.REGISTRY }}/${{ env.BACKEND_IMAGE_NAME }}:buildcache,mode=max
build-frontend:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Frontend (Next.js)
runs-on: ubuntu-latest
steps:
@@ -90,18 +113,23 @@ jobs:
type=raw,value=staging
- name: Build and push Frontend Docker image
id: build
uses: docker/build-push-action@v5
with:
context: ./nextjs-app
file: ./nextjs-app/Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-frontend.outputs.tags }}
labels: ${{ steps.meta-frontend.outputs.labels }}
build-args: |
FASTAPI_URL=http://backend:80/api
# Cache disabled due to registry size limits
build-pipeline:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Pipeline (Meltano + dbt + Airflow)
runs-on: ubuntu-latest
steps:
@@ -135,11 +163,15 @@ jobs:
type=raw,value=staging
- name: Build and push Pipeline Docker image
id: build
uses: docker/build-push-action@v5
with:
context: ./pipeline
file: ./pipeline/Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-pipeline.outputs.tags }}
labels: ${{ steps.meta-pipeline.outputs.labels }}
cache-from: type=registry,ref=${{ env.REGISTRY }}/${{ env.PIPELINE_IMAGE_NAME }}:buildcache
@@ -148,30 +180,23 @@ jobs:
deploy-staging:
name: Deploy to Staging
runs-on: ubuntu-latest
needs: [build-backend, build-frontend, build-pipeline]
needs: [prepare, build-backend, build-frontend, build-pipeline]
steps:
- name: Trigger staging stack update
run: curl -fsSk -X POST "${{ secrets.PORTAINER_STAGING_WEBHOOK }}"
- name: Wait for staging to become healthy
run: |
echo "Polling ${STAGING_BASE_URL} for up to 5 minutes..."
for i in $(seq 1 60); do
if curl -fsS -o /dev/null --max-time 10 "${STAGING_BASE_URL}/"; then
echo "Staging is up (attempt $i)"
exit 0
fi
sleep 5
done
echo "Staging did not become healthy in time" >&2
exit 1
- uses: actions/checkout@v4
- name: Verify deployed release identity
run: python3 scripts/ci/release.py wait
env:
STAGING_BASE_URL: ${{ secrets.STAGING_BASE_URL }}
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
e2e-staging:
name: E2E Journeys against Staging
runs-on: ubuntu-latest
needs: [deploy-staging]
needs: [prepare, deploy-staging, build-backend, build-frontend, build-pipeline]
steps:
- name: Checkout repository
uses: actions/checkout@v4
@@ -187,11 +212,42 @@ jobs:
npm ci
npx playwright install --with-deps chromium
- name: Verify release before journeys
run: python3 scripts/ci/release.py wait --timeout 10
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- name: Run E2E journeys
working-directory: e2e
run: npx playwright test
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- name: Verify release after journeys
run: python3 scripts/ci/release.py wait --timeout 10
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- uses: docker/setup-buildx-action@v3
- uses: docker/login-action@v3
with:
registry: ${{ env.REGISTRY }}
username: ${{ gitea.actor }}
password: ${{ secrets.REGISTRY_TOKEN }}
- name: Mark tested image digests as verified
run: python3 scripts/ci/release.py verify
env:
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
BACKEND_DIGEST: ${{ needs.build-backend.outputs.digest }}
FRONTEND_DIGEST: ${{ needs.build-frontend.outputs.digest }}
PIPELINE_DIGEST: ${{ needs.build-pipeline.outputs.digest }}
# Production deployment is a second, manual approval: see promote.yml
# ("Promote to Production (manual)") and docs/DEPLOY.md.
+2 -2
View File
@@ -68,13 +68,13 @@ jobs:
python-version: "3.12"
- name: Install dependencies
run: pip install -r requirements.txt pytest "httpx<0.28"
run: pip install -r requirements.txt pytest "httpx<0.28" pyyaml
- name: Import smoke test
run: python -c "from backend.app import app; print('backend imports OK')"
- name: Backend unit tests
run: python -m pytest backend/tests -q
run: python -m pytest backend/tests pipeline/tests scripts/ci/tests -q
build-backend:
name: Build Backend (no push)
+7 -25
View File
@@ -97,33 +97,15 @@ jobs:
username: ${{ gitea.actor }}
password: ${{ secrets.REGISTRY_TOKEN }}
- name: Retag approved images as prod (keeping rollback pointer)
run: |
SHORT_SHA="${{ steps.resolve.outputs.short }}"
for IMAGE in \
"${REGISTRY}/${BACKEND_IMAGE_NAME}" \
"${REGISTRY}/${FRONTEND_IMAGE_NAME}" \
"${REGISTRY}/${PIPELINE_IMAGE_NAME}"; do
# Keep a rollback pointer before moving :prod
docker buildx imagetools create -t "${IMAGE}:prod-previous" "${IMAGE}:prod" || true
docker buildx imagetools create -t "${IMAGE}:prod" "${IMAGE}:${SHORT_SHA}"
echo "Promoted ${IMAGE}:${SHORT_SHA} -> :prod"
done
- name: Resolve verified digests and promote the complete image set
run: python3 scripts/ci/release.py promote --output release.json
env:
EXPECTED_SHA: ${{ steps.resolve.outputs.full }}
- name: Trigger production stack update
run: curl -fsSk -X POST "${{ secrets.PORTAINER_PROD_WEBHOOK }}"
- name: Wait for production to become healthy
run: |
echo "Polling ${PROD_BASE_URL} for up to 5 minutes..."
for i in $(seq 1 60); do
if curl -fsS -o /dev/null --max-time 10 "${PROD_BASE_URL}/"; then
echo "Production is up (attempt $i)"
exit 0
fi
sleep 5
done
echo "Production did not become healthy in time" >&2
exit 1
- name: Verify production release identity
run: python3 scripts/ci/release.py wait --release release.json
env:
PROD_BASE_URL: ${{ secrets.PROD_BASE_URL }}
BASE_URL: ${{ secrets.PROD_BASE_URL }}
+13 -187
View File
@@ -1,191 +1,17 @@
# Docker Deployment Guide
# Docker deployment
## Quick Start
The maintained deployment runbook is [docs/DEPLOY.md](docs/DEPLOY.md).
Deploy the complete SchoolCompare stack (PostgreSQL + FastAPI + Next.js) with one command:
- Production: `docker-compose.portainer.yml`, using `:prod` images.
- Staging: `docker-compose.portainer.staging.yml`, using `:staging` images.
- Builds and deployment: `.gitea/workflows/deploy.yml`.
- Human-approved production promotion: `.gitea/workflows/promote.yml`.
```bash
docker-compose up -d
```
The generic `docker-compose.yml` is not a supported one-command onboarding path:
it still uses `:latest` tags that the release workflow no longer publishes and
lacks the full current CMS setup. Review the [legacy inventory](docs/LEGACY_CODE.md)
before using old compose examples. Starting an empty database does not populate
school marts.
This will start:
- **PostgreSQL** on port 5432 (database)
- **FastAPI** on port 8000 (backend API)
- **Next.js** on port 3000 (frontend)
## Service Details
### PostgreSQL Database
- **Port**: 5432
- **Container**: `schoolcompare_db`
- **Credentials**:
- User: `schoolcompare`
- Password: `schoolcompare`
- Database: `schoolcompare`
- **Volume**: `postgres_data` (persistent storage)
### FastAPI Backend
- **Port**: 8000 → 80 (container)
- **Container**: `schoolcompare_backend`
- **Built from**: Root `Dockerfile`
- **API Endpoint**: http://localhost:8000/api
- **Health Check**: http://localhost:8000/api/data-info
### Next.js Frontend
- **Port**: 3000
- **Container**: `schoolcompare_nextjs`
- **Built from**: `nextjs-app/Dockerfile`
- **URL**: http://localhost:3000
- **Connects to**: Backend via internal network
## Commands
### Start all services
```bash
docker-compose up -d
```
### View logs
```bash
# All services
docker-compose logs -f
# Specific service
docker-compose logs -f nextjs
docker-compose logs -f backend
docker-compose logs -f db
```
### Check status
```bash
docker-compose ps
```
### Stop all services
```bash
docker-compose down
```
### Rebuild after code changes
```bash
# Rebuild and restart specific service
docker-compose up -d --build nextjs
# Rebuild all services
docker-compose up -d --build
```
### Clean restart (remove volumes)
```bash
docker-compose down -v
docker-compose up -d
```
## Initial Database Setup
After first start, you may need to initialize the database:
```bash
# Enter the backend container
docker exec -it schoolcompare_backend bash
# Run migrations or data loading
python -m backend.data_loader
```
## Accessing Services
Once running:
- **Frontend**: http://localhost:3000
- **Backend API**: http://localhost:8000/api
- **API Docs**: http://localhost:8000/docs (Swagger UI)
- **Database**: localhost:5432 (use any PostgreSQL client)
## Environment Variables
Create a `.env` file in the root directory to customize:
```env
# Database
POSTGRES_USER=schoolcompare
POSTGRES_PASSWORD=your_secure_password
POSTGRES_DB=schoolcompare
# Backend
DATABASE_URL=postgresql://schoolcompare:your_secure_password@db:5432/schoolcompare
# Frontend (for client-side access)
NEXT_PUBLIC_API_URL=http://localhost:8000/api
```
Then run:
```bash
docker-compose up -d
```
## Troubleshooting
### Backend not connecting to database
```bash
# Check database health
docker-compose ps
# View backend logs
docker-compose logs backend
# Restart backend
docker-compose restart backend
```
### Frontend not connecting to backend
```bash
# Check backend health
curl http://localhost:8000/api/data-info
# Check Next.js environment variables
docker exec schoolcompare_nextjs env | grep API
```
### Port already in use
```bash
# Change ports in docker-compose.yml
# For example, change "3000:3000" to "3001:3000"
```
### Rebuild from scratch
```bash
docker-compose down -v
docker system prune -a
docker-compose up -d --build
```
## Production Deployment
For production, update the following:
1. **Use secure passwords** in `.env` file
2. **Configure reverse proxy** (Nginx) in front of Next.js
3. **Enable HTTPS** with SSL certificates
4. **Set production environment variables**:
```env
NODE_ENV=production
POSTGRES_PASSWORD=<strong-password>
```
5. **Backup database** regularly:
```bash
docker exec schoolcompare_db pg_dump -U schoolcompare schoolcompare > backup.sql
```
## Network Architecture
```
Internet
↓
Next.js (port 3000) ← User browsers
↓ (internal network)
FastAPI (port 8000) ← API calls
↓ (internal network)
PostgreSQL (port 5432) ← Data queries
```
All services communicate via the `schoolcompare-network` Docker network.
For architecture, configuration and test commands, see
[ARCHITECTURE.md](docs/ARCHITECTURE.md) and [DEVELOPMENT.md](docs/DEVELOPMENT.md).
+6
View File
@@ -24,6 +24,12 @@ RUN pip install --no-cache-dir -r requirements.txt
COPY backend/ ./backend/
COPY scripts/ ./scripts/
ARG BUILD_SHA=development
ARG BUILD_ID=development
LABEL io.schoolcompare.build-id=$BUILD_ID
LABEL io.schoolcompare.commit=$BUILD_SHA
RUN python -c 'import json,sys; open("backend/build-info.json", "w").write(json.dumps({"sha":sys.argv[1],"build_id":sys.argv[2]}))' "$BUILD_SHA" "$BUILD_ID"
# Expose the application port
EXPOSE 80
+2
View File
@@ -1,3 +1,5 @@
> Historical migration record, retained for context. Setup and architecture claims below may be obsolete. Use [README.md](README.md), [architecture](docs/ARCHITECTURE.md) and [deployment](docs/DEPLOY.md) for current guidance.
# SchoolCompare: Vanilla JS → Next.js Migration Summary
## Overview
+52 -199
View File
@@ -1,214 +1,67 @@
# Primary School Compass 🧒📚
# SchoolCompare
A modern web application for comparing **primary school (KS2)** performance data in **Wandsworth and Merton** over the last 5 years. Built with FastAPI and vanilla JavaScript with Chart.js visualizations.
SchoolCompare compares schools across England: primary (KS2), secondary (KS4),
all-through and post-16 provision, with coverage depending on the source dataset.
It provides school search, postcode maps, comparisons, rankings, place pages,
Ofsted information, admissions and destination measures. Editorial content lives
in a Payload CMS blog.
![Python](https://img.shields.io/badge/Python-3.9+-blue)
![FastAPI](https://img.shields.io/badge/FastAPI-0.109-green)
![License](https://img.shields.io/badge/License-MIT-yellow)
## Start here
## Features
- [Architecture and data flow](docs/ARCHITECTURE.md)
- [Development and validation](docs/DEVELOPMENT.md)
- [Deployment and promotion](docs/DEPLOY.md)
- [Legacy and unused-code inventory](docs/LEGACY_CODE.md)
- [Frontend conventions](nextjs-app/README.md)
- [CMS publishing](nextjs-app/docs/PUBLISHING.md)
- 📊 **Interactive Charts** - Visualize KS2 performance trends over time
- 🔍 **Smart Search** - Find primary schools by name in Wandsworth & Merton
- ⚖️ **Side-by-Side Comparison** - Compare up to 5 schools simultaneously
- 🏆 **Rankings** - View top-performing primary schools by various KS2 metrics
- 📱 **Responsive Design** - Works beautifully on desktop and mobile
## Repository map
## Key Metrics (KS2)
| Path | Responsibility |
|---|---|
| `backend/` | FastAPI routes, cached school data, read-only SQLAlchemy mappings, feature flags |
| `nextjs-app/` | Next.js App Router, React UI, Payload CMS, frontend tests |
| `pipeline/plugins/extractors/` | Custom Singer taps for GIAS, EES, Ofsted and other datasets |
| `pipeline/transform/` | dbt staging/intermediate models, marts, seeds and data tests |
| `pipeline/dags/` | Airflow extraction, transformation and publication workflows |
| `pipeline/scripts/` | Search indexing, code generation and operational diagnostics |
| `e2e/` | Playwright journeys against a running environment |
| `.gitea/workflows/` | PR checks, staging deployment and manual production promotion |
| `scripts/` | CI review tooling and historical data utilities; see the legacy inventory |
| `docs/superpowers/`, `mockups/` | Design history and prototypes, not application entry points |
The application tracks these Key Stage 2 performance indicators:
## Runtime
| Metric | Description |
|--------|-------------|
| **Reading Progress** | Progress in reading from KS1 to KS2 |
| **Writing Progress** | Progress in writing from KS1 to KS2 |
| **Maths Progress** | Progress in maths from KS1 to KS2 |
| **Reading Expected %** | Percentage meeting expected standard in reading |
| **Writing Expected %** | Percentage meeting expected standard in writing |
| **Maths Expected %** | Percentage meeting expected standard in maths |
| **Reading, Writing & Maths Combined %** | Percentage meeting expected standard in all three subjects |
The public site is **Next.js**, not the FastAPI root page. Browser `/api/*`
requests pass through a Next.js route handler to FastAPI. Server-rendered pages
call FastAPI directly using `FASTAPI_URL`, including its `/api` suffix.
## Quick Start
PostgreSQL/PostGIS stores school data. Meltano/Singer extracts source data;
dbt builds `marts.*`; FastAPI reads those tables. Typesense serves text search
and autocomplete. Payload runs inside Next.js and owns a separate `payload`
database schema and uploaded media.
### 1. Clone and Setup
There is **no automatic CSV import or sample dataset on startup**. A working
school-data environment needs populated marts from the pipeline or an approved
database snapshot. See [development](docs/DEVELOPMENT.md) before choosing a setup.
```bash
cd school_results
## Validation
# Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
```sh
cd nextjs-app
npm ci
npm run typecheck
npm test -- --runInBand
```
### 2. Run the Application
Backend checks, pipeline validation, runtime versions and E2E requirements are
listed in [DEVELOPMENT.md](docs/DEVELOPMENT.md). No `npm run lint` script is
currently defined.
```bash
# Start the server
python -m uvicorn backend.app:app --reload --port 8000
```
Then open http://localhost:8000 in your browser.
The app will run with **sample data** by default, showing **110 primary schools** (66 in Wandsworth, 44 in Merton) with 5 years of KS2 performance data.
### 3. (Optional) Use Real Data
To use real UK school performance data:
1. Visit [Compare School Performance - Download Data](https://www.compare-school-performance.service.gov.uk/download-data)
2. Download **Key Stage 2** data for the years you want (2019-2024)
- Select "Key Stage 2" as the data type
3. Place the CSV files in the `data/` folder
4. Restart the server - it will automatically load and filter to Wandsworth & Merton schools
**Note:** The app only displays schools in Wandsworth and Merton. Data from other areas will be filtered out.
See the helper script for more details:
```bash
python scripts/download_data.py
```
## Project Structure
```
school_results/
├── backend/
│ └── app.py # FastAPI application with all API endpoints
├── frontend/
│ ├── index.html # Main HTML page
│ ├── styles.css # Styling (warm, editorial design)
│ └── app.js # Frontend JavaScript
├── data/
│ └── .gitkeep # Place CSV data files here
├── scripts/
│ └── download_data.py # Helper for downloading/processing data
├── requirements.txt # Python dependencies
└── README.md
```
## API Endpoints
| Endpoint | Description |
|----------|-------------|
| `GET /api/schools` | List schools with optional search/filter |
| `GET /api/schools/{urn}` | Get detailed data for a specific school |
| `GET /api/compare?urns=...` | Compare multiple schools |
| `GET /api/rankings` | Get school rankings by metric |
| `GET /api/filters` | Get available filter options |
| `GET /api/metrics` | Get available performance metrics |
### Example API Usage
```bash
# Search for schools
curl "http://localhost:8000/api/schools?search=academy"
# Get school details
curl "http://localhost:8000/api/schools/100001"
# Compare schools
curl "http://localhost:8000/api/compare?urns=100001,100002,100003"
# Get rankings
curl "http://localhost:8000/api/rankings?metric=rwm_expected_pct&year=2024"
```
## Data Format
If using your own CSV data, ensure it includes these columns (or similar):
| Column | Type | Description |
|--------|------|-------------|
| URN | Integer | Unique Reference Number |
| SCHNAME | String | School name |
| LA | String | Local Authority (must be Wandsworth or Merton) |
| READPROG | Float | Reading progress score |
| WRITPROG | Float | Writing progress score |
| MATPROG | Float | Maths progress score |
| PTRWM_EXP | Float | % meeting expected standard in reading, writing & maths |
| PTREAD_EXP | Float | % meeting expected standard in reading |
| PTWRIT_EXP | Float | % meeting expected standard in writing |
| PTMAT_EXP | Float | % meeting expected standard in maths |
The application normalizes column names automatically and filters to only show Wandsworth and Merton schools.
## Technology Stack
- **Backend**: FastAPI (Python) - High-performance async API framework
- **Frontend**: Vanilla JavaScript with Chart.js
- **Styling**: Custom CSS with CSS variables for theming
- **Data**: Pandas for CSV processing
## Design Philosophy
The UI features a warm, editorial design inspired by quality publications:
- **Typography**: DM Sans for body text, Playfair Display for headings
- **Color Palette**: Warm cream background with coral and teal accents
- **Interactions**: Smooth animations and hover effects
- **Charts**: Clean, readable data visualizations
## Development
```bash
# Run with auto-reload
python -m uvicorn backend.app:app --reload --port 8000
# Or run directly
python backend/app.py
```
## Coverage
This application is specifically designed for:
- **School Phase**: Primary schools only (Key Stage 2)
- **Geographic Area**: Wandsworth and Merton (London boroughs)
- **Time Period**: Last 5 years of data (2020-2024)
Note: 2021 data shows as unavailable because SATs were cancelled due to COVID-19.
## Data Source
Data is sourced from the UK Government's [Compare School Performance](https://www.compare-school-performance.service.gov.uk/) service, which provides official school performance data for England.
**Important**: When using real data, please comply with the [terms of use](https://www.compare-school-performance.service.gov.uk/download-data) and data protection regulations.
## Scheduled Jobs
### Geocoding Schools (Cron Job)
School postcodes are geocoded by a scheduled job, not on-demand. This improves performance and reduces API calls.
**Setup the cron job** (runs weekly on Sunday at 2am):
```bash
# Edit crontab
crontab -e
# Add this line (adjust paths as needed):
0 2 * * 0 cd /path/to/school_compare && /path/to/venv/bin/python scripts/geocode_schools.py >> /var/log/geocode_schools.log 2>&1
```
**Manual run:**
```bash
# Geocode only schools missing coordinates
python scripts/geocode_schools.py
# Force re-geocode all schools
python scripts/geocode_schools.py --force
```
## License
MIT License - feel free to use this project for educational purposes.
---
Built with ❤️ for Wandsworth & Merton families
## Deployment
Work on a feature branch and open a PR. Merging to `main` builds images and
deploys staging. Production promotion is a separate, human-triggered Gitea
workflow. Use [DEPLOY.md](docs/DEPLOY.md) and the Portainer compose files as the
operational references. The generic compose examples still reference `:latest`,
which the current release workflow does not publish.
+105 -42
View File
@@ -26,7 +26,8 @@ from starlette.middleware.base import BaseHTTPMiddleware
import asyncio
from .config import settings
from .data_loader import (
clear_cache,
build_latest_school_data,
load_school_data_as_dataframe,
compute_benchmarks,
load_school_data,
load_latest_school_data,
@@ -39,20 +40,13 @@ from .data_loader import (
from .data_loader import get_data_info as get_db_info
from . import flags
from .places import build_place_index, build_place_registry, places_for_urn
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .schemas import METRIC_DEFINITIONS, PHASE_GROUPS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .nearby_schools import select_nearby
from .utils import clean_for_json, convert_to_native
# Values to exclude from filter dropdowns (empty strings, non-applicable labels)
EXCLUDED_FILTER_VALUES = {"", "Not applicable", "Does not apply"}
# Maps user-facing phase filter values to the GIAS PhaseOfEducation values they include.
# All-through schools appear in both primary and secondary results.
PHASE_GROUPS: dict[str, set[str]] = {
"primary": {"primary", "middle deemed primary", "all-through"},
"secondary": {"secondary", "middle deemed secondary", "all-through", "16 plus"},
"all-through": {"all-through"},
}
# Must match SITE_URL in nextjs-app/lib/site.ts. The apex 301s to www, and a
# sitemap <loc> that redirects wastes a crawl on every URL it lists.
BASE_URL = "https://www.schoolcompare.co.uk"
@@ -272,7 +266,28 @@ def _places_payload(urn: int) -> list[dict]:
return payload
def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
def _nearby_schools_payload(urn: int) -> list[dict]:
"""The nearest eligible schools this page may offer, closest first.
Phase and reach are read from the school's own row inside select_nearby,
so nothing here can hand it a phase that disagrees with the data.
Wrapped: a failure in selection must never 500 a page that is otherwise
complete, which is the posture get_supplementary_data already takes. The
section simply does not render.
"""
try:
return select_nearby(load_latest_school_data(), int(urn))
except Exception:
import logging
logging.getLogger(__name__).exception(
"Nearby schools selection failed for urn=%s", urn
)
return []
def _place_sitemap_rows(kinds: tuple[str, ...], registry=None) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
@@ -282,7 +297,9 @@ def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
linked from the place page either.
"""
rows: list[str] = []
for p in sorted(get_place_registry().values(), key=lambda p: (p.kind, p.slug)):
if registry is None:
registry = get_place_registry()
for p in sorted(registry.values(), key=lambda p: (p.kind, p.slug)):
if p.kind not in kinds:
continue
rows.append(_url_element(BASE_URL + _place_url(p)))
@@ -296,9 +313,10 @@ def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
return rows
def build_sitemaps() -> dict[str, str]:
def build_sitemaps(df=None, registry=None) -> dict[str, str]:
"""Build the sitemap index and every child, keyed by name."""
df = load_school_data()
if df is None:
df = load_school_data()
children: dict[str, str] = {
"static.xml": _urlset(
@@ -318,7 +336,7 @@ def build_sitemaps() -> dict[str, str]:
# measured apart from the school pages'.
for label, kinds in (("places", ("town", "locality", "authority")),
("outcodes", ("outcode",))):
rows = _place_sitemap_rows(kinds)
rows = _place_sitemap_rows(kinds, registry)
chunks = [rows[i:i + SITEMAP_CHUNK_SIZE]
for i in range(0, len(rows), SITEMAP_CHUNK_SIZE)] or [[]]
for n, chunk in enumerate(chunks, start=1):
@@ -700,6 +718,15 @@ async def get_config():
}
@app.get("/api/release")
async def release_identity():
import json
from pathlib import Path
path = Path(__file__).with_name("build-info.json")
identity = json.loads(path.read_text()) if path.exists() else {"sha": "development", "build_id": "development"}
return JSONResponse(identity, headers={"Cache-Control": "no-store"})
@app.get("/api/schools")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_schools(
@@ -736,7 +763,7 @@ async def get_schools(
df_latest = load_latest_school_data()
if df_latest.empty:
return {"schools": [], "total": 0, "page": page, "page_size": 0}
raise HTTPException(status_code=503, detail="School data temporarily unavailable")
# Use configured default if not specified
if page_size is None:
@@ -835,8 +862,8 @@ async def get_schools(
# Apply filters
if search:
ts_urns = search_schools_typesense(search)
if ts_urns:
ts_urns = await asyncio.to_thread(search_schools_typesense, search)
if ts_urns is not None:
urn_order = {urn: i for i, urn in enumerate(ts_urns)}
schools_df = schools_df[schools_df["urn"].isin(set(ts_urns))].copy()
schools_df["_ts_rank"] = schools_df["urn"].map(urn_order)
@@ -844,9 +871,9 @@ async def get_schools(
else:
# Fallback: Typesense unavailable, use substring match
search_lower = search.lower()
mask = schools_df["school_name"].str.lower().str.contains(search_lower, na=False)
mask = schools_df["school_name"].str.lower().str.contains(search_lower, na=False, regex=False)
if "address" in schools_df.columns:
mask = mask | schools_df["address"].str.lower().str.contains(search_lower, na=False)
mask = mask | schools_df["address"].str.lower().str.contains(search_lower, na=False, regex=False)
schools_df = schools_df[mask]
if local_authority:
@@ -905,7 +932,7 @@ async def get_school_details(request: Request, urn: int):
df = load_school_data()
if df.empty:
raise HTTPException(status_code=404, detail="No data available")
raise HTTPException(status_code=503, detail="School data temporarily unavailable")
school_data = df[df["urn"] == urn]
@@ -970,6 +997,10 @@ async def get_school_details(request: Request, urn: int):
# and authority both fall below the publish threshold has nowhere to
# point, and the page renders without the module.
"places": _places_payload(urn),
# The nearest eligible schools, closest first. Always present on a
# build with this code; the frontend treats absent and empty
# identically, which is what lets the two images deploy independently.
"nearby_schools": _nearby_schools_payload(urn),
"yearly_data": clean_for_json(school_data),
# Supplementary data (null if not yet populated by Kestra)
"ofsted": supplementary.get("ofsted"),
@@ -1542,20 +1573,51 @@ async def get_data_info(request: Request):
}
_publication_lock = asyncio.Lock()
def _prepare_publication(df):
if df.empty:
raise ValueError("Refusing to publish an empty school dataset")
if not {"urn", "year", "school_name"}.issubset(df.columns):
raise ValueError("School dataset is missing required columns")
if df["urn"].isna().any() or df.duplicated(["urn", "year"]).any():
raise ValueError("School dataset has missing URNs or duplicate school years")
latest = build_latest_school_data(df)
registry = build_place_registry(df)
index = build_place_index(registry)
sitemaps = build_sitemaps(df, registry)
return df, latest, registry, index, sitemaps
def _publish(prepared):
# Called on the event loop with no await: routes cannot observe half a swap.
# The application currently runs one worker; replicas require coordination.
from . import data_loader
global _place_registry, _place_index, _place_index_source, _sitemaps
df, latest, registry, index, sitemaps = prepared
data_loader._df_cache = df
data_loader._df_latest_cache = latest
_place_registry = registry
_place_index = index
_place_index_source = registry
_sitemaps = sitemaps
@app.post("/api/admin/reload")
@limiter.limit("5/minute")
async def reload_data(
request: Request,
_: bool = Depends(verify_admin_api_key)
):
"""
Admin endpoint to force data reload (useful after data updates).
Requires X-API-Key header with valid admin API key.
"""
clear_cache()
await asyncio.to_thread(load_school_data)
await asyncio.to_thread(load_latest_school_data)
return {"status": "reloaded"}
async def reload_data(request: Request, _: bool = Depends(verify_admin_api_key)):
"""Validate a complete replacement before publishing it; retain data on failure."""
async with _publication_lock:
try:
df = await asyncio.to_thread(load_school_data_as_dataframe)
prepared = await asyncio.to_thread(_prepare_publication, df)
except Exception as exc:
import logging
logging.getLogger(__name__).exception("Dataset reload failed")
raise HTTPException(status_code=503, detail="Dataset reload failed; previous data retained") from exc
_publish(prepared)
return {"status": "reloaded", "schools": len(prepared[1])}
@@ -1607,15 +1669,16 @@ async def regenerate_sitemap(
request: Request,
_: bool = Depends(verify_admin_api_key),
):
"""Rebuild and cache the sitemap from current school data. Called by Airflow after data updates."""
global _sitemaps, _place_registry
# Places and sitemap are rebuilt together — they read the same marts, and
# letting them drift apart would submit URLs for places that no longer
# exist.
_place_registry = None
_sitemaps = build_sitemaps()
n = sum(x.count("<url>") for x in _sitemaps.values())
return {"status": "ok", "urls": n, "sitemaps": len(_sitemaps)}
"""Rebuild derived publication data without clearing the live registry."""
async with _publication_lock:
try:
prepared = await asyncio.to_thread(_prepare_publication, load_school_data())
except Exception as exc:
raise HTTPException(status_code=503, detail="Sitemap rebuild failed; previous data retained") from exc
_publish(prepared)
n = sum(x.count("<url>") for x in prepared[4].values())
return {"status": "ok", "urls": n, "sitemaps": len(prepared[4])}
# Mount static files directly (must be after all routes to avoid catching API calls)
+55 -24
View File
@@ -84,21 +84,58 @@ def _get_typesense_client():
return None
def search_schools_typesense(query: str, limit: int = 250) -> List[int]:
"""Search Typesense. Returns URNs in relevance order, or [] if unavailable."""
SEARCH_PAGE_SIZE = 250
# Search results are filtered again by the API (authority, phase, postcode,
# etc.), so one page is too small for scoped searches. Keep the candidate set
# bounded, though: a broad query must not turn into an unbounded sequence of
# Typesense requests. Four pages is enough to preserve useful scoped matches
# while putting a hard ceiling on latency and upstream load.
SEARCH_MAX_CANDIDATES = 1_000
def search_schools_typesense(query: str) -> Optional[List[int]]:
"""Return a bounded set of matching URNs in relevance order.
``None`` means Typesense is unavailable; ``[]`` is a valid zero-match
result. The API applies its remaining filters after this search, so the
first few pages are fetched rather than only the first page. Once the
candidate ceiling is reached, the relevance-ordered prefix is returned on
purpose; fetching every match would make common or adversarial queries
unbounded.
"""
client = _get_typesense_client()
if client is None:
return []
return None
urns: list[int] = []
fetched = 0
try:
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority,postcode",
"per_page": min(limit, 250),
"typo_tokens_threshold": 1,
})
return [int(h["document"]["urn"]) for h in result.get("hits", [])]
page = 1
while fetched < SEARCH_MAX_CANDIDATES:
page_size = min(SEARCH_PAGE_SIZE, SEARCH_MAX_CANDIDATES - fetched)
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority,postcode",
"per_page": page_size,
"page": page,
"typo_tokens_threshold": 1,
})
hits = result.get("hits", [])
urns.extend(int(h["document"]["urn"]) for h in hits)
fetched += len(hits)
if fetched >= result.get("found", fetched):
return list(dict.fromkeys(urns))
if not hits:
raise ValueError("Search pagination ended before all matches arrived")
page += 1
logging.getLogger(__name__).info(
"Typesense search capped at %d candidates for query %r",
SEARCH_MAX_CANDIDATES,
query,
)
return list(dict.fromkeys(urns))
except Exception:
return []
logging.getLogger(__name__).exception("School search unavailable")
return None
# The most a public endpoint will return in one response.
@@ -188,16 +225,6 @@ def geocode_single_postcode(postcode: str) -> Optional[Tuple[float, float]]:
return None
def haversine_distance(lat1: float, lon1: float, lat2: float, lon2: float) -> float:
"""Calculate great-circle distance between two points (miles)."""
from math import radians, cos, sin, asin, sqrt
lat1, lon1, lat2, lon2 = map(radians, [lat1, lon1, lat2, lon2])
dlat = lat2 - lat1
dlon = lon2 - lon1
a = sin(dlat / 2) ** 2 + cos(lat1) * cos(lat2) * sin(dlon / 2) ** 2
return 2 * asin(sqrt(a)) * 3956
# =============================================================================
# MAIN DATA LOAD — joins dim_school + dim_location + fact_performance
# fact_performance is a merged KS2+KS4 table (one row per URN per year).
@@ -512,7 +539,12 @@ def load_latest_school_data() -> pd.DataFrame:
if _df_latest_cache is not None:
return _df_latest_cache
df = load_school_data()
_df_latest_cache = build_latest_school_data(load_school_data())
return _df_latest_cache
def build_latest_school_data(df: pd.DataFrame) -> pd.DataFrame:
"""Build a replacement snapshot without mutating the published caches."""
if df.empty:
return df
@@ -545,8 +577,7 @@ def load_latest_school_data() -> pd.DataFrame:
df_latest = pd.concat([df_latest, df_no_perf], ignore_index=True)
print(f"Latest-snapshot cache built: {len(df_latest)} schools")
_df_latest_cache = df_latest
return _df_latest_cache
return df_latest
def clear_cache():
-512
View File
@@ -1,512 +0,0 @@
"""
Database migration logic for importing CSV data.
Used by both CLI script and automatic startup migration.
"""
import re
from pathlib import Path
from typing import Dict, Optional
import numpy as np
import pandas as pd
import requests
from .config import settings
from .database import Base, engine, get_db_session
from .models import School, SchoolResult
from .schemas import (
COLUMN_MAPPINGS,
LA_CODE_TO_NAME,
NULL_VALUES,
SCHOOL_TYPE_MAP,
)
def parse_numeric(value) -> Optional[float]:
"""Parse a numeric value, handling special cases."""
if pd.isna(value):
return None
if isinstance(value, (int, float)):
return float(value) if not np.isnan(value) else None
str_val = str(value).strip().upper()
if str_val in NULL_VALUES or str_val == "":
return None
# Remove percentage signs if present
str_val = str_val.replace("%", "")
try:
return float(str_val)
except ValueError:
return None
def extract_year_from_folder(folder_name: str) -> Optional[int]:
"""Extract year from folder name like '2023-2024'."""
match = re.search(r"(\d{4})-(\d{4})", folder_name)
if match:
return int(match.group(2))
match = re.search(r"(\d{4})", folder_name)
if match:
return int(match.group(1))
return None
def geocode_postcodes_bulk(postcodes: list) -> Dict[str, tuple]:
"""
Geocode postcodes in bulk using postcodes.io API.
Returns dict of postcode -> (latitude, longitude).
"""
results = {}
valid_postcodes = [
p.strip().upper()
for p in postcodes
if p and isinstance(p, str) and len(p.strip()) >= 5
]
valid_postcodes = list(set(valid_postcodes))
if not valid_postcodes:
return results
batch_size = 100
total_batches = (len(valid_postcodes) + batch_size - 1) // batch_size
for i, batch_start in enumerate(range(0, len(valid_postcodes), batch_size)):
batch = valid_postcodes[batch_start : batch_start + batch_size]
print(
f" Geocoding batch {i + 1}/{total_batches} ({len(batch)} postcodes)..."
)
try:
response = requests.post(
"https://api.postcodes.io/postcodes",
json={"postcodes": batch},
timeout=30,
)
if response.status_code == 200:
data = response.json()
for item in data.get("result", []):
if item and item.get("result"):
pc = item["query"].upper()
lat = item["result"].get("latitude")
lon = item["result"].get("longitude")
if lat and lon:
results[pc] = (lat, lon)
except Exception as e:
print(f" Warning: Geocoding batch failed: {e}")
return results
def load_csv_data(data_dir: Path) -> pd.DataFrame:
"""Load all CSV data from data directory."""
all_data = []
for folder in sorted(data_dir.iterdir()):
if not folder.is_dir():
continue
year = extract_year_from_folder(folder.name)
if not year:
continue
# Specifically look for the KS2 results file
ks2_file = folder / "england_ks2final.csv"
if not ks2_file.exists():
continue
csv_file = ks2_file
print(f" Loading {csv_file.name} (year {year})...")
try:
df = pd.read_csv(csv_file, encoding="latin-1", low_memory=False)
except Exception as e:
print(f" Error loading {csv_file}: {e}")
continue
# Rename columns
df.rename(columns=COLUMN_MAPPINGS, inplace=True)
df["year"] = year
# Handle local authority name
la_name_cols = ["LANAME", "LA (name)", "LA_NAME", "LA NAME"]
la_name_col = next((c for c in la_name_cols if c in df.columns), None)
if la_name_col and la_name_col != "local_authority":
df["local_authority"] = df[la_name_col]
elif "LEA" in df.columns:
df["local_authority_code"] = pd.to_numeric(df["LEA"], errors="coerce")
df["local_authority"] = (
df["local_authority_code"]
.map(LA_CODE_TO_NAME)
.fillna(df["LEA"].astype(str))
)
# Store LEA code
if "LEA" in df.columns:
df["local_authority_code"] = pd.to_numeric(df["LEA"], errors="coerce")
# Map school type
if "school_type_code" in df.columns:
df["school_type"] = (
df["school_type_code"]
.map(SCHOOL_TYPE_MAP)
.fillna(df["school_type_code"])
)
# Create combined address
addr_parts = ["address1", "address2", "town", "postcode"]
for col in addr_parts:
if col not in df.columns:
df[col] = None
df["address"] = df.apply(
lambda r: ", ".join(
str(v)
for v in [
r.get("address1"),
r.get("address2"),
r.get("town"),
r.get("postcode"),
]
if pd.notna(v) and str(v).strip()
),
axis=1,
)
all_data.append(df)
print(f" Loaded {len(df)} records")
if all_data:
result = pd.concat(all_data, ignore_index=True)
print(f"\nTotal records loaded: {len(result)}")
print(f"Unique schools: {result['urn'].nunique()}")
print(f"Years: {sorted(result['year'].unique())}")
return result
return pd.DataFrame()
def migrate_data(df: pd.DataFrame, geocode: bool = False, geocode_cache: dict = None):
"""Migrate DataFrame data to database."""
if geocode_cache is None:
geocode_cache = {}
# Clean URN column - convert to integer, drop invalid values
df = df.copy()
df["urn"] = pd.to_numeric(df["urn"], errors="coerce")
df = df.dropna(subset=["urn"])
df["urn"] = df["urn"].astype(int)
# Group by URN to get unique schools (use latest year's data)
school_data = (
df.sort_values("year", ascending=False).groupby("urn").first().reset_index()
)
print(f"\nMigrating {len(school_data)} unique schools...")
# Geocode postcodes that aren't already in the cache
geocoded = dict(geocode_cache) # start with preserved coordinates
if geocode and "postcode" in df.columns:
cached_postcodes = {
str(row.get("postcode", "")).strip().upper()
for _, row in school_data.iterrows()
if int(float(str(row.get("urn", 0) or 0))) in geocode_cache
}
postcodes_needed = [
p for p in df["postcode"].dropna().unique()
if str(p).strip().upper() not in cached_postcodes
]
if postcodes_needed:
print(f"\nGeocoding {len(postcodes_needed)} postcodes ({len(geocode_cache)} restored from cache)...")
fresh = geocode_postcodes_bulk(postcodes_needed)
geocoded.update(fresh)
print(f" Successfully geocoded {len(fresh)} new postcodes")
else:
print(f"\nAll {len(geocode_cache)} postcodes restored from cache, skipping geocoding.")
with get_db_session() as db:
# Create schools
urn_to_school_id = {}
schools_created = 0
for _, row in school_data.iterrows():
# Safely parse URN - handle None, NaN, whitespace, and invalid values
urn_val = row.get("urn")
urn = None
if pd.notna(urn_val):
try:
urn_str = str(urn_val).strip()
if urn_str:
urn = int(float(urn_str)) # Handle "12345.0" format
except (ValueError, TypeError):
pass
if not urn:
continue
# Skip if we've already added this URN (handles duplicates in source data)
if urn in urn_to_school_id:
continue
# Get geocoding data
postcode = row.get("postcode")
lat, lon = None, None
if postcode and pd.notna(postcode):
coords = geocoded.get(str(postcode).strip().upper())
if coords:
lat, lon = coords
# Safely parse local_authority_code
la_code = None
la_code_val = row.get("local_authority_code")
if pd.notna(la_code_val):
try:
la_code_str = str(la_code_val).strip()
if la_code_str:
la_code = int(float(la_code_str))
except (ValueError, TypeError):
pass
school = School(
urn=urn,
school_name=row.get("school_name")
if pd.notna(row.get("school_name"))
else "Unknown",
local_authority=row.get("local_authority")
if pd.notna(row.get("local_authority"))
else None,
local_authority_code=la_code,
school_type=row.get("school_type")
if pd.notna(row.get("school_type"))
else None,
school_type_code=row.get("school_type_code")
if pd.notna(row.get("school_type_code"))
else None,
religious_denomination=row.get("religious_denomination")
if pd.notna(row.get("religious_denomination"))
else None,
age_range=row.get("age_range")
if pd.notna(row.get("age_range"))
else None,
address1=row.get("address1") if pd.notna(row.get("address1")) else None,
address2=row.get("address2") if pd.notna(row.get("address2")) else None,
town=row.get("town") if pd.notna(row.get("town")) else None,
postcode=row.get("postcode") if pd.notna(row.get("postcode")) else None,
latitude=lat,
longitude=lon,
)
db.add(school)
db.flush() # Get the ID
urn_to_school_id[urn] = school.id
schools_created += 1
if schools_created % 1000 == 0:
print(f" Created {schools_created} schools...")
print(f" Created {schools_created} schools")
# Create results
print(f"\nMigrating {len(df)} yearly results...")
results_created = 0
for _, row in df.iterrows():
# Safely parse URN
urn_val = row.get("urn")
urn = None
if pd.notna(urn_val):
try:
urn_str = str(urn_val).strip()
if urn_str:
urn = int(float(urn_str))
except (ValueError, TypeError):
pass
if not urn or urn not in urn_to_school_id:
continue
school_id = urn_to_school_id[urn]
# Safely parse year
year_val = row.get("year")
year = None
if pd.notna(year_val):
try:
year = int(float(str(year_val).strip()))
except (ValueError, TypeError):
pass
if not year:
continue
result = SchoolResult(
school_id=school_id,
year=year,
total_pupils=parse_numeric(row.get("total_pupils")),
eligible_pupils=parse_numeric(row.get("eligible_pupils")),
# Expected Standard
rwm_expected_pct=parse_numeric(row.get("rwm_expected_pct")),
reading_expected_pct=parse_numeric(row.get("reading_expected_pct")),
writing_expected_pct=parse_numeric(row.get("writing_expected_pct")),
maths_expected_pct=parse_numeric(row.get("maths_expected_pct")),
gps_expected_pct=parse_numeric(row.get("gps_expected_pct")),
science_expected_pct=parse_numeric(row.get("science_expected_pct")),
# Higher Standard
rwm_high_pct=parse_numeric(row.get("rwm_high_pct")),
reading_high_pct=parse_numeric(row.get("reading_high_pct")),
writing_high_pct=parse_numeric(row.get("writing_high_pct")),
maths_high_pct=parse_numeric(row.get("maths_high_pct")),
gps_high_pct=parse_numeric(row.get("gps_high_pct")),
# Progress
reading_progress=parse_numeric(row.get("reading_progress")),
writing_progress=parse_numeric(row.get("writing_progress")),
maths_progress=parse_numeric(row.get("maths_progress")),
# Averages
reading_avg_score=parse_numeric(row.get("reading_avg_score")),
maths_avg_score=parse_numeric(row.get("maths_avg_score")),
gps_avg_score=parse_numeric(row.get("gps_avg_score")),
# Context
disadvantaged_pct=parse_numeric(row.get("disadvantaged_pct")),
eal_pct=parse_numeric(row.get("eal_pct")),
sen_support_pct=parse_numeric(row.get("sen_support_pct")),
sen_ehcp_pct=parse_numeric(row.get("sen_ehcp_pct")),
stability_pct=parse_numeric(row.get("stability_pct")),
# Absence
reading_absence_pct=parse_numeric(row.get("reading_absence_pct")),
gps_absence_pct=parse_numeric(row.get("gps_absence_pct")),
maths_absence_pct=parse_numeric(row.get("maths_absence_pct")),
writing_absence_pct=parse_numeric(row.get("writing_absence_pct")),
science_absence_pct=parse_numeric(row.get("science_absence_pct")),
# Gender
rwm_expected_boys_pct=parse_numeric(row.get("rwm_expected_boys_pct")),
rwm_expected_girls_pct=parse_numeric(row.get("rwm_expected_girls_pct")),
rwm_high_boys_pct=parse_numeric(row.get("rwm_high_boys_pct")),
rwm_high_girls_pct=parse_numeric(row.get("rwm_high_girls_pct")),
# Disadvantaged
rwm_expected_disadvantaged_pct=parse_numeric(
row.get("rwm_expected_disadvantaged_pct")
),
rwm_expected_non_disadvantaged_pct=parse_numeric(
row.get("rwm_expected_non_disadvantaged_pct")
),
disadvantaged_gap=parse_numeric(row.get("disadvantaged_gap")),
# 3-Year
rwm_expected_3yr_pct=parse_numeric(row.get("rwm_expected_3yr_pct")),
reading_avg_3yr=parse_numeric(row.get("reading_avg_3yr")),
maths_avg_3yr=parse_numeric(row.get("maths_avg_3yr")),
)
db.add(result)
results_created += 1
if results_created % 10000 == 0:
print(f" Created {results_created} results...")
db.flush()
print(f" Created {results_created} results")
# Commit all changes
db.commit()
print("\nMigration complete!")
def _apply_schema_alterations():
"""
Add new columns to existing tables using ALTER TABLE … ADD COLUMN IF NOT EXISTS.
Safe to run on every migration — no-ops if the column already exists.
Add entries here whenever models.py gains new columns on an existing table.
"""
alterations = [
# v4: Ofsted Report Card columns
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS framework VARCHAR(20)",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_safeguarding_met BOOLEAN",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_inclusion INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_curriculum_teaching INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_achievement INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_attendance_behaviour INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_personal_development INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_leadership_governance INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_early_years INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_sixth_form INTEGER",
]
from sqlalchemy import text as sa_text
with engine.connect() as conn:
for stmt in alterations:
try:
conn.execute(sa_text(stmt))
except Exception as e:
print(f" Warning: alteration skipped ({e})")
conn.commit()
def _apply_schema_drops():
"""
Drop tables retired from the schema. Idempotent (DROP … IF EXISTS), so it's
safe to run on every migration. Add entries here when a model is removed.
"""
drops = [
# v6: Ofsted Parent View feature removed
"DROP TABLE IF EXISTS marts.fact_parent_view CASCADE",
]
from sqlalchemy import text as sa_text
with engine.connect() as conn:
for stmt in drops:
try:
conn.execute(sa_text(stmt))
except Exception as e:
print(f" Warning: drop skipped ({e})")
conn.commit()
def run_full_migration(geocode: bool = False) -> bool:
"""
Run a complete migration: drop all tables and reimport from CSV.
Returns True if successful, False if no data found.
Raises exception on error.
"""
# Preserve existing geocoding so a reimport doesn't throw away coordinates
# that took a long time to compute.
geocode_cache: dict[int, tuple[float, float]] = {}
inspector = __import__("sqlalchemy").inspect(engine)
if "schools" in inspector.get_table_names():
try:
with get_db_session() as db:
rows = db.execute(
__import__("sqlalchemy").text(
"SELECT urn, latitude, longitude FROM schools "
"WHERE latitude IS NOT NULL AND longitude IS NOT NULL"
)
).fetchall()
geocode_cache = {r.urn: (r.latitude, r.longitude) for r in rows}
print(f" Saved {len(geocode_cache)} existing geocoded coordinates.")
except Exception as e:
print(f" Warning: could not save geocode cache: {e}")
# Only drop the core KS2 tables — leave supplementary tables (ofsted, census,
# finance, etc.) intact so a reimport doesn't wipe integrator-populated data.
# schema_version is NOT dropped: it persists so restarts don't re-trigger migration.
ks2_tables = ["school_results", "schools"]
print(f"Dropping core tables: {ks2_tables} ...")
inspector = __import__("sqlalchemy").inspect(engine)
existing = set(inspector.get_table_names())
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
print("Creating all tables...")
Base.metadata.create_all(bind=engine)
# ALTER existing supplementary tables to add any new columns.
# create_all() only creates missing tables; it won't add columns to tables
# that already exist from an older schema version. These statements are
# idempotent (IF NOT EXISTS) so they're safe to run on every migration.
print("Applying column additions to supplementary tables...")
_apply_schema_alterations()
print("Dropping retired tables...")
_apply_schema_drops()
print("\nLoading CSV data...")
df = load_csv_data(settings.data_dir)
if df.empty:
print("Warning: No CSV data found to migrate!")
return False
migrate_data(df, geocode=geocode, geocode_cache=geocode_cache)
return True
+259
View File
@@ -0,0 +1,259 @@
"""Which nearby schools a detail page may offer as alternatives.
HARD FILTERS decide eligibility, and encode claims the section is not allowed
to make. A selective school is not an alternative to a non-selective one, a
special school is not comparable to a mainstream one, and a Girls school is not
an option for a Boys school's reader. They never relax, at any distance, even
where that means the section does not render at all.
DISTANCE decides the order, and nothing else does.
An earlier version ranked by intake similarity first and used distance only as
a tiebreak. That put a Catholic school 2.9 miles away above the community
school 0.3 miles down the road, and — because the row filled from the best tier
before widening — filled all six slots with faith matches while omitting every
school a parent could actually walk to. For a primary, a school that far is not
a weaker option; it is not an option. Distance is a constraint and intake is a
preference, and the ranking now says so.
Similarity survives as `shared`: what a candidate genuinely has in common with
this school, reported on its card, so a reader applies their own weighting
instead of having ours applied for them.
Pure functions over a DataFrame: no I/O, no FastAPI, no database.
"""
from __future__ import annotations
import re
import numpy as np
import pandas as pd
from .schemas import PHASE_GROUPS
# Three fit the row; the rest are behind the carousel arrows.
MAX_SCHOOLS = 6
MINIMUM = 2
# How far the section will reach, in miles, when nothing closer exists.
#
# A sanity bound rather than a target: ordering by distance already handles
# density, so a school in a dense area fills all six slots inside a mile and
# never sees this. It decides one thing — what happens where the area is
# sparse — and the answer differs by phase because catchments do. Primary
# catchments are routinely under a mile; beyond two, a primary is not a weaker
# option but not an option, and no section is the honest answer.
PRIMARY_RADIUS_MILES = 2.0
SECONDARY_RADIUS_MILES = 6.0
POST16_RADIUS_MILES = 10.0
EARTH_RADIUS_MILES = 3958.8
_SPECIAL = re.compile(r"\bspecial\b|pupil referral|alternative provision", re.I)
# Values that mean "this school has no religious character".
_NO_FAITH = {"", "none", "does not apply", "not applicable"}
def is_special_provision(school_type: str | None) -> bool:
"""Mirror of isSpecialSchool() in nextjs-app/lib/utils.ts.
Special schools carry a mainstream phase, so phase alone cannot identify
them. The two implementations must agree: a school the frontend treats as
special for benchmarking but this treats as mainstream would be dropped
from its own England comparison and then offered as a peer to a mainstream
school on the next page along.
"""
return bool(_SPECIAL.search(school_type or ""))
def is_selective(admissions_policy: str | None) -> bool:
"""Strictly selective. Unknown counts as non-selective, which is the safe
direction: it can only ever exclude a pairing, never invent one."""
return (admissions_policy or "").strip().lower() == "selective"
def faith_key(denomination: str | None) -> str:
value = (denomination or "").strip().lower()
return "" if value in _NO_FAITH else value
def faith_label(denomination: str | None) -> str:
return denomination.strip() if faith_key(denomination) else "No religious character"
def genders_compatible(a: str | None, b: str | None) -> bool:
single = {"boys", "girls"}
left, right = (a or "").strip().lower(), (b or "").strip().lower()
return not (left in single and right in single and left != right)
def is_secondary_phase(phase: str | None) -> bool:
"""Whether this phase takes the secondary side: secondary group membership,
minus all-through.
Membership is read from PHASE_GROUPS rather than tested with `"secondary" in
phase`, because that substring misses "16 plus" — GIAS phase 6, which
PHASE_GROUPS deliberately files as secondary. The substring version fails
silently rather than loudly: a sixth-form college is simply handed the
primary bucket and offered infant schools as peers.
All-through is the exception. PHASE_GROUPS lists it on both sides because it
belongs on both phases' place pages, but the detail page renders it with the
primary template, and the metric follows the template.
"""
text = (phase or "").strip().lower()
return text != "all-through" and text in PHASE_GROUPS["secondary"]
def radius_miles(phase: str | None) -> float:
"""How far this phase's section will reach when nothing closer exists."""
if (phase or "").strip().lower() == "16 plus":
return POST16_RADIUS_MILES
return SECONDARY_RADIUS_MILES if is_secondary_phase(phase) else PRIMARY_RADIUS_MILES
def _phase_group(is_secondary: bool) -> set[str]:
return PHASE_GROUPS["secondary" if is_secondary else "primary"]
def _haversine_miles(lat1: float, lon1: float, lat2, lon2):
"""Vectorised, matching the postcode search in app.py."""
lat1_r, lon1_r = np.radians(lat1), np.radians(lon1)
lat2_r, lon2_r = np.radians(lat2.astype(float)), np.radians(lon2.astype(float))
dlat, dlon = lat2_r - lat1_r, lon2_r - lon1_r
a = np.sin(dlat / 2) ** 2 + np.cos(lat1_r) * np.cos(lat2_r) * np.sin(dlon / 2) ** 2
return 2 * EARTH_RADIUS_MILES * np.arcsin(np.sqrt(a))
def _native(value):
"""NaN and numpy scalars both reach JSONResponse badly; normalise here so
the caller never has to remember to."""
if value is None:
return None
if isinstance(value, np.generic):
value = value.item()
if isinstance(value, float) and np.isnan(value):
return None
return value
def _mask(series: pd.Series, predicate) -> pd.Series:
"""A boolean mask that survives an empty frame.
`Series.apply` on an empty Series returns an empty *DataFrame*, and using
that as a mask silently drops every column — so the next column lookup
raises KeyError rather than yielding no rows. This is not hypothetical: a
special school with no special school near it empties the frame at the
provision filter, which is the ordinary case for most special schools.
"""
return pd.Series([predicate(value) for value in series], index=series.index, dtype=bool)
def _shared(subject: pd.Series, candidate: pd.Series, is_secondary: bool) -> list[str]:
"""What this candidate genuinely has in common with the subject.
Empty is a real answer, and renders no chips at all. A card claiming a
shared characteristic it does not have would be worse than a bare one —
and since these no longer affect the order, an empty list costs the school
nothing but its place in the row, which distance already decided.
"""
shared: list[str] = []
gender = str(subject.get("gender") or "").strip()
if gender and str(candidate.get("gender") or "").strip().lower() == gender.lower():
shared.append(gender)
if is_secondary:
policy = str(candidate.get("admissions_policy") or "").strip()
subject_policy = str(subject.get("admissions_policy") or "").strip()
if (
policy
and policy.lower() == subject_policy.lower()
and policy.lower() not in {"not applicable", "unknown"}
):
shared.append(policy)
if faith_key(candidate.get("religious_denomination")) == faith_key(
subject.get("religious_denomination")
):
shared.append(faith_label(candidate.get("religious_denomination")))
return shared
def select_nearby(frame: pd.DataFrame, urn: int) -> list[dict]:
"""The nearest eligible schools, closest first — at most MAX_SCHOOLS, and
none at all below MINIMUM.
The phase is read from the subject's own row rather than passed in, so a
caller cannot hand this a phase that disagrees with the data it selects
from.
"""
subject_rows = frame[frame["urn"] == urn]
if subject_rows.empty:
return []
subject = subject_rows.iloc[0]
lat, lon = _native(subject.get("latitude")), _native(subject.get("longitude"))
if lat is None or lon is None:
return []
phase = subject.get("phase")
is_secondary = is_secondary_phase(phase)
reach = radius_miles(phase)
metric_key = "attainment_8_score" if is_secondary else "rwm_expected_pct"
candidates = frame[frame["urn"] != urn].copy()
for column in ("latitude", "longitude"):
candidates = candidates[candidates[column].notna()]
if candidates.empty:
return []
# ── Hard filters ────────────────────────────────────────────────────
allowed_phases = _phase_group(is_secondary)
candidates = candidates[
candidates["phase"].fillna("").str.lower().isin(allowed_phases)
]
candidates = candidates[candidates["status"].fillna("").str.lower().str.startswith("open")]
subject_special = is_special_provision(subject.get("school_type"))
special = _mask(candidates["school_type"], is_special_provision)
candidates = candidates[special if subject_special else ~special]
subject_selective = is_selective(subject.get("admissions_policy"))
selective = _mask(candidates["admissions_policy"], is_selective)
candidates = candidates[selective if subject_selective else ~selective]
subject_gender = subject.get("gender")
candidates = candidates[
_mask(candidates["gender"], lambda g: genders_compatible(subject_gender, g))
]
if candidates.empty:
return []
candidates["distance_miles"] = _haversine_miles(
lat, lon, candidates["latitude"].values, candidates["longitude"].values
).round(1)
# ── Nearest first, and nothing else has a say ───────────────────────
within = candidates[candidates["distance_miles"] <= reach]
if len(within) < MINIMUM:
return []
selected = within.sort_values(["distance_miles", "urn"]).head(MAX_SCHOOLS)
return [
{
"urn": int(row["urn"]),
"school_name": str(row.get("school_name") or ""),
"distance_miles": float(row["distance_miles"]),
"school_type": _native(row.get("school_type")),
"age_range": _native(row.get("age_range")),
"shared": _shared(subject, row, is_secondary),
"metric_value": _native(row.get(metric_key)),
"metric_key": metric_key,
"metric_year": _native(row.get("year")),
}
for _, row in selected.iterrows()
]
+12
View File
@@ -532,6 +532,18 @@ RANKING_COLUMNS = [
"gcse_grade_91_pct",
]
# Maps user-facing phase filter values to the GIAS PhaseOfEducation values they
# include. All-through schools appear in both primary and secondary results,
# which is why this is a set per phase rather than a single string comparison.
#
# Lives here rather than in app.py because nearby_schools.py needs it too, and
# importing app from there would be a cycle.
PHASE_GROUPS: dict[str, set[str]] = {
"primary": {"primary", "middle deemed primary", "all-through"},
"secondary": {"secondary", "middle deemed secondary", "all-through", "16 plus"},
"all-through": {"all-through"},
}
# School listing columns
SCHOOL_COLUMNS = [
"urn",
+383
View File
@@ -0,0 +1,383 @@
"""Selection rules for the nearby-schools section.
Hard filters encode claims the section is not allowed to make — that a
selective school is an alternative to a non-selective one, that a special
school is comparable to a mainstream one, or that a Girls school is an option
for a Boys school's reader. They decide who is eligible.
Distance decides the order, and nothing else does. An earlier version ranked by
intake similarity first, which put a Catholic school 2.9 miles away above the
community school 0.3 miles down the road — for a primary, a school that far is
not a weaker option, it is not an option. Similarity is now reported on the
card and never reorders the row.
"""
import numpy as np
import pandas as pd
from backend.nearby_schools import (
is_secondary_phase,
radius_miles,
select_nearby,
)
BASE_LAT, BASE_LON = 51.5000, -0.1000
def _row(urn, name, **overrides):
base = {
"urn": urn,
"school_name": name,
"local_authority": "Testshire",
"school_type": "Community school",
"phase": "Primary",
"age_range": "4-11",
"status": "Open",
"gender": "Mixed",
"religious_denomination": "None",
"admissions_policy": "Not applicable",
"latitude": BASE_LAT,
"longitude": BASE_LON,
"year": 202425,
"rwm_expected_pct": 70.0,
"attainment_8_score": np.nan,
}
base.update(overrides)
return base
def _frame(*rows):
return pd.DataFrame(list(rows))
def _at(miles):
"""A latitude `miles` north of BASE_LAT."""
return BASE_LAT + miles / 69.0
# ---------------------------------------------------------------------------
# Order: distance, and only distance
# ---------------------------------------------------------------------------
def test_returns_nearest_first():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Mid", latitude=_at(1.0)),
_row(100003, "Near", latitude=_at(0.4)),
_row(100004, "Far", latitude=_at(1.8)),
)
result = select_nearby(frame, 100001)
assert [s["urn"] for s in result] == [100003, 100002, 100004]
assert result[0]["distance_miles"] == 0.4
def test_a_faith_match_never_outranks_a_closer_school():
"""The reported defect. A Catholic primary surrounded by Catholic primaries
showed six of them and omitted the community school down the road."""
frame = _frame(
_row(100001, "St Jude's RC Primary", religious_denomination="Roman Catholic"),
_row(100002, "Elm Grove Primary", religious_denomination="None", latitude=_at(0.3)),
_row(100003, "Holy Cross RC", religious_denomination="Roman Catholic", latitude=_at(0.8)),
_row(100004, "Sacred Heart RC", religious_denomination="Roman Catholic", latitude=_at(1.2)),
_row(100005, "St Peter's RC", religious_denomination="Roman Catholic", latitude=_at(1.6)),
)
result = select_nearby(frame, 100001)
assert result[0]["urn"] == 100002, "the nearest school leads, whatever its intake"
assert [s["distance_miles"] for s in result] == sorted(s["distance_miles"] for s in result)
def test_the_nearest_eligible_school_is_always_shown():
"""Whatever else changes, a section titled "nearby" cannot omit the nearest
school while listing one four times further away."""
frame = _frame(
_row(100001, "Subject", gender="Boys", religious_denomination="Roman Catholic"),
_row(100002, "Nearest", gender="Mixed", religious_denomination="None", latitude=_at(0.2)),
*[
_row(100010 + n, f"Match {n}", gender="Boys",
religious_denomination="Roman Catholic", latitude=_at(0.9 + n * 0.1))
for n in range(6)
],
)
assert select_nearby(frame, 100001)[0]["urn"] == 100002
def test_caps_at_six_taking_the_nearest():
frame = _frame(
_row(100001, "Subject"),
*[_row(100010 + n, f"Peer {n}", latitude=_at(0.1 * (n + 1))) for n in range(7)],
)
result = select_nearby(frame, 100001)
assert len(result) == 6
assert 100016 not in {s["urn"] for s in result}, "the seventh-nearest is the one dropped"
def test_fewer_than_two_matches_returns_empty():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Only neighbour", latitude=_at(0.5)),
)
assert select_nearby(frame, 100001) == []
def test_excludes_the_subject_school():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
assert 100001 not in {s["urn"] for s in select_nearby(frame, 100001)}
def test_a_school_is_never_listed_twice():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
result = select_nearby(frame, 100001)
assert len(result) == len({s["urn"] for s in result})
# ---------------------------------------------------------------------------
# Reach: a sanity bound, not a target
# ---------------------------------------------------------------------------
def test_primary_does_not_reach_past_two_miles():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Just inside", latitude=_at(1.9)),
_row(100003, "Just outside", latitude=_at(2.4)),
_row(100004, "Miles away", latitude=_at(4.0)),
)
# One inside the cap is below the minimum, so nothing renders at all —
# a primary with nothing within two miles has no nearby schools.
assert select_nearby(frame, 100001) == []
def test_secondary_reaches_further_than_primary():
frame = _frame(
_row(100001, "Subject", phase="Secondary"),
_row(100002, "A", phase="Secondary", latitude=_at(3.0)),
_row(100003, "B", phase="Secondary", latitude=_at(5.5)),
)
assert {s["urn"] for s in select_nearby(frame, 100001)} == {100002, 100003}
def test_the_cap_follows_the_phase():
assert radius_miles("Primary") == 2.0
assert radius_miles("Middle deemed primary") == 2.0
assert radius_miles("All-through") == 2.0
assert radius_miles("Secondary") == 6.0
assert radius_miles("Middle deemed secondary") == 6.0
# Post-16 is the phase people travel furthest for.
assert radius_miles("16 plus") == 10.0
# ---------------------------------------------------------------------------
# Hard filters: eligibility, never order
# ---------------------------------------------------------------------------
def test_selective_never_meets_non_selective():
frame = _frame(
_row(100001, "Grammar", phase="Secondary", admissions_policy="Selective"),
_row(100002, "Comp A", phase="Secondary", admissions_policy="Non-selective", latitude=_at(0.5)),
_row(100003, "Comp B", phase="Secondary", admissions_policy="Non-selective", latitude=_at(0.6)),
)
assert select_nearby(frame, 100001) == []
assert 100001 not in {s["urn"] for s in select_nearby(frame, 100002)}
def test_special_schools_match_only_each_other():
frame = _frame(
_row(100001, "Special", school_type="Community special school"),
_row(100002, "Mainstream A", latitude=_at(0.5)),
_row(100003, "Mainstream B", latitude=_at(0.6)),
)
assert select_nearby(frame, 100001) == []
assert select_nearby(frame, 100002) == []
def test_boys_never_meets_girls():
frame = _frame(
_row(100001, "Boys School", gender="Boys"),
_row(100002, "Girls School", gender="Girls", latitude=_at(0.5)),
_row(100003, "Mixed School", gender="Mixed", latitude=_at(0.6)),
_row(100004, "Another Mixed", gender="Mixed", latitude=_at(0.7)),
)
urns = {s["urn"] for s in select_nearby(frame, 100001)}
assert 100002 not in urns
assert urns == {100003, 100004}
def test_closed_schools_and_missing_coordinates_are_dropped():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Closed", status="Closed", latitude=_at(0.5)),
_row(100003, "No coords", latitude=np.nan, longitude=np.nan),
_row(100004, "Good A", latitude=_at(0.6)),
_row(100005, "Good B", latitude=_at(0.7)),
)
assert {s["urn"] for s in select_nearby(frame, 100001)} == {100004, 100005}
def test_all_through_is_offered_on_both_phase_sides():
frame = _frame(
_row(100001, "Primary subject", phase="Primary"),
_row(100002, "All through", phase="All-through", latitude=_at(0.5)),
_row(100003, "Primary peer", phase="Primary", latitude=_at(0.6)),
)
assert 100002 in {s["urn"] for s in select_nearby(frame, 100001)}
secondary = _frame(
_row(100010, "Secondary subject", phase="Secondary"),
_row(100002, "All through", phase="All-through", latitude=_at(0.5)),
_row(100011, "Secondary peer", phase="Secondary", latitude=_at(0.6)),
)
assert 100002 in {s["urn"] for s in select_nearby(secondary, 100010)}
def test_sixteen_plus_is_matched_against_secondary_not_primary():
"""GIAS phase 6 is "16 plus", and PHASE_GROUPS puts it in the secondary
group — a sixth-form college's peers are secondaries and other colleges,
never primary schools. A substring test for "secondary" misses it silently:
no crash, just a page offering infant schools to a sixth form."""
frame = _frame(
_row(100001, "Sixth Form College", phase="16 plus", age_range="16-19"),
_row(100002, "Nearby Secondary", phase="Secondary", latitude=_at(0.5),
attainment_8_score=52.0),
_row(100003, "Nearby College", phase="16 plus", latitude=_at(0.6)),
_row(100004, "Nearby Primary", phase="Primary", latitude=_at(0.1)),
)
result = select_nearby(frame, 100001)
urns = {s["urn"] for s in result}
assert 100004 not in urns, "a primary school is not a peer for a sixth form"
assert urns == {100002, 100003}
assert all(s["metric_key"] == "attainment_8_score" for s in result)
def test_is_secondary_phase_agrees_with_the_phase_groups_it_selects_from():
for phase in ("Secondary", "Middle deemed secondary", "16 plus"):
assert is_secondary_phase(phase) is True, phase
for phase in ("Primary", "Middle deemed primary", "Nursery", "", None):
assert is_secondary_phase(phase) is False, phase
# In PHASE_GROUPS an all-through school is on both sides, but it renders
# with the primary template, and the metric follows the phase side.
assert is_secondary_phase("All-through") is False
# ---------------------------------------------------------------------------
# What the card reports
# ---------------------------------------------------------------------------
def test_shared_lists_only_what_is_actually_shared():
frame = _frame(
_row(100001, "Subject", phase="Secondary", gender="Mixed",
religious_denomination="None", admissions_policy="Non-selective"),
_row(100002, "Full match", phase="Secondary", gender="Mixed",
religious_denomination="None", admissions_policy="Non-selective", latitude=_at(0.5)),
_row(100003, "Faith differs", phase="Secondary", gender="Mixed",
religious_denomination="Church of England", admissions_policy="Non-selective", latitude=_at(0.6)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert by_urn[100002]["shared"] == ["Mixed", "Non-selective", "No religious character"]
assert by_urn[100003]["shared"] == ["Mixed", "Non-selective"]
def test_a_shared_faith_is_named():
frame = _frame(
_row(100001, "Subject", religious_denomination="Roman Catholic"),
_row(100002, "Also RC", religious_denomination="Roman Catholic", latitude=_at(0.4)),
_row(100003, "Secular", religious_denomination="None", latitude=_at(0.5)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert "Roman Catholic" in by_urn[100002]["shared"]
assert by_urn[100003]["shared"] == ["Mixed"]
def test_shared_is_empty_when_nothing_is_shared():
frame = _frame(
_row(100001, "Subject", gender="Boys", religious_denomination="Roman Catholic"),
_row(100002, "A", gender="Mixed", religious_denomination="None", latitude=_at(0.4)),
_row(100003, "B", gender="Mixed", religious_denomination="Church of England", latitude=_at(0.5)),
)
assert all(s["shared"] == [] for s in select_nearby(frame, 100001))
def test_no_tier_is_reported_because_there_are_no_tiers():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.4)),
_row(100003, "B", latitude=_at(0.5)),
)
assert all("tier" not in s for s in select_nearby(frame, 100001))
def test_metric_follows_the_phase_side_not_the_neighbour():
frame = _frame(
_row(100001, "Subject", phase="Secondary", attainment_8_score=50.0),
_row(100002, "A", phase="Secondary", attainment_8_score=52.8, latitude=_at(0.5)),
_row(100003, "B", phase="Secondary", attainment_8_score=np.nan, latitude=_at(0.6)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert by_urn[100002]["metric_key"] == "attainment_8_score"
assert by_urn[100002]["metric_value"] == 52.8
assert by_urn[100002]["metric_year"] == 202425
assert by_urn[100003]["metric_value"] is None
def test_values_are_json_safe_native_types():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
for school in select_nearby(frame, 100001):
assert isinstance(school["urn"], int)
assert isinstance(school["distance_miles"], float)
assert not isinstance(school["metric_value"], np.generic)
# ---------------------------------------------------------------------------
# The endpoint
# ---------------------------------------------------------------------------
import pytest
from fastapi.testclient import TestClient
def _endpoint_frame():
return _frame(
_row(100001, "Subject Primary"),
_row(100002, "Neighbour A", latitude=_at(0.5)),
_row(100003, "Neighbour B", latitude=_at(0.6)),
)
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_latest_school_data", _endpoint_frame)
monkeypatch.setattr(app_module, "load_school_data", _endpoint_frame)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
return TestClient(app_module.app, raise_server_exceptions=False)
def test_detail_payload_carries_nearby_schools(client):
resp = client.get("/api/schools/100001")
assert resp.status_code == 200, resp.text
similar = resp.json()["nearby_schools"]
assert [s["school_name"] for s in similar] == ["Neighbour A", "Neighbour B"]
assert similar[0]["metric_key"] == "rwm_expected_pct"
def test_a_failure_in_selection_does_not_break_the_page(client, monkeypatch):
from backend import app as app_module
def _explode(*args, **kwargs):
raise ValueError("selection blew up")
monkeypatch.setattr(app_module, "select_nearby", _explode)
resp = client.get("/api/schools/100001")
assert resp.status_code == 200, resp.text
assert resp.json()["nearby_schools"] == []
+68
View File
@@ -0,0 +1,68 @@
"""Publication must preserve the current dataset until every replacement is ready."""
import asyncio
import pandas as pd
import pytest
from fastapi.testclient import TestClient
from backend import app as api, data_loader
from backend.tests.test_sixth_form_flag import _schools_df
@pytest.fixture
def client(monkeypatch):
old = _schools_df()
monkeypatch.setattr(data_loader, '_df_cache', old)
monkeypatch.setattr(data_loader, '_df_latest_cache', old)
monkeypatch.setattr(api, '_place_registry', {'old': 'registry'})
monkeypatch.setattr(api, '_place_index', {'old': 'index'})
monkeypatch.setattr(api, '_place_index_source', api._place_registry)
monkeypatch.setattr(api, '_sitemaps', {'old.xml': 'old sitemap'})
monkeypatch.setattr(api, '_publication_lock', asyncio.Lock())
monkeypatch.setattr(api.limiter, 'enabled', False)
api.app.dependency_overrides[api.verify_admin_api_key] = lambda: True
yield TestClient(api.app, raise_server_exceptions=False)
api.app.dependency_overrides.clear()
def state():
return (data_loader._df_cache, data_loader._df_latest_cache, api._place_registry,
api._place_index, api._place_index_source, api._sitemaps)
@pytest.mark.parametrize('failure', ['empty', 'database', 'sitemap', 'duplicate'])
def test_failed_reload_preserves_every_published_object(client, monkeypatch, failure):
before = state()
df = _schools_df()
if failure == 'empty':
df = pd.DataFrame()
if failure == 'duplicate':
df = pd.concat([df, df.iloc[:1]], ignore_index=True)
def load():
if failure == 'database':
raise RuntimeError('database unavailable')
return df
monkeypatch.setattr(api, 'load_school_data_as_dataframe', load)
if failure == 'sitemap':
monkeypatch.setattr(api, 'build_sitemaps', lambda *args: (_ for _ in ()).throw(RuntimeError('bad XML')))
response = client.post('/api/admin/reload')
assert response.status_code == 503
assert all(a is b for a, b in zip(before, state()))
def test_success_publishes_school_data_places_and_sitemaps(client, monkeypatch):
df = _schools_df()
df.loc[0, 'school_name'] = 'Replacement School'
monkeypatch.setattr(api, 'load_school_data_as_dataframe', lambda: df)
response = client.post('/api/admin/reload')
assert response.status_code == 200
assert data_loader.load_school_data() is df
assert data_loader.load_latest_school_data().iloc[0].school_name == 'Replacement School'
assert api._place_index_source is api._place_registry
assert 'old.xml' not in api._sitemaps
assert 'replacement-school' in api._sitemaps['schools-1.xml']
def test_failed_sitemap_regeneration_keeps_existing_publication(client, monkeypatch):
before = state()
monkeypatch.setattr(api, 'build_sitemaps', lambda *args: (_ for _ in ()).throw(RuntimeError('bad XML')))
assert client.post('/api/admin/regenerate-sitemap').status_code == 503
assert all(a is b for a, b in zip(before, state()))
+114
View File
@@ -0,0 +1,114 @@
from types import SimpleNamespace
import pytest
from fastapi.testclient import TestClient
from backend import app as api, data_loader
from backend.tests.test_sixth_form_flag import _schools_df
def client_for(monkeypatch, search):
client = SimpleNamespace(collections={'schools': SimpleNamespace(documents=SimpleNamespace(search=search))})
monkeypatch.setattr(data_loader, '_get_typesense_client', lambda: client)
def test_search_returns_matches_beyond_first_page(monkeypatch):
pages = []
def search(params):
pages.append(params['page'])
urns = range(100000, 100250) if params['page'] == 1 else [100999]
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in urns]}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('academy')
assert len(result) == 251
assert result[-1] == 100999
assert pages == [1, 2]
def test_search_caps_broad_queries_at_a_bounded_number_of_pages(monkeypatch):
requests = []
def search(params):
requests.append(params)
return {
'found': 10_000,
'hits': [
{'document': {'urn': 100000 + params['page'] * 1000 + i}}
for i in range(params['per_page'])
],
}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('school')
assert len(result) == data_loader.SEARCH_MAX_CANDIDATES
assert len(requests) == data_loader.SEARCH_MAX_CANDIDATES // data_loader.SEARCH_PAGE_SIZE
assert all(request['per_page'] == data_loader.SEARCH_PAGE_SIZE for request in requests)
assert requests[-1]['page'] == len(requests)
def test_search_uses_a_smaller_final_page_when_the_cap_is_not_a_page_multiple(monkeypatch):
monkeypatch.setattr(data_loader, 'SEARCH_MAX_CANDIDATES', 251)
requests = []
def search(params):
requests.append(params)
return {
'found': 10_000,
'hits': [{'document': {'urn': 100000 + len(requests) * 1000 + i}}
for i in range(params['per_page'])],
}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('school')
assert len(result) == 251
assert [request['per_page'] for request in requests] == [250, 1]
def test_later_page_failure_does_not_return_partial_results(monkeypatch):
def search(params):
if params['page'] == 2:
raise RuntimeError('timeout')
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in range(100000, 100250)]}
client_for(monkeypatch, search)
assert data_loader.search_schools_typesense('academy') is None
def test_zero_matches_are_distinct_from_unavailable(monkeypatch):
client_for(monkeypatch, lambda _: {'found': 0, 'hits': []})
assert data_loader.search_schools_typesense('academy') == []
monkeypatch.setattr(data_loader, '_get_typesense_client', lambda: None)
assert data_loader.search_schools_typesense('academy') is None
@pytest.mark.parametrize('matches, expected', [([], []), (None, [100001])])
def test_fallback_only_on_dependency_failure(monkeypatch, matches, expected):
monkeypatch.setattr(api.limiter, 'enabled', False)
monkeypatch.setattr(api, 'load_latest_school_data', _schools_df)
monkeypatch.setattr(api, 'search_schools_typesense', lambda _: matches)
response = TestClient(api.app).get('/api/schools?search=Alpha')
assert response.status_code == 200
assert [s['urn'] for s in response.json()['schools']] == expected
def test_filtered_api_keeps_match_from_second_search_page(monkeypatch):
monkeypatch.setattr(api.limiter, 'enabled', False)
df = _schools_df()
monkeypatch.setattr(api, 'load_latest_school_data', lambda: df)
def search(params):
urns = range(200000, 200250) if params['page'] == 1 else [100001]
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in urns]}
client_for(monkeypatch, search)
response = TestClient(api.app).get('/api/schools?search=Alpha&local_authority=Testshire')
assert response.status_code == 200
assert response.json()['total'] == 1
assert response.json()['schools'][0]['urn'] == 100001
def test_unavailable_dataset_is_not_a_missing_school_or_empty_search(monkeypatch):
import pandas as pd
monkeypatch.setattr(api.limiter, 'enabled', False)
monkeypatch.setattr(api, 'load_school_data', lambda: pd.DataFrame())
monkeypatch.setattr(api, 'load_latest_school_data', lambda: pd.DataFrame())
client = TestClient(api.app)
assert client.get('/api/schools/100001').status_code == 503
assert client.get('/api/schools?search=school').status_code == 503
-26
View File
@@ -1,26 +0,0 @@
"""
Schema versioning for database migrations.
HOW TO USE:
- Bump SCHEMA_VERSION when making changes to database models
- This triggers an automatic full data reimport on next app startup
WHEN TO BUMP:
- Adding/removing columns in models.py
- Changing column types or constraints
- Modifying CSV column mappings in schemas.py
- Any change that requires fresh data import
"""
# Current schema version - increment when models change
SCHEMA_VERSION = 6
# Changelog for documentation
SCHEMA_CHANGELOG = {
1: "Initial schema with School and SchoolResult tables",
2: "Added pupil absence fields (reading, maths, gps, writing, science)",
3: "Added supplementary data tables: ofsted, parent_view, census, admissions, sen_detail, phonics, deprivation, finance; GIAS columns on schools",
4: "Added Ofsted Report Card columns to ofsted_inspections (new framework from Nov 2025)",
5: "Apply ALTER TABLE additions for RC columns missed by create_all on existing tables",
6: "Removed the Ofsted Parent View feature: dropped fact_parent_view table and model",
}
+29 -172
View File
@@ -1,180 +1,37 @@
# SchoolCompare.co.uk - Project Context
# SchoolCompare project context
## Overview
## Maintained documentation
SchoolCompare is a web application for comparing UK primary school (KS2) performance data. It allows users to:
- Search and browse schools by name, location (postcode), or local authority
- Compare multiple schools side-by-side with charts and tables
- View school rankings by various KS2 metrics
- See historical performance trends across years
Read [README.md](README.md), [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) and
[docs/DEVELOPMENT.md](docs/DEVELOPMENT.md) for the current implementation.
[docs/LEGACY_CODE.md](docs/LEGACY_CODE.md) records obsolete paths and deliberate
compatibility code. Historical design documents are not current setup instructions.
## Architecture
## Architecture constraints
### Backend (Python/FastAPI)
- **Framework**: FastAPI with uvicorn
- **Database**: PostgreSQL with SQLAlchemy ORM
- **Data Source**: UK Government "Compare School Performance" CSV downloads
Key files:
- `backend/app.py` - Main FastAPI application, API routes
- `backend/config.py` - Configuration via pydantic-settings (env vars, .env file)
- `backend/database.py` - SQLAlchemy engine, session management
- `backend/models.py` - Database models (School, SchoolResult)
- `backend/data_loader.py` - Data queries, geocoding, legacy DataFrame compatibility
- `backend/schemas.py` - Column mappings, metric definitions, LA code mappings
### Content / CMS (Payload)
Payload CMS runs **inside** the Next.js app — one image, one container, no
separate service. It powers `/blog`; `/about` is a plain coded page.
- **Admin panel:** `/admin`. The only authenticated surface on the site.
`noindex` via both `robots.txt` and `X-Robots-Tag`.
- **CMS API:** `/cms-api`, **not** `/api`. `/api/*` is a catch-all proxy to
FastAPI (`app/(frontend)/api/[...path]`) which would silently swallow every
admin call and forward it to the backend. Mount points are defined once in
`lib/payloadRoutes.ts`.
- **Database:** the existing Postgres, in its own `payload` schema, so no
pipeline operation on `public` — including
`scripts/migrate_csv_to_db.py --drop` — can reach blog content.
- **Uploads:** the `payload_media` Docker volume at `/app/media`. Not
reproducible from the pipeline; must be backed up.
- **New env vars:** `DATABASE_URL` and `PAYLOAD_SECRET` on the frontend service.
Staging must use a different `PAYLOAD_SECRET` from production.
- Publishing workflow and house style: `nextjs-app/docs/PUBLISHING.md`.
- **Admin field components resolve through a generated import map**
(`app/(payload)/admin/importMap.js`). Payload hands the client a *path* per
field and looks it up there; a missing entry renders no field and reports no
error, while `required` still blocks the save. After adding or changing any
field, editor or lexical feature, run `npm run generate:importmap` in
`nextjs-app/` and commit the result.
### Two route groups
`nextjs-app/app/` has no root `layout.tsx`. It cannot: Payload's admin panel
ships its own root layout rendering `<html>`/`<body>`, and Next permits
multiple root layouts only when no `app/layout.tsx` exists.
- `app/(frontend)/` — the site. Its `layout.tsx` is the site's root layout.
- `app/(payload)/` — the admin panel and `/cms-api`.
Route groups are invisible to routing, so every public URL is unchanged.
**The metadata file conventions stay at the `app/` root** — `robots.ts`,
`opengraph-image.tsx`, `icon.png`, `apple-icon.png`. Inside a route group Next
treats them as segment-scoped: it renames `/icon.png` to `/icon-<hash>.png` and
drops `/robots.txt` entirely. Route handlers are unaffected.
The build must succeed with `DATABASE_URL` unset, because CI builds it that
way. Never call `getCachedPayload()` at module scope, and never add
`generateStaticParams` to a DB-backed route.
### Frontend (Vanilla JS)
- Single-page application with hash-based routing
- Chart.js for data visualization
- No build step required
Key files:
- `frontend/index.html` - Main HTML structure
- `frontend/app.js` - All application logic, API calls, rendering
- `frontend/styles.css` - Styling (CSS variables, responsive design)
### Database Schema
```
schools school_results
├── id (PK) ├── id (PK)
├── urn (unique, indexed) ├── school_id (FK → schools.id)
├── school_name ├── year (indexed)
├── local_authority ├── rwm_expected_pct
├── school_type ├── reading_expected_pct
├── postcode ├── ... (all KS2 metrics)
├── latitude, longitude └── unique(school_id, year)
└── results → SchoolResult[]
```
## Configuration
Environment variables (or `.env` file):
- `DATABASE_URL` - PostgreSQL connection string (default: `postgresql://schoolcompare:schoolcompare@localhost:5432/schoolcompare`)
- `HOST`, `PORT` - Server binding (default: `0.0.0.0:80`)
- `ALLOWED_ORIGINS` - CORS origins
## Running Locally
1. Start PostgreSQL:
```bash
docker compose up -d db
```
2. Run migration to import CSV data:
```bash
python scripts/migrate_csv_to_db.py --drop
# Add --geocode to geocode postcodes (slower, adds lat/long)
```
3. Start the app:
```bash
uvicorn backend.app:app --host 0.0.0.0 --port 8000
```
## Docker Deployment
```bash
docker compose up -d
```
This starts:
- `db` - PostgreSQL 16 with persistent volume
- `app` - FastAPI application on port 80
## Data
- Source: UK Government Compare School Performance downloads
- Location: `data/` directory with year folders (e.g., `2023-2024/england_ks2final.csv`)
- The `scripts/download_data.py` can fetch data from the government website
## Key Features
- **Location Search**: Enter postcode to find nearby schools (uses postcodes.io API)
- **Multi-school Comparison**: Select multiple schools, view metrics across years
- **Rankings**: Top schools by any KS2 metric, filterable by local authority
- **Variability Analysis**: Shows standard deviation of scores across years
## API Endpoints
- `GET /api/schools` - List/search schools (supports pagination, location search)
- `GET /api/schools/{urn}` - School details with all yearly data
- `GET /api/compare?urns=123,456` - Compare multiple schools
- `GET /api/rankings` - School rankings by metric
- `GET /api/filters` - Available filter options (LAs, types, years)
- `GET /api/metrics` - Metric definitions (single source of truth)
- `GET /api/data-info` - Database stats
- Next.js serves the public UI. FastAPI serves school data from dbt-built `marts.*`.
The backend does not create school tables or import CSVs at startup.
- School coverage spans England and multiple phases, not only primary schools in
Wandsworth and Merton.
- `/api/*` belongs to the FastAPI proxy. Payload uses `/cms-api` and `/admin`.
- Payload runs inside Next.js, with its own `payload` schema and persistent media.
Keep CMS migrations independent of school-data transformations.
- Public and Payload route groups have separate root layouts. Do not introduce
`app/layout.tsx`. Keep site-wide metadata files at the `app/` root.
- Builds must succeed with `DATABASE_URL` unset. Do not call `getCachedPayload()`
at module scope or add DB-backed `generateStaticParams`.
- After changing CMS fields/editors, run `npm run generate:importmap` and commit
the generated import map. See `nextjs-app/docs/PUBLISHING.md`.
- The backend and pipeline GIAS dictionary copies are generated together; preserve
their parity. Tests enforce it.
## SDLC
Full details in `docs/DEPLOY.md`. The short version:
Follow [docs/DEPLOY.md](docs/DEPLOY.md).
- **Never push to `main` directly.** Work on a feature branch and open a PR;
branch protection requires the PR checks (typecheck, tests, builds, AI review)
to pass before merge.
- Merging to `main` deploys automatically **to staging only**: images are
built once, deployed to the staging Portainer stack, and verified by the
Playwright journeys in `e2e/`. Production is a second, manual approval:
the "Promote to Production (manual)" workflow in Gitea Actions, run after
testing the feature on staging. It refuses commits whose staging E2E gate
isn't green. Never trigger it yourself — promotion is the human's call.
- If you change user-facing behaviour, update or extend the `e2e/` journey
tests in the same PR — they gate whether staging is fit for human testing
and whether a commit is promotable.
## Recent Changes
- Added staging environment + automated staging→prod pipeline (Gitea Actions)
- Migrated from CSV file storage to PostgreSQL database
- Added location-based search using postcode geocoding
- Added local authority filter to rankings
- Improved frontend with featured schools, loading states, API caching
# Important
- Do not attempt to start a local server to test the application, it does not work
- Never push directly to `main`. Use a feature branch and a PR with passing checks.
- Merges deploy staging only. Production promotion is a separate human decision;
do not trigger the promotion workflow yourself.
- Update E2E journeys in the same PR when changing user-facing behaviour.
- Do not attempt to start a local server to test the application; use unit checks
and the configured integration environment.
+125
View File
@@ -0,0 +1,125 @@
# Architecture
This describes the implementation as reviewed on 2026-09-14. It distinguishes
current behaviour from improvements still to be implemented.
## Request flow
```text
Browser → Next.js public routes
├─ /api/* proxy → FastAPI → cached DataFrames / PostgreSQL marts
│ ├─ Typesense (search and suggestions)
│ └─ postcodes.io (postcode lookup)
└─ /admin, /cms-api, /blog → Payload → payload schema + media volume
Next.js server rendering → FastAPI directly through FASTAPI_URL
```
`nextjs-app/lib/api.ts` contains typed fetch wrappers and revalidation defaults.
The proxy is `nextjs-app/app/(frontend)/api/[...path]/route.ts`. Payload uses
`/cms-api` so its routes do not collide with the FastAPI proxy. The proxy denies
`/api/flags`; server-side rendering reads flags directly from FastAPI.
## Data ownership
| Layer | Owner and role |
|---|---|
| Source data | GIAS, DfE EES, Ofsted, finance, deprivation and council admission-distance sources |
| `raw` | Singer taps and the PostgreSQL target configured in `pipeline/meltano.yml` |
| Staging/intermediate/marts | dbt models in `pipeline/transform`; marts are materialized tables |
| `marts.dim_school`, `marts.dim_location` | School identity and location, filtered to supported England establishments |
| `marts.fact_*` | Performance and supplementary datasets; coverage and years vary |
| Typesense `schools` alias | Search documents built by `pipeline/scripts/sync_typesense.py` |
| `payload` | CMS collections and migrations in `nextjs-app/`; independent of dbt |
| Media volume | Uploaded blog media; requires backup and cannot be regenerated from school datasets |
`backend/models.py` maps existing marts for reading. It does not create the school
schema. There is no startup schema-version migration or CSV reimport. Payload's
`nextjs-app/migrations/` is active and must not be confused with the removed
legacy backend migration code.
Coordinates normally come from GIAS British National Grid coordinates transformed
by PostGIS in `dim_location.sql`. `pipeline/scripts/geocode_postcodes.py` is a
manual fallback utility, not a task wired into the current school-data DAG.
Backend postcode searches also use postcodes.io; that lookup does not populate
school coordinates in the database.
## Backend boundaries
- `app.py`: routes, middleware, search filtering, sitemap/place publication and response assembly.
- `data_loader.py`: SQL loading, process-local DataFrame caches, Typesense calls,
postcode lookups, supplementary queries and benchmark calculation.
- `database.py`: synchronous SQLAlchemy engine and sessions.
- `schemas.py`: metric definitions, column mappings and display metadata; despite
its name this is not a collection of Pydantic API response models.
- `places.py` and `localities.py`: place registry and curated locality information.
- `flags.py`: Unleash-backed feature flags, disabled when no server is configured.
- `gias_codes.py` / `ofsted_codes.py`: source-code translation and display rules.
Search starts from a cached latest-row-per-school snapshot. Detail pages read
history from the full DataFrame and supplementary data from marts. Comparisons
batch supplementary queries across selected URNs. Async routes still contain
synchronous dependency calls; a fully asynchronous database layer is not present.
## Frontend boundaries
`app/(frontend)` owns the public root layout and pages. `app/(payload)` owns the
CMS root layout. Do not add a shared `app/layout.tsx`: these groups deliberately
have separate root layouts. Root metadata files remain in `app/`.
Server pages fetch initial data and pass it to client views. Client state uses
React hooks, URL search parameters and the comparison context/localStorage.
There is no SWR dependency. Leaflet maps are loaded through dynamic wrappers;
Chart.js renders performance and comparison charts.
`components/school/` contains detail sections, with section decisions and data
preparation in `lib/schoolSections.ts`. The nearby-schools section is selected in
`backend/nearby_schools.py` — hard filters decide eligibility (phase, provision,
selectivity, gender) and distance alone decides the order, capped per phase —
and served on `/api/schools/{urn}`. Its rules are presentation logic,
deliberately kept out of `marts.*` so they can be tuned by deploy rather than by
pipeline run. `lib/types.ts` contains manually maintained
API types. `payload-types.ts` and the Payload import map are generated artifacts.
## Publication and caching today
1. Airflow DAGs extract and validate source data, then run selected dbt builds.
2. Relevant DAGs rebuild Typesense and swap the `schools` alias.
3. They call `POST /api/admin/reload` with `X-API-Key`. It builds and validates
replacement DataFrames, places, reverse membership and sitemaps off the request
loop, then publishes them together. Failure returns 503 and preserves live data.
4. A separate weekly sitemap DAG can regenerate the derived publication from the
current DataFrame without clearing the live registry first.
GIAS is scheduled daily, Ofsted monthly, and annual datasets are manually
triggered. The DAG definitions are authoritative for selectors and dependencies.
Caches exist in several independent layers: backend DataFrames and registries,
backend HTTP Cache-Control/ETags, Next.js fetch/page revalidation, and browser or
shared HTTP caches where configured. Place fetches request a one-week revalidation
interval. HTTP ETags are computed after route execution, not before database work.
Typesense publication validates every import response and the final document
count before switching aliases. A session-scoped PostgreSQL advisory lock
serialises index reads/publication across DAGs. The previous collection remains
available for rollback; old unaliased collections are pruned after success.
Failed drafts are retained until a later successful cleanup, because an uncertain
alias-update response must never cause deletion of a potentially live index.
The backend snapshot swap is process-local and assumes the current single-worker
deployment. It is not an atomic transaction spanning PostgreSQL marts, Typesense
and Next.js caches. Next.js caches are not explicitly purged by the pipeline.
School search retrieves a relevance-ordered candidate prefix (currently capped at
1,000 URNs) before applying API filters. This keeps scoped searches useful while
putting a hard ceiling on Typesense round trips; only a dependency failure invokes
substring fallback, not a valid empty match set.
## Deployment references
See [DEPLOY.md](DEPLOY.md). PR checks include frontend typechecking/tests, backend
unit tests, image builds and AI review. Staging journeys run after merging.
Staging runs are serialised across builds, deployment and E2E. Build-stamped
frontend/backend identities are checked before and after journeys. Only then are
the captured image digests marked verified. Promotion resolves and validates the
complete verified image set before retagging production. See the runbook for
first-rollout requirements and remaining integration checks.
+53 -4
View File
@@ -19,8 +19,9 @@ PR checks (.gitea/workflows/pr-checks.yml)
▼
Stage pipeline (.gitea/workflows/deploy.yml) — automatic
1. build & push images → tags sha-<sha>, staging
2. staging Portainer webhook → wait for staging health
2. staging Portainer webhook → verify frontend/backend SHA + build ID
3. Playwright E2E journeys against staging ← gate before human testing
4. verify identity again; tag tested digests verified-<full-sha>
▼
Manual testing on staging (stx.schoolcompare.co.uk)
│ Actions → "Promote to Production (manual)" ← approval #2
@@ -28,14 +29,15 @@ Manual testing on staging (stx.schoolcompare.co.uk)
Promote pipeline (.gitea/workflows/promote.yml) — manual dispatch
1. resolve target sha (input, or latest main if empty)
2. REFUSE unless that commit's "E2E Journeys against Staging" status is green
3. retag sha-<sha> → :prod (same bytes — build once, promote the image)
3. resolve verified-<full-sha> digests, validate labels, retag digests → :prod
previous :prod saved as :prod-previous
4. prod Portainer webhook → wait for prod health
4. prod Portainer webhook → verify expected SHA + build ID
```
Key principle: **build once, promote the exact image**. Production pins `:prod`,
which only moves when a human runs the promote workflow — and the workflow
only accepts commits that passed the staging E2E gate. Nothing tags `:latest`
only accepts commits that passed the staging E2E gate and have a complete verified
image set. Nothing tags `:latest`
anymore.
## Branch & PR workflow
@@ -256,3 +258,50 @@ how long any feature is exposed to this.
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
attempted. That is the correct behaviour for local development and CI, and it
means the test suites need no flag server.
## Release identity and the P1 reliability gate
Every staging run creates a random build ID before building its three images.
Each image carries the commit and build ID as labels. Frontend/backend images
also contain a build-time JSON file; environment overrides cannot rewrite it.
`/release.json` returns both identities with `Cache-Control: no-store`. It fails
with 503 when either identity cannot be read. FastAPI's internal endpoint is
`/api/release`.
The entire staging workflow shares one concurrency group, with cancellation
disabled. This needs Gitea 1.26 or newer, where workflow concurrency is supported
([release notes](https://blog.gitea.com/release-of-1.26.0/)); the configured server
reported 1.27.3 during this change. Do not run the workflow on an older server
that ignores the concurrency key. Manual deployments outside this workflow must
also avoid changing staging during journeys.
The gate checks both identities before and after Playwright. It then validates
labels on the captured build output digests and tags them `verified-<full-sha>`.
The manual promotion script resolves all three verified tags to immutable digests
and confirms one matching commit/build ID before moving any `:prod` tag. It polls
production for that same identity using a locally saved release manifest.
A registry error can still interrupt the three tag writes; the Portainer webhook
only runs after successful promotion, and rerunning promotion resolves the full
verified set again. There is no cross-registry atomic tag transaction.
**First rollout:** old green commits without verified tags/build identities are
not promotable through this gate. Build and test a commit containing the new
workflow first. The release route must be reachable through the configured
`STAGING_BASE_URL`/`PROD_BASE_URL`; it deliberately avoids the public staging
`/api` proxy limitation. No new deployment secret is required.
`scripts/ci/release.py` implements identity polling and digest verification.
The poller identifies itself as `SchoolCompare-Release-Check/1.0`: the public
staging proxy has returned HTTP 403 to Python's default urllib user agent even
while the release endpoint was healthy. It logs changes in HTTP/connection
failures or observed release identities, and includes the last observation in
the timeout error. If verification fails, use that observation to distinguish
proxy rejection (403), an unavailable release endpoint (503), and containers
still reporting an older SHA/build ID. Check the configured base URL from the
CI runner; a successful request from another machine does not establish runner
connectivity. Do not bypass identity verification to unblock a deployment.
Its mocked tests run in PR checks alongside backend and index-publication tests.
The new Playwright journeys also check deployed identity and stale pagination.
Local unit checks do not validate registry credentials, Portainer behaviour,
proxy routing or a deployed image; those require the staging run. Production
promotion remains a separate human action.
+94
View File
@@ -0,0 +1,94 @@
# Development and validation
## Prerequisites and environment boundaries
Use a feature branch. The deployed stack is the integration environment; do not
assume a local server can run from a fresh checkout. This cleanup did not start
local servers or provision databases. Unit tests use fixtures and mocks.
The current versions are not yet aligned:
| Component | Container | PR checks |
|---|---|---|
| Backend | Python 3.11 | Python 3.12 |
| Frontend | Node 24 | Node 22 |
| Pipeline | Python 3.13 | Pipeline image build |
Use the component's container version when reproducing deployment behaviour.
The backend dependency pins predate Python 3.14; do not assume the system Python
can install or run them. Version alignment is a separate maintenance task.
## Frontend checks
```sh
cd nextjs-app
npm ci
npm run typecheck
npm test -- --runInBand
```
`npm run build` is the production build check. There is no `lint` script.
Tests live in `__tests__/` and use Jest/React Testing Library. These checks do not
prove that live PostgreSQL queries, Typesense or a deployed proxy work.
The frontend `.env.example` documents runtime variables. Browser traffic normally
uses `/api`; `FASTAPI_URL` is an absolute server-side URL ending in `/api`.
Payload additionally needs `DATABASE_URL` and `PAYLOAD_SECRET` when used at runtime.
Never commit credentials or real `.env` files.
## Backend checks
From the repository root, using an available Python 3.11 or 3.12 interpreter:
```sh
python3.11 -m venv /tmp/schoolcompare-backend-venv
/tmp/schoolcompare-backend-venv/bin/python -m pip install -r requirements.txt pytest 'httpx<0.28' pyyaml
/tmp/schoolcompare-backend-venv/bin/python -m pytest backend/tests pipeline/tests scripts/ci/tests -q
```
Substitute `python3.12` if matching PR CI. The test dependencies above match the
current workflow; they are not yet captured in a dedicated development lockfile.
Backend configuration is defined in `backend/config.py`; `.env.example` documents
commonly used values. `ALLOWED_ORIGINS` uses a JSON array, not a comma-separated string.
## Data and pipeline work
The app needs populated `marts.*` tables. A new Postgres instance alone is not a
working school-data environment. Use the existing managed pipeline or an approved
snapshot; the removed CSV importer cannot build the current schema.
The pipeline container includes Meltano, dbt/Postgres, Airflow and the custom taps.
Airflow commands/selectors live in `pipeline/dags/`. Schema tests live in
`pipeline/transform/tests/` and model YAML files. Run the relevant `dbt build`
selector in an isolated data environment for model changes; it writes tables and
is not a read-only smoke test. Prefer `python -m dbt.cli.main` as the DAGs do.
GIAS dictionaries are generated together by
`pipeline/scripts/generate_gias_codes.py`. The backend and pipeline copies are
intentional; `backend/tests/test_gias_codes.py` checks that they stay identical.
For Payload collection/editor changes, run `npm run generate:importmap` in
`nextjs-app/` and include the generated map. Preserve CMS migrations and the
separate `payload` schema. See [publishing](../nextjs-app/docs/PUBLISHING.md).
## End-to-end checks
Against an existing, authorised test environment:
```sh
cd e2e
npm ci
npx playwright install chromium
BASE_URL=https://your-test-environment.example npx playwright test
```
The suite does not start a web server. CI installs Chromium with system dependencies
and runs against staging. Use the configured staging target: `docs/DEPLOY.md`
records the public staging proxy limitation. User-visible behaviour changes should
update the corresponding journeys.
## Before requesting review
Run checks relevant to the change, inspect `git diff --check`, and report checks
that could not run. Do not publish or promote as part of local validation.
[DEPLOY.md](DEPLOY.md) documents the PR and human promotion gates.
+89
View File
@@ -0,0 +1,89 @@
# Legacy and unused-code inventory
Reviewed 2026-09-14. This inventory records source evidence, not production usage
telemetry. A command with no repository caller may still be run manually or from
an external scheduler. Historical specs and prototypes are not runtime imports.
## Method and scope
Searched backend imports, tests, CLI scripts, Airflow DAGs, Meltano configuration,
Gitea workflows, Dockerfiles and documentation. For frontend candidates, inspected
TypeScript imports, re-exports, literal dynamic imports and `require` calls,
resolving relative and `@/` paths while excluding tests, dependencies and build
output. Checked candidates again with text searches including tests.
Next.js route files, generated Payload import-map entries and plugin discovery
are entry points even without ordinary imports. This is why a zero-import count
alone is not sufficient grounds for deletion. Computed imports and external
operators are outside this static audit.
## Removed in this cleanup
These names are recorded for Git-history lookup; they are no longer file links.
| Removed path or symbol | Evidence and replacement |
|---|---|
| `backend/migration.py` | Imported `School` and `SchoolResult`, which no longer exist in `backend/models.py`. Only the legacy CSV CLI imported it. Current tables are built by dbt. |
| `backend/version.py` | Only the legacy importer consumed `SCHEMA_VERSION`. FastAPI lifespan does not perform version-triggered imports. This is unrelated to active Payload migrations. |
| `scripts/migrate_csv_to_db.py` | Imported removed `init_db`/`set_db_schema_version` helpers and the obsolete models indirectly. No runtime, DAG or workflow calls it. Use the managed pipeline for current marts. |
| `scripts/geocode_schools.py` | Imported the removed `School` ORM model. No pipeline/workflow calls it. Coordinates now come from GIAS/PostGIS; a separate mart-aware manual utility remains under `pipeline/scripts/`. |
| `backend.data_loader.haversine_distance` | No callers. Search uses its inline vectorised NumPy calculation. |
| `nextjs-app/lib/api.ts: fetcher` | No callers; SWR is not installed. Application fetches use the named API wrappers. |
| `nextjs-app/lib/api.ts: kmToMiles` | No callers. `calculateDistance` remains because `CutoffMapPanel` uses it. |
The removed command files could not import successfully against the current
backend. This cleanup does not run replacements, migrate data or modify databases.
Their previous implementations remain recoverable from Git history.
## Unused candidates retained for a separate cleanup
| Candidate | Evidence | Recommended next step |
|---|---|---|
| `nextjs-app/components/LoadingSkeleton.tsx` and its CSS | No application or test imports found. | Remove together after confirming no planned use. |
| `nextjs-app/components/Pagination.tsx` and its CSS | No application or test imports found; HomeView implements load-more behaviour. | Remove as a pair if numbered pagination will not return. |
| `nextjs-app/components/SchoolCard.tsx` and its CSS | Imported by its own tests, not application code. HomeView uses SchoolRow/SecondarySchoolRow. | Decide whether to retire the card design; if removed, remove its dedicated tests as well. Passing tests do not establish runtime use. |
| `backend/database.py: get_db`, `get_db_session` | No remaining callers after removing the importer. Current code creates SessionLocal directly. | Either adopt these helpers during session-lifecycle cleanup or remove them; do not rewrite active sessions in a documentation change. |
| `backend/schemas.py: COLUMN_MAPPINGS`, `NULL_VALUES`, `LA_CODE_TO_NAME` | No remaining Python consumers found after importer removal. Other constants in this module are active. | Remove individual constants after checking external data utilities; retain the module. |
| `backend/config.py: data_dir`, `max_page_size`, `rate_limit_burst` | No active consumers found. `default_page_size` appears only in a branch that expects None, although the route supplies a concrete default. | Reconcile settings with route validation in a focused API change. |
## Legacy/manual paths requiring operational verification
| Path | Status and reason to retain for now |
|---|---|
| FastAPI `/`, `/compare`, `/rankings`, `/favicon.svg`, `/robots.txt`, and conditional `/static` | Old frontend-serving routes reference a `frontend/` directory absent from the checkout and backend image. Next.js owns these public surfaces. Removal changes externally callable routes, so first check proxy/operator usage and define replacement responses. |
| `scripts/fetch_real_data.py`, `scripts/download_data.py` | Historical standalone CSV utilities. The fetch script targets Wandsworth/Merton; neither is wired into the managed pipeline. Marked historical, retained pending confirmation of manual use. |
| `pipeline/scripts/geocode_postcodes.py` | Mart-aware postcode fallback, not called by the current DAGs. Do not confuse it with the removed legacy ORM geocoder. Verify the target schema before manual use. |
| `docker-compose.yml` | Uses unpublished `:latest` release tags and lacks frontend Payload DB/secret/media configuration. Retained as an old development topology, not recommended onboarding. |
| `nextjs-app/docker-compose.yml` | Standalone legacy recipe with old backend port assumptions and no CMS persistence setup. Retained until its consumers are checked. |
| `MIGRATION_SUMMARY.md`, `docs/superpowers/`, `mockups/` | Historical designs and prototypes. Retain as history; do not follow as current deployment instructions. |
| `scripts/sql/drop_fact_parent_view.sql` | One-off maintenance SQL. Not an application entry point; repository call-site searches cannot establish whether it is still needed operationally. |
## Active code that can look obsolete
- `backend/data_loader.py` older-mart query fallbacks are covered by backend tests
and support databases at different migration stages. Remove only after verifying
the deployed schemas in every supported environment.
- `backend/gias_codes.py` and `pipeline/scripts/gias_codes.py` are intentionally
generated copies for separate runtime images. Their parity is tested.
- `nextjs-app/migrations/`, `payload-types.ts` and the Payload import map are active
CMS artifacts, not remnants of the removed school importer.
- `get_available_years`, `get_available_local_authorities` and `get_schools_count`
in `data_loader.py` are called through `get_data_info`, which serves the backend
data-info endpoint. They are not dead functions.
- `get_supplementary_data` is an intentional single-school wrapper around the
batch implementation.
- `pipeline/transform` models named `legacy` can be active data sources: annual
DAG selectors explicitly include legacy KS2/KS4 lineage. Names alone do not
establish obsolescence.
## Suggested next passes
1. Decide the fate of the three unused UI components and remove paired assets/tests.
2. Consolidate backend session usage and remove abandoned settings/constants.
3. Verify external consumers, then retire static-serving API routes and old compose recipes.
4. Audit manual data utilities with pipeline operators before deleting them.
5. Revisit compatibility fallbacks only after documenting supported schema versions.
Validation for this cleanup should include frontend typechecking/tests, Python
syntax checks, reference searches and documentation link checks. Live database,
external scheduler and deployed route usage require separate integration evidence.
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,432 @@
# Other Schools Nearby — Design
**Date:** 2026-09-21, revised 2026-09-22
**Status:** revised after staging review
**Scope:** school detail pages, both phase templates
> **Revision, 2026-09-22.** The first build ranked by intake similarity and used
> distance as a tiebreak. On staging a Catholic primary showed six Catholic
> primaries, none of them close enough to be a real option, and omitted the
> community school down the road. Distance now decides the order and nothing
> else does; the tier system is gone. The reasoning is kept below rather than
> quietly overwritten, because the mistake is the instructive part.
## Goal
Give a school detail page an answer to the question every reader arrives with
after the results tables: *and what else is around here?*
Today a school page links outward to its place pages through
`components/school/NearbyPlaces.tsx` and nowhere else. It never links to another
school. This section adds that edge — up to six nearby schools of the same phase
and a comparable intake, three at a time in a carousel, each a crawlable link and
each addable to the comparison basket in one click.
Mockup, in both themes (drawn against the original tiered design, so its ledes
and chip fallbacks are one revision behind the copy specified below):
<https://claude.ai/artifact/168KdUMcfkUeGWW2FGjuec>
Source of the same page in the repo: `mockups/similar-schools-nearby.html`.
## The constraint that shapes everything
**A nearby school is not automatically a comparable school.**
The section's whole value is that a reader treats what it shows as a shortlist.
That makes every card an implicit claim that the school is a realistic
alternative, and there are three ways that claim goes wrong:
1. A **selective** school beside a non-selective one. Their intakes are
different by construction, so putting their Attainment 8 figures side by side
invites a conclusion the data cannot support.
2. A **special school, PRU or AP** beside a mainstream school. This is the same
error PR #70 fixed for the England benchmark, where Greenmead (URN 101099)
rendered "0% — 62 below England".
3. A **single-sex** school of the opposite sex. Not a weak match — not an option
at all.
So the design separates two kinds of fact, and never confuses them:
- **Hard filters** encode the claims above. They decide eligibility, and are
never relaxed at any distance, even if that means the section does not render.
- **Shared characteristics** — gender, religious character, selectivity —
describe how closely an intake resembles this school's. They are *reported on
the card and never ranked on*, so the reader weighs them rather than having
them weighed for them.
Everything below follows from that split. The revision at the top of this
document is what happens when the second kind is treated as the first.
## Selection algorithm
A backend helper, `_nearby_schools_payload(urn)` in `backend/app.py`, modelled
on the existing `_places_payload(urn)` and delegating to
`backend/nearby_schools.select_nearby(frame, urn)`, which operates on the cached
`load_latest_school_data()` frame — one row per URN, already carrying
`latitude`, `longitude`, `phase`, `gender`, `religious_denomination`,
`admissions_policy`, `school_type` and `status`.
### Hard filters
| Filter | Rule |
|---|---|
| Self | `urn` is excluded |
| Status | GIAS status must be open |
| Coordinates | both `latitude` and `longitude` present on both schools |
| Phase | same phase group via the existing `PHASE_GROUPS` map |
| Provision | special/PRU/AP match only each other |
| Selectivity | selective matches selective; non-selective matches non-selective |
| Gender | Boys never matches Girls; Mixed is compatible with both |
`PHASE_GROUPS` is reused rather than re-derived so an all-through school is
offered correctly on both the primary and secondary sides, exactly as it already
behaves in search.
The provision filter needs a backend counterpart to the frontend's
`isSpecialSchool()` in `nextjs-app/lib/utils.ts:897`, reading the same GIAS
establishment types through `backend/gias_codes.py`. The two must agree: a
school the frontend treats as special for benchmarking but the backend treats as
mainstream for matching would be dropped from its own England comparison and
then offered as a peer to a mainstream school on the next page along.
**Up to six cards, three visible.** Three fit the row; the rest are reached with
the carousel arrows. Two is the minimum that renders at all.
### Order: distance, and nothing else
The nearest eligible schools, closest first. Similarity does not enter the
ranking at any point.
**Why not, having built it the other way first.** The original design ranked by
tiers — same gender and faith within 3 miles, then same gender within 5, then
anything within 10 — and used distance only to order the result. Two things
followed, and both showed up on the first Catholic primary anyone looked at:
- A faith match at 2.9 miles outranked a community school at 0.3 miles. For a
primary, whose catchment is routinely under a mile, the far school is not a
weaker option; it is not an option.
- Because the row filled from the best tier before widening, three Catholic
schools within 3 miles were enough to fill all six slots with Catholic
schools. The stopping rule that produced this had been added to prevent the
*opposite* failure — padding a row with weak distant matches — and made this
one certain.
The premise was backwards. **Distance is a constraint and intake is a
preference.** A parent cannot act on a school outside their reach however well
it matches, and they are perfectly capable of noticing a shared denomination
for themselves if we show it to them. So similarity moved from the ranking to
the card: `shared` reports what a school genuinely has in common, and the reader
applies their own weighting.
The hard filters above were always where the defensibility lived. They are
untouched.
### Reach: a sanity bound, not a target
| Phase | Reach |
|---|---|
| Primary, middle deemed primary, all-through | 2 miles |
| Secondary, middle deemed secondary | 6 miles |
| 16 plus | 10 miles |
Ordering by distance already handles density — a school in inner London fills
all six slots inside a mile and never approaches the cap. The cap decides one
thing: what happens where the area is sparse. It differs by phase because
catchments do, and because people travel furthest for post-16.
**A primary with nothing inside two miles renders no section**, and that is the
intended answer rather than a gap. The alternative is a section headed "nearby"
listing a school four miles from a five-year-old.
**Past the sixth school, the rest are dropped without a count.** The section
does not try to be the list: `NearbyPlaces` sits directly beneath and already
leads to the place pages, which are built for browsing a full set and which the
school page exists to feed.
**Fewer than two results renders nothing.** Not an empty state, not a single
lonely card. The section is absent, the nav item is absent, and the page is
unchanged from before it existed.
### Distance
Straight-line, from the vectorised haversine already used for postcode search at
`backend/app.py:831`, computed over the ~27k-row frame in numpy. Reported to one
decimal place in miles, consistent with the rest of the site.
Straight-line distance is not road distance and is not measured from the
reader's home. The section says so in its disclosure rather than leaving the
reader to assume otherwise.
## API
`/api/schools/{urn}` gains a `similar_schools` array. Each row:
| Field | Notes |
|---|---|
| `urn` | for the link and the compare basket |
| `school_name` | link text |
| `distance_miles` | one decimal place |
| `school_type` | GIAS type, translated, for the card's meta line |
| `age_range` | for the meta line |
| `shared` | what this school genuinely shares with the subject; may be empty |
| `metric_value` | the phase-appropriate headline figure, or null |
| `metric_key` | `rwm_expected_pct` or `attainment_8_score` — see below |
| `metric_year` | the year the figure is from |
The metric follows the subject school's phase side, not the neighbour's own
phase, so a row of cards never mixes two scales. The secondary
side uses `attainment_8_score`; the primary side uses `rwm_expected_pct`. Where
the neighbour has no value for that key, the card reads "Not published" rather
than falling back to the other key.
Which side a school takes is decided once, in
`similar_schools.is_secondary_phase`, by membership of `PHASE_GROUPS["secondary"]`
minus all-through — never by testing for the substring "secondary", which misses
`16 plus` (GIAS phase 6) and hands a sixth-form college the primary bucket.
All-through is the exception in the other direction: `PHASE_GROUPS` lists it on
both sides, but it takes the primary metric.
This is usually the same thing as "the template the page renders", but not
always. `computeSchoolFlags` decides the template with that same substring test,
so a `16 plus` school renders `PrimarySchoolSections` while being matched —
correctly — against secondaries. The section therefore takes its lede noun from
the school's own phase rather than from its template, or it would print "Other
primary schools near <sixth form college>" above a row of secondaries.
Up to six rows of roughly 130 bytes each. It rides in the existing detail payload
rather than a new endpoint because the page already makes exactly one server
fetch for its data, and `/school/[slug]` regenerates at most weekly
(`revalidate = 604800`), so the per-request cost is paid once per school per
week.
**The key is absent, not null, on a backend that does not have this code.** The
frontend treats absent and empty identically, which is what allowed
`NearbyPlaces` to ship without a lockstep deploy of the two images.
`shared` is computed on the backend, beside the data it is derived from, not
re-derived on the frontend. Deriving it twice is how a card comes to claim
something the selection never established. An empty list is a real answer and
renders no chips: a bare card costs a school nothing but the likeness it does
not have, since the order was already settled by distance.
## Frontend
### Components
`components/school/SimilarSchoolsSection.tsx` — a server component wrapped in
the shared `Section` shell from `sectionShared.tsx`. It renders the heading,
the lede, the card grid, the footer CTA and one caption line. Every
card's title is an `<a>` to the school's canonical slug URL via `schoolUrl()`.
`components/school/AddToCompareButton.tsx` — calls `addSchool` from
`ComparisonProvider` and reports the selection with a `from: 'similar_schools'`
attribution, mirroring `addSchoolFromSearch` in `HomeView.tsx:442`.
`components/school/SimilarSchoolsCarousel.tsx` — the scroller and its arrows. It
takes the server-rendered cards as `children` and the server-rendered heading and
lede as a `header` prop, so those stay server components while the client
component owns only the ref, the scroll handler and the arrows' disabled state.
The split matters: the links — the part with SEO value and the part that must
work without JavaScript — are server-rendered into the initial HTML, and only
the basket interaction and the arrows are hydrated.
### The carousel
**Every card is in the initial HTML.** The arrows scroll a list; they never swap
a view. Six `<a>` elements are in the markup whether or not anything is
hydrated, which is the whole reason the section exists — a paginated widget that
mounts cards on click would put four of the six links beyond a crawler and
beyond a reader with no JavaScript.
So the scroller is a plain overflowing `<ul>` with `scroll-snap-type: x
mandatory`, and the arrows call `scrollBy` on it. With no JavaScript it
degrades to a horizontally scrollable row that still works by touch and by
trackpad. Three cards are visible at desktop width and two below 820px.
**Arrows appear only when there is somewhere to go** — that is, only when more
than three schools were found. Each disables itself at its own end of the
travel.
#### Below 640px the arrows go away
This follows [MOBILE.md](../../../MOBILE.md), which makes 360px the design
floor and mobile the primary target at ≥55% of traffic.
Kept in the heading's flex row at 360px, the two arrow buttons take 96px from a
328px card and crush the lede into a four-line column — measured, not guessed.
And swiping already does what they do. So below 640px the header becomes a
single column, the arrows are not rendered, one card shows at 86% width so the
next one peeks, and the affordance is carried by the right-edge scroll-fade that
MOBILE.md documents for exactly this case:
```css
mask-image: linear-gradient(to right, #000 calc(100% - 28px), transparent);
```
The fade lifts at the end of the travel, where there is nothing left to hint
at. That means the at-end state must be computed whether or not an arrow exists
to consume it — on mobile it drives the mask alone.
**Every interactive element clears 44×44px**, per MOBILE.md's iOS HIG check: the
arrow buttons and the add-to-compare button are both 44px, up from the 40px they
were first drawn at. A card title's own box is shorter than that, but its hit
area is the whole card through the `::after` overlay, so it passes on the target
that actually receives the tap.
**The edge test needs a tolerance, and this is not fussiness.** The scroller
carries 2px of padding so focus rings are not clipped, and scroll-snap treats
that padding as the first card's snap position: a scroller sitting at its start
reports `scrollLeft` of 2, not 0. Sub-pixel rounding moves it again at other
zoom levels. Testing `scrollLeft === 0` therefore leaves the back arrow live and
pointing nowhere on first paint — confirmed in the mockup before it was fixed.
Both ends compare against an 8px tolerance.
**Selecting a school must not move the row.** Adding to the basket re-renders
the footer; the scroll offset lives in the DOM rather than in React state, so
the carousel must not remount or reset on that render. A reader who ticks the
fifth school and is thrown back to the first has been punished for using the
feature.
### Placement and navigation
Rendered as the last section **inside** `SchoolDetailShell`, from both
`PrimarySchoolSections` and `SecondarySchoolSections`. Inside, not after, because
the sticky nav's scroll-spy locates sections with `document.getElementById` and
can only reach a section that lives in the shell.
`NearbyPlaces` stays where it is, outside the shell, immediately below. The
resulting order — this school, then similar schools, then the places containing
them — narrows before it widens, which is the order a reader leaves a page in.
`buildNavItems` and `buildSecondaryNavItems` both gain
`{ id: 'similar', label: 'Similar schools' }`, gated on the section rendering.
The id must match the `Section` id or the scroll-spy silently breaks.
### The comparison CTA
A plain `<a href="/compare?urns=…">`, built from this school's URN plus the
selected ones. `/compare` already parses `urns` from the query string
(`app/(frontend)/compare/page.tsx:55`), so this needs no new compare plumbing.
With nothing selected the CTA is disabled; the button also adds to the shared
basket so the site-wide comparison state stays consistent with what the page
shows.
## Copy, and what the section is allowed to claim
**The lede never claims an intake.** It reads "Other primary schools near X." —
one sentence, no variants. The earlier version varied the wording by tier, which
only existed to soften a claim the section should not have been making.
**The heading is "Other schools nearby", not "Similar schools nearby".** The
hard filters do guarantee a comparable set — same phase, same selectivity,
mainstream never beside special — but nothing ranks on likeness, so the heading
does not say it does. The nav item reads "Nearby schools" and the section id is
`nearby`.
**Chips state only what is shared, and may be absent entirely.** A card with
nothing in common renders no chip row rather than falling back to a filler.
Since chips no longer affect the order, an empty one costs that school nothing
except a claim it cannot support — and a Catholic parent scanning the row still
spots "Roman Catholic" on the card that carries it, and weighs it themselves.
**The neighbour's metric carries no valence colour.** Green and terracotta are
reserved site-wide for comparison against the England average. Colouring a
neighbour's figure against this school's would read as ranking the neighbours
against each other, which is precisely the endorsement this section must not
make. The figure sits in neutral ink above a plain "72% at this school"
reference line, and the reader draws their own conclusion.
**A missing figure reads "Not published".** Never 0, never blank, never an
em dash. This follows the same rule the rest of the detail page uses: a school
with no published result has not scored zero.
**There is no "how these are chosen" disclosure.** The method is visible in what
the section already shows — the phase in the lede, the shared characteristics on
each card, the distance above each name — and a collapsed panel restating it
earns less than the space it costs.
**One caption line survives, and only one:** that distances are straight-line
from the school and not road distance. This is not a method note. A reader who
sees "0.6 miles away" and takes it for the walk has been misled by us, and no
other element on the card corrects that. The remaining notes — that listing is
not a recommendation, that special schools only meet special schools — are
statements the selection rules already keep true without being narrated.
## Degradation
| Condition | Behaviour |
|---|---|
| `similar_schools` absent (older backend image) | no section, no nav item |
| fewer than 2 qualifying schools | no section, no nav item |
| this school has no coordinates | no section |
| the helper raises | returns `[]`; the page renders without the section |
The helper is wrapped so a failure inside it never 500s a page that is otherwise
complete — the posture `get_supplementary_data` already takes for its own
queries.
## Testing
**Backend**, in a new `backend/tests/test_similar_schools.py`, against a
synthetic frame rather than live marts:
- a selective school never returns a non-selective one, and vice versa
- a special school returns only special schools; a mainstream school returns none
- a Boys school never returns a Girls school; Mixed matches both
- closed schools and schools without coordinates are never returned
- results are ordered by distance ascending, always
- a faith match never outranks a closer school (the staging defect, pinned)
- the nearest eligible school is always present
- more than six qualifying schools returns the six nearest
- reach is capped per phase, and a primary beyond two miles returns `[]`
- an all-through school is offered on both phase sides
- a `16 plus` school is matched against secondaries and colleges, never primaries
- `is_secondary_phase` and `PHASE_GROUPS` agree on every GIAS phase value
- fewer than two qualifying schools returns `[]`
- distances match a hand-computed haversine for a known pair
**Frontend**, in `nextjs-app/__tests__`:
- the section renders nothing for absent, empty and single-row inputs
- the lede never claims a similar intake
- an empty `shared` renders no chips rather than a filler
- a null metric renders "Not published"
- the nav item appears only alongside the section
- every card is in the DOM, including the ones scrolled out of view
- arrows render only when more than three schools were found
jsdom has no layout, so `scrollWidth` and `clientWidth` are both 0 there and the
arrows' disabled state cannot be meaningfully asserted in Jest. That behaviour is
covered in the journey instead, against a real engine, rather than by a unit test
that would pass on a measurement that does not exist.
**E2E**, added to the existing journeys in `e2e/tests` in the same PR, per the
repository's rule on user-facing behaviour:
- the section renders on a known staging URN, with resolving links
- where arrows are present, the back arrow starts disabled and the forward arrow
moves the row
- selecting a school does not reset the scroll position
- add-to-compare reaches `/compare` with the expected `urns`
- at 360, 390 and 430px: no horizontal overflow, every interactive element in the
section clears 44×44px, and no arrows are rendered
The E2E gate runs after merge on this project, so these journeys are not
provable in the PR checks; the PR is verified on the unit tests, and the
journeys are confirmed on the post-merge staging run.
## Out of scope
- A map of the nearby schools. The section is a list; the page already has a map.
- Autoplay, dots, or an infinite loop on the carousel. It is a short list a
reader scans deliberately, not a banner competing for attention, and a row
that moves on its own is a row that moves while someone is reading it.
- Statistical neighbours on deprivation, size or cohort profile. If plain
distance proves too blunt, that is the trigger to move this computation into a
dbt mart — `select_nearby` is a deliberate seam for exactly that swap.
- Precomputing neighbours in `marts.*`. Rejected for now: a new mart is inert
until Airflow runs, so the feature would ship dark, and every tuning change to
the rules would become a pipeline round-trip instead of a deploy. The revision
at the top of this document is the argument for keeping that loop short.
- Any change to `/api/compare`, the compare page, or the comparison basket.
+189 -1
View File
@@ -1,4 +1,4 @@
import { test, expect, Page } from '@playwright/test';
import { test, expect, Locator, Page } from '@playwright/test';
/**
* Journey tests for SchoolCompare, run against the staging environment as the
@@ -19,6 +19,31 @@ function schoolLinks(page: Page) {
return page.locator('a[href^="/school/"]');
}
/**
* A scroll offset that has stopped moving.
*
* The carousel arrows scroll with `behavior: 'smooth'`, so a reading taken
* straight after a click lands mid-animation. Measured against staging: the
* animation runs ~700ms, and a poll for "has it moved at all" is satisfied
* 50ms in, at 13px of a 1300px journey. A test that then records an offset,
* does something, and records again is measuring the tail of the arrow's
* animation rather than the effect of whatever it did in between.
*
* Two identical readings in a row is the cheapest sound definition of settled.
*/
async function settledScrollLeft(scroller: Locator): Promise<number> {
let previous = -1;
await expect
.poll(async () => {
const current = await scroller.evaluate((node: HTMLElement) => Math.round(node.scrollLeft));
const settled = current === previous;
previous = current;
return settled;
}, { timeout: 10_000 })
.toBe(true);
return previous;
}
/**
* Two URNs guaranteed to be pure-primary (same phase). The compare page's
* phase tabs split all-through schools (which carry KS4 data) onto the
@@ -2734,3 +2759,166 @@ test('the content sitemap lists the about page and is advertised in robots', asy
expect(body).toContain('/sitemap.xml');
expect(body).toContain('/content-sitemap.xml');
});
/**
* Other schools nearby.
*
* The section is absent by design where fewer than two schools qualify, and the
* arrows are absent where three cards fit, so this asserts each part of the
* contract only where it applies.
*
* Two things here cannot be tested anywhere else: the arrows' disabled state,
* which jsdom cannot measure because it has no layout, and the scroll position
* surviving a selection, which is DOM state rather than React state.
*/
test('nearby schools link on to other schools and into compare', async ({ page }) => {
await searchByName(page, 'Primary');
await schoolLinks(page).first().click();
await page.waitForURL(/\/school\//);
const section = page.locator('#nearby');
if ((await section.count()) === 0) {
test.skip(true, 'No qualifying similar schools for this school');
}
// Every card is a real link to another school page — including the ones
// behind the arrows, which is the whole reason this is a scroller and not a
// paginated widget.
const links = section.locator('a[href^="/school/"]');
const linkCount = await links.count();
expect(linkCount).toBeGreaterThanOrEqual(2);
expect(linkCount).toBeLessThanOrEqual(6);
expect(await links.first().getAttribute('href')).toMatch(/^\/school\/\d{6}-/);
await expect(section.getByText(/miles away/).first()).toBeVisible();
const scroller = section.locator('ul').first();
// The carousel, where this school had more than three matches.
const forward = section.getByRole('button', { name: 'More schools' });
if (await forward.count()) {
const back = section.getByRole('button', { name: 'Previous schools' });
await expect(back).toBeDisabled();
await forward.click();
expect(await settledScrollLeft(scroller)).toBeGreaterThan(8);
await expect(back).toBeEnabled();
}
// The compare hand-off, and the row must not jump back to the start when the
// footer re-renders underneath it.
//
// Click the LAST card's button, not the first. Playwright scrolls a target
// into view before clicking it, so clicking card one while the row is paged
// to the end scrolls the container back to the start — and the assertion
// below then measures Playwright's own scrolling rather than the app's.
// That is what this test did on its first staging run: 537 → 2, reproduced
// afterwards on a static page with no React on it at all.
//
// A few pixels of snap or sub-pixel adjustment are fine; a reset to the
// start is not, which is the whole point of the check.
const offsetBefore = await settledScrollLeft(scroller);
await section.getByRole('button', { name: /Add to compare/ }).last().click();
await expect(
section.getByRole('button', { name: /Added to compare/ }).first(),
).toBeVisible();
const offsetAfter = await settledScrollLeft(scroller);
expect(Math.abs(offsetAfter - offsetBefore)).toBeLessThanOrEqual(8);
});
/**
* Every section in the mobile jump sheet can actually be reached.
*
* The sheet is a fixed bottom sheet, and the app has a fixed bottom tab bar.
* `position: sticky` with a z-index on the sticky nav makes it a stacking
* context, so the sheet's own z-index orders it only within that context —
* against the tab bar, the nav's value is what counts. The last item in the
* sheet was therefore painted over and untappable as soon as the list grew
* long enough to reach the bar, which adding "Nearby schools" is what did.
*
* Bounding boxes are not enough to catch this: the item is in the viewport and
* the right size, it is simply underneath something. So this asks the question
* a thumb asks — what is on top at this point.
*/
test('every section in the mobile jump sheet is tappable, not under the tab bar', async ({ page }) => {
await page.setViewportSize({ width: 390, height: 844 });
await searchByName(page, 'Primary');
await schoolLinks(page).first().click();
await page.waitForURL(/\/school\//);
// Scroll down so the sticky nav is docked and the sheet has somewhere to open.
await page.evaluate(() => window.scrollTo({ top: 1200 }));
// Two controls carry aria-haspopup: the mobile "Section" button and the
// desktop "All" one, which is display:none here but still in the DOM.
await page.locator('[aria-haspopup="menu"]:visible').click();
const sheet = page.locator('[role="menu"]');
await expect(sheet).toBeVisible();
const covered = await sheet.evaluate((panel: HTMLElement) =>
Array.from(panel.querySelectorAll('[role="menuitem"]'))
.map((el) => {
const box = el.getBoundingClientRect();
const hit = document.elementFromPoint(
Math.round(box.left + box.width / 2),
Math.round(box.top + box.height / 2),
);
return { label: (el as HTMLElement).innerText.trim().replace(/\s+/g, ' '), reachable: !!(hit && hit.closest('[role="menuitem"]')) };
})
.filter((item) => !item.reachable)
.map((item) => item.label),
);
expect(covered).toEqual([]);
});
/**
* The nearby-schools section at MOBILE.md's three reference widths.
*
* MOBILE.md asks for exactly this check and records that it was not written
* because "Playwright isn't currently in the project dependency set". That is
* no longer true — this suite is Playwright — so the check exists now, scoped
* to the page this feature touches.
*/
for (const width of [360, 390, 430]) {
test(`nearby schools survives a ${width}px viewport`, async ({ page }) => {
await page.setViewportSize({ width, height: 800 });
await searchByName(page, 'Primary');
await schoolLinks(page).first().click();
await page.waitForURL(/\/school\//);
const section = page.locator('#nearby');
if ((await section.count()) === 0) {
test.skip(true, 'No qualifying similar schools for this school');
}
// 1. Nothing bleeds past the right edge.
expect(
await page.evaluate(() => document.documentElement.scrollWidth - window.innerWidth),
).toBe(0);
// 2. No arrows on touch widths — swiping does the job, and they would take
// 96px from a 328px card.
await expect(section.getByRole('button', { name: 'More schools' })).toHaveCount(0);
// 3. Every tap target in the section clears 44px. A card title's own box is
// shorter, but its hit area is the whole card via ::after.
const failing = await section.evaluate((root: HTMLElement) =>
Array.from(root.querySelectorAll('a, button'))
.filter((el) => (el as HTMLElement).offsetParent)
.map((el) => {
const card = el.closest('li');
const box = el.matches('h3 a') && card
? card.getBoundingClientRect()
: el.getBoundingClientRect();
return {
text: (el as HTMLElement).innerText.trim().slice(0, 24),
w: box.width,
h: box.height,
};
})
.filter((o) => o.w < 44 || o.h < 44),
);
expect(failing).toEqual([]);
});
}
+41
View File
@@ -0,0 +1,41 @@
import { test, expect, Route } from '@playwright/test';
test('the deployed frontend and backend report the tested build', async ({ request }) => {
const response = await request.get('/release.json');
expect(response.ok()).toBeTruthy();
expect(response.headers()['cache-control']).toContain('no-store');
const identity = await response.json();
expect(identity.frontend).toEqual(identity.backend);
expect(identity.frontend.sha).toMatch(/^[a-f0-9]{40}$/);
expect(identity.frontend.build_id).toMatch(/^[a-f0-9]{32}$/);
if (process.env.EXPECTED_SHA) expect(identity.frontend.sha).toBe(process.env.EXPECTED_SHA);
if (process.env.EXPECTED_BUILD_ID) expect(identity.frontend.build_id).toBe(process.env.EXPECTED_BUILD_ID);
});
test('changing search while loading another page does not append old results', async ({ page }) => {
await page.goto('/?phase=primary');
await expect(page.getByRole('button', { name: 'Load more schools' })).toBeVisible();
let received!: (route: Route) => void;
const pending = new Promise<Route>(resolve => { received = resolve; });
await page.route('**/api/schools?**', async route => {
if (new URL(route.request().url()).searchParams.get('page') === '2') {
received(route);
return;
}
await route.continue();
});
await page.getByRole('button', { name: 'Load more schools' }).click();
const oldRequest = await pending;
const search = page.getByPlaceholder('School name or postcode').first();
await search.fill('secondary');
await search.press('Enter');
await page.waitForURL(/search=secondary/);
// A cancelled fetch may prevent route fulfilment altogether; either way,
// this deliberately late response must not become part of the new results.
await oldRequest.fulfill({ json: {
schools: [{ urn: 999998, school_name: 'P1 stale result sentinel', phase: 'Primary' }],
total: 2, page: 2, page_size: 1, total_pages: 2,
} }).catch(() => {});
await expect(page.getByText('P1 stale result sentinel')).toHaveCount(0);
await expect(page.getByRole('button', { name: 'Loading...' })).toHaveCount(0);
});
+373
View File
@@ -0,0 +1,373 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Similar Schools Nearby</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600&family=Manrope:wght@500;600;700&display=swap" rel="stylesheet">
<style>
/* Tokens copied verbatim from nextjs-app/app/(frontend)/globals.css so this
mockup cannot drift from the shipped palette. Light values first, dark
under prefers-color-scheme, both overridable by the theme switch. */
:root {
color-scheme: light;
--bg-primary:#FAFAF8; --bg-secondary:#F5EFE6; --bg-card:#FFFFFF;
--text-primary:#1C2731; --text-secondary:#4A5560; --text-muted:#5F6A75;
--border:#E5E7EB; --border-strong:#D3D7DD;
--brand:#0F766E; --brand-strong:#0C5F58; --brand-bg:rgba(15,118,110,.10); --brand-on:#FFFFFF;
--action:#BE3C27; --action-strong:#A33320; --action-on:#FFFFFF;
--sand:#F5EFE6;
--font-display:Manrope,-apple-system,BlinkMacSystemFont,sans-serif;
--font-ui:Inter,-apple-system,BlinkMacSystemFont,sans-serif;
--radius-md:8px; --radius-lg:16px;
--shadow:0 1px 2px rgba(28,39,49,.06),0 1px 3px rgba(28,39,49,.05);
}
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) {
color-scheme: dark;
--bg-primary:#111A20; --bg-secondary:#16222A; --bg-card:#18242C;
--text-primary:#E9EEF0; --text-secondary:#B4C2C7; --text-muted:#8B9AA1;
--border:#26343D; --border-strong:#35454F;
--brand:#5FC7BB; --brand-strong:#7BD6CC; --brand-bg:rgba(95,199,187,.14); --brand-on:#0A1418;
--action:#F08A72; --action-strong:#F5A492; --action-on:#241009;
--sand:#1B2730;
--shadow:0 1px 2px rgba(0,0,0,.3),0 1px 3px rgba(0,0,0,.25);
}
}
:root[data-theme="dark"] {
color-scheme: dark;
--bg-primary:#111A20; --bg-secondary:#16222A; --bg-card:#18242C;
--text-primary:#E9EEF0; --text-secondary:#B4C2C7; --text-muted:#8B9AA1;
--border:#26343D; --border-strong:#35454F;
--brand:#5FC7BB; --brand-strong:#7BD6CC; --brand-bg:rgba(95,199,187,.14); --brand-on:#0A1418;
--action:#F08A72; --action-strong:#F5A492; --action-on:#241009;
--sand:#1B2730;
--shadow:0 1px 2px rgba(0,0,0,.3),0 1px 3px rgba(0,0,0,.25);
}
* { box-sizing:border-box; }
body {
margin:0; padding:32px 16px 80px; background:var(--bg-primary);
color:var(--text-primary); font:15px/1.55 var(--font-ui);
-webkit-font-smoothing:antialiased;
}
.page { max-width:960px; margin:0 auto; }
.page > header { margin-bottom:28px; display:flex; flex-wrap:wrap; gap:16px; align-items:flex-start; justify-content:space-between; }
.page > header h1 { font:700 25px/1.25 var(--font-display); letter-spacing:-.6px; margin:0 0 6px; }
.page > header p { margin:0; color:var(--text-muted); font-size:14px; max-width:60ch; }
.theme-switch { border:1px solid var(--border-strong); background:var(--bg-card); color:var(--text-secondary); border-radius:999px; padding:8px 14px; font:500 13px var(--font-ui); cursor:pointer; min-height:44px; }
.theme-switch:hover { border-color:var(--brand); color:var(--brand); }
/* ── The page context each variant is shown inside ─────────────────── */
.variant { margin-bottom:40px; }
.variant > .context { padding:0 4px 12px; }
.variant .eyebrow { margin:0 0 4px; font-size:12px; letter-spacing:.04em; text-transform:uppercase; color:var(--text-muted); }
.variant .context h2 { font:600 18px/1.35 var(--font-display); margin:0; color:var(--text-secondary); }
.variant .note { margin:10px 4px 0; font-size:12.5px; color:var(--text-muted); }
.variant .note b { color:var(--text-secondary); font-weight:600; }
/* ── The section itself — mirrors components/school/Section ────────── */
.card {
background:var(--bg-card); border:1px solid var(--border);
border-radius:var(--radius-lg); padding:28px; box-shadow:var(--shadow);
}
.top { display:flex; align-items:flex-start; justify-content:space-between; gap:16px; }
.top h2 { font:700 22px/1.25 var(--font-display); letter-spacing:-.4px; margin:0; }
.lede { margin:8px 0 20px; color:var(--text-secondary); font-size:14.5px; max-width:64ch; }
/* ── Carousel ───────────────────────────────────────────────────────
Every card is in the DOM and in the initial HTML — the arrows scroll a
list, they do not swap a view. That keeps all six links crawlable and
keeps the section usable with no JavaScript, where it degrades to a
plain horizontally scrollable row. */
.arrows { display:flex; gap:8px; flex:none; }
.arrow {
width:44px; height:44px; display:grid; place-items:center; cursor:pointer;
border:1px solid var(--border-strong); border-radius:999px;
background:var(--bg-card); color:var(--brand);
}
.arrow:hover:not(:disabled) { border-color:var(--brand); background:var(--brand-bg); }
.arrow:disabled { opacity:.35; cursor:default; }
.arrow:focus-visible { outline:2px solid var(--brand); outline-offset:2px; }
.arrow svg { width:17px; height:17px; }
.scroller {
display:grid; grid-auto-flow:column;
grid-auto-columns:calc((100% - 28px) / 3);
gap:14px; overflow-x:auto; scroll-snap-type:x mandatory;
padding:2px; margin:-2px; /* room for focus rings */
scrollbar-width:none; -ms-overflow-style:none;
list-style:none;
}
.scroller::-webkit-scrollbar { display:none; }
.scroller:focus-visible { outline:2px solid var(--brand); outline-offset:4px; border-radius:var(--radius-md); }
@media (max-width:820px) { .scroller { grid-auto-columns:calc((100% - 14px) / 2); } }
/* Touch widths: the arrows would squeeze the lede into a four-line column for a
control that swiping already provides, so they go and the documented
right-edge fade carries the affordance instead (MOBILE.md). The fade lifts at
the end of the travel, where there is nothing more to hint at. */
@media (max-width:640px) {
.top { display:block; }
.arrows { display:none; }
.scroller { grid-auto-columns:86%; mask-image:linear-gradient(to right, #000 calc(100% - 28px), transparent); }
.scroller[data-at-end=true] { mask-image:none; }
.card { padding:20px; }
}
.school {
position:relative; display:flex; flex-direction:column; scroll-snap-align:start;
border:1px solid var(--border); border-radius:var(--radius-md);
padding:16px; background:var(--bg-card);
}
.school:has(.add[aria-pressed=true]) { border-color:var(--brand); background:var(--brand-bg); }
.distance { display:flex; align-items:center; gap:5px; font-size:12px; color:var(--text-muted); margin:0 0 10px; }
.distance svg { width:13px; height:13px; flex:none; }
.school h3 { font:600 16px/1.35 var(--font-display); margin:0 0 6px; }
/* The whole card is the link target; the button sits above it on z-index so
it stays independently clickable. */
.school h3 a { color:var(--text-primary); text-decoration:none; }
.school h3 a::after { content:""; position:absolute; inset:0; border-radius:var(--radius-md); }
.school:hover { border-color:var(--border-strong); }
.school h3 a:hover { color:var(--brand); text-decoration:underline; }
.school h3 a:focus-visible { outline:none; }
.school:has(h3 a:focus-visible) { outline:2px solid var(--brand); outline-offset:2px; }
.meta { margin:0 0 12px; font-size:12.5px; color:var(--text-muted); }
.shared { display:flex; flex-wrap:wrap; gap:6px; margin:0 0 14px; padding:0; list-style:none; }
.shared li { font-size:11.5px; line-height:1.4; padding:4px 8px; border-radius:999px; background:var(--brand-bg); color:var(--brand); border:1px solid transparent; }
.shared li.loose { background:transparent; color:var(--text-muted); border-color:var(--border); }
.metric { margin-top:auto; padding-top:13px; border-top:1px solid var(--border); }
.value { font:700 26px/1.1 var(--font-display); letter-spacing:-.6px; margin:0; }
.value.absent { font-size:15px; font-weight:600; color:var(--text-muted); letter-spacing:0; }
.metric .label { margin:4px 0 0; font-size:12px; color:var(--text-secondary); }
.metric .ref { margin:2px 0 0; font-size:12px; color:var(--text-muted); }
.add {
position:relative; z-index:1; margin-top:14px; width:100%; min-height:44px;
font:500 13px var(--font-ui); cursor:pointer; border-radius:var(--radius-md);
border:1px solid var(--border-strong); background:var(--bg-card); color:var(--brand);
}
.add:hover { border-color:var(--brand); background:var(--brand-bg); }
.add[aria-pressed=true] { border-color:var(--brand); background:var(--brand-bg); font-weight:600; }
.add:focus-visible { outline:2px solid var(--brand); outline-offset:2px; }
.footer {
display:flex; flex-wrap:wrap; align-items:center; justify-content:space-between;
gap:14px; margin-top:20px; padding-top:18px; border-top:1px solid var(--border);
}
.footer p { margin:0; font-size:12.5px; color:var(--text-muted); }
.footer strong { display:block; font:600 14px var(--font-ui); color:var(--text-primary); }
/* Coral: the one decisive action in this section, and there is only one. */
.compare {
min-height:44px; padding:0 20px; border-radius:var(--radius-md); cursor:pointer;
font:600 14px var(--font-ui); background:var(--action); color:var(--action-on);
border:1px solid var(--action); text-decoration:none; display:inline-flex; align-items:center; gap:8px;
}
.compare:hover { background:var(--action-strong); border-color:var(--action-strong); }
.compare[aria-disabled=true] { opacity:.45; pointer-events:none; }
.caption { margin:16px 0 0; font-size:11.5px; color:var(--text-muted); }
.sr { position:absolute; width:1px; height:1px; padding:0; margin:-1px; overflow:hidden; clip:rect(0 0 0 0); white-space:nowrap; border:0; }
</style>
</head>
<body>
<div class="page">
<header>
<div>
<h1>Similar schools nearby</h1>
<p>A new section on the school detail page. Fictional schools and figures; shipped
colour, type and section shell taken from <code>globals.css</code>.</p>
</div>
<button class="theme-switch" type="button" id="theme">Dark theme</button>
</header>
<div id="variants"></div>
</div>
<script>
const PIN = '<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M20 10c0 6-8 12-8 12s-8-6-8-12a8 8 0 0 1 16 0Z"/><circle cx="12" cy="10" r="3"/></svg>';
const CHEV = (dir) => `<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="${dir === 'prev' ? 'M15 18l-6-6 6-6' : 'M9 18l6-6-6-6'}"/></svg>`;
const variants = [
{
id: 'dense',
eyebrow: 'Variant 1 · Dense urban primary — six matches, carousel active',
context: 'Meadowbrook Primary School — Ages 4–11 · Mixed · No religious character · Community school',
lede: 'Other primary schools near Meadowbrook Primary School, with a similar intake.',
metric: 'Reading, writing & maths',
caption: 'Meeting the expected standard at key stage 2, 2025.',
thisValue: '72%',
note: 'Fourteen schools cleared <b>tier 1</b> within three miles, so the section takes the six nearest and stops there. The arrows scroll a list that is entirely in the HTML — all six links are crawlable, and with JavaScript off the row still scrolls.',
schools: [
{ name:'Willow Lane Primary School', distance:'0.4', meta:'Community school · Ages 4–11', shared:['Mixed','No religious character'], value:'74%' },
{ name:'Oakfield Primary School', distance:'0.6', meta:'Academy converter · Ages 3–11', shared:['Mixed','No religious character'], value:'69%' },
{ name:'Brookside Primary School', distance:'0.9', meta:'Community school · Ages 4–11', shared:['Mixed','No religious character'], value:'Not published' },
{ name:'Hollytree Primary School', distance:'1.3', meta:'Academy converter · Ages 4–11', shared:['Mixed','No religious character'], value:'81%' },
{ name:'Marsh Green Primary School', distance:'1.8', meta:'Community school · Ages 3–11', shared:['Mixed','No religious character'], value:'64%' },
{ name:'Kingsway Primary School', distance:'2.2', meta:'Foundation school · Ages 4–11', shared:['Mixed','No religious character'], value:'77%' },
],
},
{
id: 'secondary',
eyebrow: 'Variant 2 · Secondary — four matches, mixed tiers',
context: 'Meadowbrook High School — Ages 11–18 · Mixed · Non-selective · Academy',
lede: 'Other secondary schools near Meadowbrook High School, with a similar intake.',
metric: 'Attainment 8',
caption: 'Average GCSE attainment score across eight qualifications, 2025.',
thisValue: '51.2',
note: 'Only two schools cleared tier 1, so the search widened to <b>tier 2</b> and found two more. It stops there rather than widening again to reach six — tiers relax to reach a usable set, never to fill the last slots. Selectivity never relaxes, so no grammar school can appear here.',
schools: [
{ name:'Rivermead High School', distance:'0.9', meta:'Academy converter · Ages 11–18', shared:['Mixed','Non-selective','No religious character'], value:'52.8' },
{ name:'Oakfield Academy', distance:'1.7', meta:'Academy sponsor led · Ages 11–16', shared:['Mixed','Non-selective','No religious character'], value:'49.6' },
{ name:'St Aidan’s Catholic High School', distance:'2.4', meta:'Voluntary aided · Ages 11–18', shared:['Mixed','Non-selective'], value:'53.4' },
{ name:'Parkside Community School', distance:'3.8', meta:'Community school · Ages 11–16', shared:['Mixed','Non-selective'], value:'50.9' },
],
},
{
id: 'sparse',
eyebrow: 'Variant 3 · Rural — two matches, no arrows',
context: 'Little Ashby Church of England Primary School — Ages 4–11 · Mixed · Church of England · Voluntary controlled',
lede: 'Other primary schools near Little Ashby Church of England Primary School.',
metric: 'Reading, writing & maths',
caption: 'Meeting the expected standard at key stage 2, 2025.',
thisValue: '66%',
note: 'Nothing matched on religious character within range. At <b>tier 3</b> the lede drops the phrase “with a similar intake” and the chips fall back to the plain phase. Two cards fit the row, so the arrows are not rendered at all. One school fewer and the section would not render either.',
schools: [
{ name:'Great Marden Primary School', distance:'4.2', meta:'Community school · Ages 4–11', shared:['Primary school'], loose:true, value:'71%' },
{ name:'Ashby Vale Academy', distance:'7.8', meta:'Academy converter · Ages 4–11', shared:['Primary school'], loose:true, value:'58%' },
],
},
];
const selected = Object.fromEntries(variants.map((v) => [v.id, new Set()]));
function card(v, s, i) {
const on = selected[v.id].has(i);
const absent = s.value === 'Not published';
return `
<li class="school">
<p class="distance">${PIN}${s.distance} miles away</p>
<h3><a href="#">${s.name}</a></h3>
<p class="meta">${s.meta}</p>
<ul class="shared">${s.shared.map((c) => `<li class="${s.loose ? 'loose' : ''}">${c}</li>`).join('')}</ul>
<div class="metric">
<p class="value ${absent ? 'absent' : ''}">${s.value}</p>
<p class="label">${v.metric}</p>
<p class="ref">${v.thisValue} at this school</p>
</div>
<button class="add" type="button" data-variant="${v.id}" data-index="${i}" aria-pressed="${on}">
${on ? '✓ Added to compare' : '+ Add to compare'}<span class="sr"> — ${s.name}</span>
</button>
</li>`;
}
function render() {
document.getElementById('variants').innerHTML = variants.map((v) => {
const count = selected[v.id].size;
// Three fit the row, so anything more is what the arrows are for.
const scrollable = v.schools.length > 3;
return `
<section class="variant">
<div class="context">
<p class="eyebrow">${v.eyebrow}</p>
<h2>${v.context}</h2>
</div>
<div class="card">
<div class="top">
<div>
<h2 id="h-${v.id}">Similar schools nearby</h2>
<p class="lede">${v.lede}</p>
</div>
${scrollable ? `<div class="arrows">
<button class="arrow" type="button" data-scroll="prev" data-variant="${v.id}" aria-label="Previous schools" aria-controls="sc-${v.id}">${CHEV('prev')}</button>
<button class="arrow" type="button" data-scroll="next" data-variant="${v.id}" aria-label="More schools" aria-controls="sc-${v.id}">${CHEV('next')}</button>
</div>` : ''}
</div>
<ul class="scroller" id="sc-${v.id}" ${scrollable ? `tabindex="0" role="group" aria-labelledby="h-${v.id}"` : ''}>
${v.schools.map((s, i) => card(v, s, i)).join('')}
</ul>
<div class="footer">
<p aria-live="polite">
<strong>${count ? `${count} school${count === 1 ? '' : 's'} selected` : 'Compare side by side'}</strong>
${count ? 'This school is included automatically.' : 'Add a school to compare it with this one.'}
</p>
<a class="compare" href="#" aria-disabled="${count ? 'false' : 'true'}">
${count ? `Compare ${count + 1} schools` : 'Compare'} →
</a>
</div>
<p class="caption">Distances are straight-line from this school, not road distance.
${v.caption} Fictional schools and figures for this mockup.</p>
</div>
<p class="note">${v.note}</p>
</section>`;
}).join('');
variants.forEach((v) => {
const scroller = document.getElementById(`sc-${v.id}`);
if (scroller) syncArrows(v.id, scroller);
});
}
/** An arrow that scrolls nowhere is a dead control, so each end disables its own.
*
* EDGE is not paranoia. The scroller carries 2px of padding so focus rings are
* not clipped, and scroll-snap treats that padding as the first card's snap
* position — so a scroller sitting at its start reports scrollLeft 2, not 0.
* Sub-pixel rounding at other zoom levels moves it again. Testing against an
* exact 0 leaves the back arrow live at the start, pointing nowhere. */
const EDGE = 8;
function syncArrows(id, scroller) {
const max = scroller.scrollWidth - scroller.clientWidth;
const atStart = scroller.scrollLeft <= EDGE;
const atEnd = scroller.scrollLeft >= max - EDGE;
// Drives the mobile scroll-fade, so it is computed even where no arrow is
// rendered to consume it.
scroller.dataset.atEnd = String(atEnd);
const prev = document.querySelector(`.arrow[data-scroll="prev"][data-variant="${id}"]`);
const next = document.querySelector(`.arrow[data-scroll="next"][data-variant="${id}"]`);
if (!prev || !next) return;
prev.disabled = atStart;
next.disabled = atEnd;
}
document.getElementById('variants').addEventListener('click', (event) => {
const arrow = event.target.closest('.arrow');
if (arrow) {
const scroller = document.getElementById(`sc-${arrow.dataset.variant}`);
// A page is what the reader can see, so the viewport is the step.
scroller.scrollBy({ left: (arrow.dataset.scroll === 'next' ? 1 : -1) * scroller.clientWidth, behavior: 'smooth' });
return;
}
const button = event.target.closest('.add');
if (!button) return;
const { variant, index } = button.dataset;
const set = selected[variant];
const i = Number(index);
// Scroll position is DOM state, not React state; keep it across the re-render.
const offset = document.getElementById(`sc-${variant}`).scrollLeft;
set.has(i) ? set.delete(i) : set.add(i);
render();
const scroller = document.getElementById(`sc-${variant}`);
scroller.scrollLeft = offset;
syncArrows(variant, scroller);
document.querySelector(`.add[data-variant="${variant}"][data-index="${index}"]`).focus();
}, true);
document.getElementById('variants').addEventListener('scroll', (event) => {
const scroller = event.target.closest('.scroller');
if (scroller) syncArrows(scroller.id.replace('sc-', ''), scroller);
}, true);
const themeButton = document.getElementById('theme');
themeButton.addEventListener('click', () => {
const dark = document.documentElement.dataset.theme === 'dark';
document.documentElement.dataset.theme = dark ? 'light' : 'dark';
themeButton.textContent = dark ? 'Dark theme' : 'Light theme';
});
render();
</script>
</body>
</html>
+12 -4
View File
@@ -1,8 +1,16 @@
# API Configuration
NEXT_PUBLIC_API_URL=http://localhost:8000/api
# Browser requests use the same-origin Next.js proxy.
NEXT_PUBLIC_API_URL=/api
# Production API URL (for deployment)
# NEXT_PUBLIC_API_URL=https://api.schoolcompare.co.uk/api
# Absolute URL for server-side fetching and the proxy; include /api.
# In the managed container network this is http://backend:80/api (staging differs).
FASTAPI_URL=http://localhost:8000/api
# Payload CMS runtime configuration. Use the managed environment's database;
# Payload owns the payload schema, independently of the school marts.
DATABASE_URL=postgresql://schoolcompare:CHANGE_THIS_PASSWORD@localhost:5432/schoolcompare
# Generate a secret: python -c "import secrets; print(secrets.token_urlsafe(32))"
# Use distinct secrets for staging and production.
PAYLOAD_SECRET=CHANGE_THIS_TO_A_SECURE_RANDOM_SECRET
# Node Environment
NODE_ENV=development
+12 -288
View File
@@ -1,291 +1,15 @@
# Deployment Guide
# Frontend deployment
This guide covers deployment options for the SchoolCompare Next.js application.
Next.js and Payload run in the same frontend container. The maintained deployment
procedure is [docs/DEPLOY.md](../docs/DEPLOY.md), with the production and staging
Portainer compose files at the repository root.
## Deployment Options
The frontend Dockerfile builds a standalone Next.js image. Runtime configuration
supplies `FASTAPI_URL`, `DATABASE_URL` and `PAYLOAD_SECRET`; uploaded CMS media is
persisted in a volume. Promote the built image through the repository's Gitea
workflow after human staging approval.
### Option 1: Vercel (Recommended for Next.js)
Vercel is the easiest and most optimized platform for Next.js applications.
#### Steps:
1. **Install Vercel CLI**:
```bash
npm install -g vercel
```
2. **Login to Vercel**:
```bash
vercel login
```
3. **Deploy**:
```bash
vercel --prod
```
4. **Configure Environment Variables** in Vercel dashboard:
- `NEXT_PUBLIC_API_URL`: Your FastAPI endpoint (e.g., `https://api.schoolcompare.co.uk/api`)
- `FASTAPI_URL`: Same as above for server-side requests
#### Benefits:
- Automatic HTTPS
- Global CDN
- Zero-config deployment
- Automatic preview deployments
- Built-in analytics
---
### Option 2: Docker (Self-hosted)
Deploy using Docker containers for full control.
#### Prerequisites:
- Docker 20+
- Docker Compose 2+
#### Steps:
1. **Build Docker Image**:
```bash
docker build -t schoolcompare-nextjs:latest .
```
2. **Run with Docker Compose**:
```bash
# Create .env file with production variables
echo "NEXT_PUBLIC_API_URL=https://api.schoolcompare.co.uk/api" > .env
echo "FASTAPI_URL=http://backend:8000/api" >> .env
# Start services
docker-compose up -d
```
3. **Verify Deployment**:
```bash
curl http://localhost:3000
```
#### Environment Variables:
- `NEXT_PUBLIC_API_URL`: Public API endpoint (client-side)
- `FASTAPI_URL`: Internal API endpoint (server-side)
- `NODE_ENV`: `production`
---
### Option 3: PM2 (Node.js Process Manager)
Deploy directly on a Node.js server using PM2.
#### Prerequisites:
- Node.js 24+
- PM2 (`npm install -g pm2`)
#### Steps:
1. **Build Application**:
```bash
npm run build
```
2. **Create PM2 Ecosystem File** (`ecosystem.config.js`):
```javascript
module.exports = {
apps: [{
name: 'schoolcompare-nextjs',
script: 'npm',
args: 'start',
cwd: '/path/to/nextjs-app',
instances: 'max',
exec_mode: 'cluster',
env: {
NODE_ENV: 'production',
PORT: 3000,
NEXT_PUBLIC_API_URL: 'https://api.schoolcompare.co.uk/api',
FASTAPI_URL: 'http://localhost:8000/api',
},
}],
};
```
3. **Start with PM2**:
```bash
pm2 start ecosystem.config.js
pm2 save
pm2 startup
```
---
### Option 4: Nginx Reverse Proxy
Use Nginx as a reverse proxy in front of Next.js.
#### Nginx Configuration:
```nginx
server {
listen 80;
server_name schoolcompare.co.uk;
# Redirect to HTTPS
return 301 https://$server_name$request_uri;
}
server {
listen 443 ssl http2;
server_name schoolcompare.co.uk;
# SSL Configuration
ssl_certificate /etc/ssl/certs/schoolcompare.crt;
ssl_certificate_key /etc/ssl/private/schoolcompare.key;
# Security Headers
# frame-ancestors replaces X-Frame-Options so the analytics subdomain
# (Umami heatmap/recorder) can embed the site in an iframe.
add_header Content-Security-Policy "frame-ancestors 'self' https://analytics.schoolcompare.co.uk" always;
add_header X-Content-Type-Options "nosniff" always;
add_header X-XSS-Protection "1; mode=block" always;
# Proxy to Next.js
location / {
proxy_pass http://localhost:3000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection 'upgrade';
proxy_set_header Host $host;
proxy_cache_bypass $http_upgrade;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
# Proxy to FastAPI
location /api/ {
proxy_pass http://localhost:8000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
# Cache static files
location /_next/static/ {
proxy_pass http://localhost:3000;
add_header Cache-Control "public, max-age=31536000, immutable";
}
}
```
---
## Pre-Deployment Checklist
- [ ] Run `npm run build` successfully
- [ ] Run `npm test` - all tests pass
- [ ] Environment variables configured
- [ ] FastAPI backend accessible
- [ ] Database migrations applied
- [ ] SSL certificates configured (production)
- [ ] Domain DNS configured
- [ ] Monitoring/logging set up
- [ ] Backup strategy in place
---
## Post-Deployment Verification
1. **Health Check**:
```bash
curl https://schoolcompare.co.uk
```
2. **Test Routes**:
- Home: `https://schoolcompare.co.uk/`
- School Page: `https://schoolcompare.co.uk/school/100001`
- Compare: `https://schoolcompare.co.uk/compare`
- Rankings: `https://schoolcompare.co.uk/rankings`
3. **Check SEO**:
- Sitemap: `https://schoolcompare.co.uk/sitemap.xml`
- Robots: `https://schoolcompare.co.uk/robots.txt`
4. **Performance Audit**:
- Run Lighthouse in Chrome DevTools
- Target scores: 90+ for Performance, Accessibility, Best Practices, SEO
---
## Monitoring
### Recommended Tools:
- **Vercel Analytics** (if using Vercel)
- **Sentry** for error tracking
- **Google Analytics** for user analytics
- **Uptime Robot** for uptime monitoring
### Health Check Endpoint:
The application automatically serves health data at the root route.
---
## Rollback Procedure
### Vercel:
```bash
vercel rollback
```
### Docker:
```bash
docker-compose down
docker-compose up -d --force-recreate
```
### PM2:
```bash
pm2 stop schoolcompare-nextjs
# Restore previous build
pm2 start schoolcompare-nextjs
```
---
## Troubleshooting
### Issue: API requests failing
- **Solution**: Check `NEXT_PUBLIC_API_URL` and `FASTAPI_URL` environment variables
- **Verify**: FastAPI backend is accessible from Next.js container/server
### Issue: Build fails
- **Solution**: Check Node.js version (requires 24+)
- **Clear cache**: `rm -rf .next node_modules && npm install && npm run build`
### Issue: Slow page loads
- **Solution**: Enable caching in API calls
- **Check**: Network latency to FastAPI backend
- **Verify**: CDN is serving static assets
---
## Security Considerations
- ✅ HTTPS enabled
- ✅ Security headers configured (X-Frame-Options, CSP, etc.)
- ✅ API keys in environment variables (never in code)
- ✅ CORS properly configured
- ✅ Rate limiting on API endpoints
- ✅ Regular security updates
- ✅ Dependency vulnerability scanning
---
## Support
For deployment issues, contact the DevOps team or refer to:
- [Next.js Deployment Docs](https://nextjs.org/docs/deployment)
- [Vercel Documentation](https://vercel.com/docs)
- [Docker Documentation](https://docs.docker.com/)
Earlier Vercel and standalone deployment recipes have been retired from this file
because they do not describe the current CMS, persistence and promotion setup.
See [development](../docs/DEVELOPMENT.md) for checks and
[publishing](docs/PUBLISHING.md) for CMS operations.
+10
View File
@@ -28,6 +28,10 @@ ENV NODE_ENV=production
ARG FASTAPI_URL=http://backend:80/api
ENV FASTAPI_URL=${FASTAPI_URL}
ARG BUILD_SHA=development
ARG BUILD_ID=development
RUN node -e 'require("fs").writeFileSync("build-info.json", JSON.stringify({sha:process.argv[1],build_id:process.argv[2]}))' "$BUILD_SHA" "$BUILD_ID"
# Build application
RUN npm run build
@@ -70,6 +74,12 @@ USER nextjs
EXPOSE 3000
# Set environment variables
ARG BUILD_SHA=development
ARG BUILD_ID=development
LABEL io.schoolcompare.build-id=$BUILD_ID
LABEL io.schoolcompare.commit=$BUILD_SHA
COPY --from=builder /app/build-info.json ./build-info.json
ENV PORT=3000
ENV HOSTNAME="0.0.0.0"
+43 -141
View File
@@ -1,156 +1,58 @@
# SchoolCompare Next.js Application
# SchoolCompare frontend and CMS
Modern Next.js application for comparing primary school KS2 performance across England.
Next.js App Router with React, TypeScript, CSS Modules, Chart.js, Leaflet and
Payload CMS. It serves school search, comparisons, rankings, school/place detail
pages and editorial content across England.
## Features
Start with the [repository overview](../README.md),
[architecture](../docs/ARCHITECTURE.md) and [development checks](../docs/DEVELOPMENT.md).
- **Server-Side Rendering (SSR)**: Fast initial page loads with pre-rendered content
- **Individual School Pages**: Dedicated pages for each school with full SEO optimization
- **Side-by-Side Comparison**: Compare up to 5 schools simultaneously
- **School Rankings**: Top-performing schools by various metrics
- **Interactive Maps**: Leaflet integration for geographic visualization
- **Performance Charts**: Chart.js visualizations for historical data
- **Responsive Design**: Mobile-first approach with full responsive support
- **SEO Optimized**: Dynamic sitemaps, meta tags, and structured data
## Source map
## Tech Stack
| Path | Purpose |
|---|---|
| `app/(frontend)/` | Public root layout, server pages and FastAPI proxy |
| `app/(payload)/` | Payload root layout, `/admin` and `/cms-api` |
| `app/robots.ts`, `app/opengraph-image.tsx`, root icons | Site-wide metadata endpoints |
| `components/` | Client views and reusable display components |
| `components/school/` | School detail sections |
| `lib/api.ts`, `lib/types.ts` | Fetch wrappers and manual school API types |
| `lib/schoolSections.ts`, `lib/compareLogic.ts` | Presentation decisions and data preparation |
| `context/`, `hooks/` | Comparison state, suggestion state and responsive behaviour |
| `collections/`, `blocks/`, `migrations/` | CMS schema and production migrations |
| `__tests__/` | Jest and React Testing Library tests |
- **Framework**: Next.js 16 (App Router)
- **Language**: TypeScript 5
- **Styling**: CSS Modules + CSS Variables
- **State Management**: React Context API + URL state
- **Data Fetching**: SWR (client-side) + Next.js fetch (server-side)
- **Charts**: Chart.js + react-chartjs-2
- **Maps**: Leaflet + react-leaflet
- **Testing**: Jest + React Testing Library
- **Validation**: Zod
Do not introduce a shared `app/layout.tsx`: public pages and Payload have separate
root layouts. Keep root metadata files outside the route groups.
## Getting Started
## Data and state
### Prerequisites
Server pages fetch initial data directly from `FASTAPI_URL`. Browser fetches use
`/api` by default, forwarded by `app/(frontend)/api/[...path]/route.ts`.
`FASTAPI_URL` must include `/api`. See `.env.example` for CMS and API settings.
- Node.js 24+ (using nvm recommended)
- FastAPI backend running on port 8000
State uses React hooks/context, URL search parameters and localStorage for the
comparison basket. SWR is not installed. Maps use dynamic Leaflet wrappers.
Revalidation intervals are configured in fetch wrappers and pages; they vary by
resource. Backend reloads do not automatically invalidate every Next.js cache.
### Installation
## Commands
```bash
# Install dependencies
npm install
# Copy environment variables
cp .env.example .env.local
# Update .env.local with your configuration
```
### Development
```bash
# Start development server
npm run dev
# Open http://localhost:3000
```
### Building
```bash
# Build for production
```sh
npm ci
npm run typecheck
npm test -- --runInBand
npm run build
# Start production server
npm start
```
### Testing
`test:watch` and `test:coverage` are also available. There is no `lint` script.
A running application needs the backend/data environment described in the
[development guide](../docs/DEVELOPMENT.md).
```bash
# Run tests
npm test
After CMS field or editor changes, run `npm run generate:importmap`. Keep
`payload-types.ts` generated from the CMS schema rather than editing it by hand.
The build must work without a database connection; avoid module-scope CMS queries
and DB-backed `generateStaticParams` functions.
# Run tests in watch mode
npm run test:watch
# Run tests with coverage
npm run test:coverage
```
### Linting
```bash
# Run ESLint
npm run lint
```
## Project Structure
```
nextjs-app/
├── app/ # App Router pages
│ ├── layout.tsx # Root layout
│ ├── page.tsx # Home page
│ ├── compare/ # Compare page
│ ├── rankings/ # Rankings page
│ ├── school/[urn]/ # Individual school pages
│ ├── sitemap.ts # Dynamic sitemap
│ └── robots.ts # Robots.txt
├── components/ # React components
│ ├── SchoolCard.tsx # School card component
│ ├── FilterBar.tsx # Search/filter controls
│ ├── ComparisonView.tsx # Comparison interface
│ ├── RankingsView.tsx # Rankings table
│ └── ...
├── lib/ # Utility libraries
│ ├── api.ts # API client
│ ├── types.ts # TypeScript types
│ └── utils.ts # Helper functions
├── hooks/ # Custom React hooks
├── context/ # React Context providers
├── styles/ # Global styles
├── public/ # Static assets
└── __tests__/ # Test files
```
## Environment Variables
| Variable | Description | Default |
|----------|-------------|---------|
| `NEXT_PUBLIC_API_URL` | Public API endpoint (client-side) | `http://localhost:8000/api` |
| `FASTAPI_URL` | Server-side API endpoint | `http://localhost:8000/api` |
| `NODE_ENV` | Environment mode | `development` |
## Performance Optimizations
- **Server-Side Rendering**: Initial HTML rendered on server
- **Static Generation**: Where possible, pages are pre-generated
- **Image Optimization**: Next.js Image component with AVIF/WebP support
- **Code Splitting**: Automatic route-based code splitting
- **Dynamic Imports**: Heavy components loaded on demand
- **API Caching**: Configurable revalidation for data fetching
- **Bundle Optimization**: Tree shaking and minification
- **Compression**: Gzip compression enabled
## SEO Features
- **Dynamic Meta Tags**: Generated per page with Next.js Metadata API
- **Open Graph**: Social media optimization
- **JSON-LD**: Structured data for search engines
- **Sitemap**: Auto-generated from database
- **Robots.txt**: Search engine crawling rules
- **Canonical URLs**: Duplicate content prevention
## Browser Support
- Chrome (latest)
- Firefox (latest)
- Safari (latest)
- Edge (latest)
## License
Proprietary - SchoolCompare
## Support
For issues and questions, please contact the development team.
See [publishing](docs/PUBLISHING.md) for CMS operations and
[deployment](../docs/DEPLOY.md) for staging and production promotion.
@@ -0,0 +1,51 @@
import { fireEvent, render, screen } from '@testing-library/react';
import HomePage from '@/app/(frontend)/page';
import SchoolPage from '@/app/(frontend)/school/[slug]/page';
import ErrorPage from '@/app/(frontend)/error';
import { APIFetchError, fetchSchools, fetchFilters, fetchSchoolDetails } from '@/lib/api';
import { fetchPlace, fetchPlaces } from '@/lib/places';
jest.mock('@/lib/api', () => ({
...jest.requireActual('@/lib/api'),
fetchSchools: jest.fn(),
fetchSchoolDetails: jest.fn(),
fetchFilters: jest.fn(async () => ({})),
fetchDataInfo: jest.fn(async () => null),
fetchNationalAverages: jest.fn(async () => null),
}));
jest.mock('@/lib/flags', () => ({ getFlags: jest.fn(async () => ({})) }));
jest.mock('next/navigation', () => ({
notFound: () => { throw new Error('NEXT_NOT_FOUND'); },
redirect: jest.fn(),
}));
const realFetch = global.fetch;
afterEach(() => { global.fetch = realFetch; jest.clearAllMocks(); });
test('school outages propagate; only a real 404 becomes not found', async () => {
const request = { params: Promise.resolve({ slug: '100001-school' }) };
const outage = new APIFetchError('unavailable', 503);
jest.mocked(fetchSchoolDetails).mockRejectedValueOnce(outage);
await expect(SchoolPage(request)).rejects.toBe(outage);
jest.mocked(fetchSchoolDetails).mockRejectedValueOnce(new APIFetchError('missing', 404));
await expect(SchoolPage(request)).rejects.toThrow('NEXT_NOT_FOUND');
});
test('homepage search failure is not returned as an empty successful page', async () => {
jest.mocked(fetchSchools).mockRejectedValueOnce(new APIFetchError('unavailable', 503));
await expect(HomePage({ searchParams: Promise.resolve({ search: 'school' }) })).rejects.toThrow('unavailable');
});
test('a place is absent only on 404; other failures propagate', async () => {
global.fetch = jest.fn().mockResolvedValue({ ok: false, status: 404 });
await expect(fetchPlace('town', 'example')).resolves.toBeNull();
jest.mocked(global.fetch).mockResolvedValue({ ok: false, status: 503 } as Response);
await expect(fetchPlace('town', 'example')).rejects.toMatchObject({ status: 503 });
await expect(fetchPlaces()).rejects.toMatchObject({ status: 503 });
});
test('the error boundary offers a retry without showing an empty search', () => {
const reset = jest.fn();
render(<ErrorPage error={new Error('offline')} reset={reset} />);
fireEvent.click(screen.getByRole('button', { name: 'Try again' }));
expect(reset).toHaveBeenCalledTimes(1);
});
@@ -0,0 +1,30 @@
/** @jest-environment node */
import { GET } from '@/app/(frontend)/release.json/route';
import { readFile } from 'node:fs/promises';
jest.mock('node:fs/promises', () => ({ readFile: jest.fn() }));
const realFetch = global.fetch;
const identity = { sha: 'a'.repeat(40), build_id: 'b'.repeat(32) };
beforeEach(() => {
jest.mocked(readFile).mockResolvedValue(JSON.stringify(identity));
global.fetch = jest.fn(async () => Response.json(identity));
});
afterEach(() => { global.fetch = realFetch; jest.resetAllMocks(); });
test('reports immutable file identity and backend identity without caching', async () => {
const response = await GET();
expect(response.status).toBe(200);
expect(response.headers.get('Cache-Control')).toBe('no-store');
expect(await response.json()).toEqual({ frontend: identity, backend: identity });
expect(fetch).toHaveBeenCalledWith(expect.stringMatching(/\/api\/release$/), expect.objectContaining({ cache: 'no-store', signal: expect.anything() }));
});
test('missing build metadata cannot pass the release gate', async () => {
jest.mocked(readFile).mockRejectedValueOnce(new Error('missing file'));
expect((await GET()).status).toBe(503);
});
test('backend failure cannot pass the release gate', async () => {
jest.mocked(fetch).mockResolvedValueOnce(new Response('', { status: 503 }));
expect((await GET()).status).toBe(503);
});
@@ -0,0 +1,81 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import { HomeView } from '@/components/HomeView';
import { fetchSchools } from '@/lib/api';
import { primaryFixture } from '../support/schoolFixtures';
import type { SchoolsResponse, School } from '@/lib/types';
let params = new URLSearchParams('postcode=SW1A+1AA');
jest.mock('next/navigation', () => ({
useSearchParams: () => params,
usePathname: () => '/',
useRouter: () => ({ push: jest.fn(), replace: jest.fn() }),
}));
jest.mock('@/context/ComparisonContext', () => ({
useComparisonContext: () => ({ addSchool: jest.fn(), removeSchool: jest.fn(), selectedSchools: [] }),
}));
jest.mock('@/lib/api', () => ({
fetchSchools: jest.fn(),
fetchNationalAverages: jest.fn(async () => ({})),
fetchLAaverages: jest.fn(async () => ({ secondary: { attainment_8_by_la: {} } })),
}));
jest.mock('@/components/FilterBar', () => ({ FilterBar: () => null }));
jest.mock('@/components/SchoolRow', () => ({ SchoolRow: ({ school }: {school: School}) => <div>{school.school_name}</div> }));
jest.mock('@/components/SchoolMap', () => ({ SchoolMap: ({ schools }: {schools: School[]}) => <div data-testid="map">{schools.map(s => s.school_name).join(',')}</div> }));
const filters = { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
function response(name: string): SchoolsResponse {
return { schools: [{ ...primaryFixture.schoolInfo, school_name: name }],
total: 2, page: 1, page_size: 1, total_pages: 2 };
}
function deferred() {
let resolve!: (value: SchoolsResponse) => void;
let reject!: (error: Error) => void;
const promise = new Promise<SchoolsResponse>((yes, no) => { resolve = yes; reject = no; });
return { promise, resolve, reject };
}
beforeEach(() => {
params = new URLSearchParams('postcode=SW1A+1AA');
jest.mocked(fetchSchools).mockReset();
});
test('load-more results from an old search are discarded, even after returning to it', async () => {
const pending = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(pending.promise);
const view = render(<HomeView initialSchools={response('Initial A')} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Load more schools' }));
const signal = jest.mocked(fetchSchools).mock.calls[0][1]?.signal;
params = new URLSearchParams('postcode=SW2+1AA');
view.rerender(<HomeView initialSchools={response('Initial B')} filters={filters} />);
expect(signal?.aborted).toBe(true);
params = new URLSearchParams('postcode=SW1A+1AA');
view.rerender(<HomeView initialSchools={response('Fresh A')} filters={filters} />);
await act(async () => pending.resolve(response('Stale append')));
expect(screen.queryByText('Stale append')).not.toBeInTheDocument();
expect(screen.getByRole('button', { name: 'Load more schools' })).toBeEnabled();
});
test('an older map response cannot overwrite the current search', async () => {
const first = deferred(), second = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(first.promise).mockReturnValueOnce(second.promise);
const initial = response('Initial A');
const view = render(<HomeView initialSchools={initial} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Map' }));
params = new URLSearchParams('postcode=SW2+1AA');
view.rerender(<HomeView initialSchools={response('Initial B')} filters={filters} />);
await act(async () => second.resolve(response('Current map')));
await act(async () => first.resolve(response('Stale map')));
expect(screen.getByTestId('map')).toHaveTextContent('Current map');
expect(screen.getByTestId('map')).not.toHaveTextContent('Stale map');
});
test('failed map requests can be retried by reopening the map', async () => {
const pending = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(pending.promise).mockResolvedValue(response('Retry result'));
render(<HomeView initialSchools={response('Initial')} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Map' }));
await act(async () => pending.reject(new Error('offline')));
fireEvent.click(screen.getByRole('button', { name: 'List' }));
await act(async () => fireEvent.click(screen.getByRole('button', { name: 'Map' })));
expect(fetchSchools).toHaveBeenCalledTimes(2);
expect(screen.getByTestId('map')).toHaveTextContent('Retry result');
});
@@ -0,0 +1,192 @@
/**
* The section's job is to be honest about what it is showing. These tests pin
* the ways it could mislead: rendering below the minimum, claiming a likeness
* it does not rank on, showing a missing figure as a number, or hiding a card
* behind an arrow where a crawler cannot reach it.
*/
import { render, screen } from '@testing-library/react';
import {
nearbyNoun,
NearbySchoolsSection,
shouldRenderNearby,
} from '@/components/school/NearbySchoolsSection';
import type { NearbySchool } from '@/lib/types';
jest.mock('@/components/school/AddToCompareButton', () => ({
AddToCompareButton: ({ school }: { school: NearbySchool }) => (
<button type="button">Add {school.school_name} to compare</button>
),
}));
jest.mock('@/components/school/NearbySchoolsCompareBar', () => ({
NearbySchoolsCompareBar: ({ thisUrn }: { thisUrn: number }) => (
<div data-testid="compare-bar">bar for {thisUrn}</div>
),
}));
function school(overrides: Partial<NearbySchool> = {}): NearbySchool {
return {
urn: 100002,
school_name: 'Willow Lane Primary School',
distance_miles: 0.6,
school_type: 'Community school',
age_range: '4-11',
shared: ['Mixed', 'No religious character'],
metric_value: 74,
metric_key: 'rwm_expected_pct',
metric_year: 202425,
...overrides,
};
}
function renderSection(nearby: NearbySchool[]) {
return render(
<NearbySchoolsSection
urn={100001}
schoolName="Meadowbrook Primary School"
phase="Primary"
thisMetricValue={72}
nearby={nearby}
/>,
);
}
describe('render gates', () => {
it.each([
['undefined', undefined],
['null', null],
['empty', []],
['a single school', [school()]],
])('renders nothing for %s', (_label, value) => {
expect(shouldRenderNearby(value as NearbySchool[] | null | undefined)).toBe(false);
});
it('renders for two or more schools', () => {
expect(shouldRenderNearby([school(), school({ urn: 100003 })])).toBe(true);
});
it('returns null rather than an empty shell below the minimum', () => {
const { container } = renderSection([school()]);
expect(container).toBeEmptyDOMElement();
});
});
describe('what the section claims', () => {
it('never claims a similar intake, because it does not rank on one', () => {
renderSection([school(), school({ urn: 100003, shared: [] })]);
expect(screen.queryByText(/similar intake/i)).not.toBeInTheDocument();
});
it('is headed "Other schools nearby", not "similar"', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByRole('heading', { name: 'Other schools nearby' })).toBeInTheDocument();
});
it('shows chips for what is shared', () => {
renderSection([school({ shared: ['Mixed', 'Roman Catholic'] }), school({ urn: 100003 })]);
expect(screen.getAllByText('Roman Catholic').length).toBe(1);
});
it('shows no chips at all when nothing is shared, rather than inventing one', () => {
const { container } = render(
<NearbySchoolsSection
urn={100001}
schoolName="Meadowbrook Primary School"
phase="Primary"
thisMetricValue={72}
nearby={[school({ shared: [] }), school({ urn: 100003, shared: [] })]}
/>,
);
// The card still carries its distance, name, type and figure — just no
// claim of likeness.
expect(container.querySelectorAll('li ul').length).toBe(0);
expect(screen.getAllByText(/miles away/).length).toBe(2);
});
});
describe('what the lede calls the set', () => {
it.each([
['Primary', 'primary schools'],
['Middle deemed primary', 'primary schools'],
['Secondary', 'secondary schools'],
['Middle deemed secondary', 'secondary schools'],
['All-through', 'all-through schools'],
// GIAS phase 6. Its candidates span the whole secondary group, so no
// single noun fits and it takes the honest general one.
['16 plus', 'schools and colleges'],
['', 'schools'],
[null, 'schools'],
])('calls a %s school\'s neighbours "%s"', (phase, expected) => {
expect(nearbyNoun(phase)).toBe(expected);
});
it('never calls a sixth form college\'s neighbours primary schools', () => {
render(
<NearbySchoolsSection
urn={100001}
schoolName="Barnet Sixth Form College"
phase="16 plus"
thisMetricValue={null}
nearby={[school(), school({ urn: 100003 })]}
/>,
);
expect(screen.getByText(/Other schools and colleges near Barnet Sixth Form College/)).toBeInTheDocument();
expect(screen.queryByText(/primary schools/)).not.toBeInTheDocument();
});
});
describe('cards', () => {
it('links each school to its canonical slug', () => {
renderSection([school(), school({ urn: 100003, school_name: 'Oakfield Primary School' })]);
const link = screen.getByRole('link', { name: /Willow Lane Primary School/ });
expect(link).toHaveAttribute('href', '/school/100002-willow-lane-primary-school');
});
it('shows the distance and the shared characteristics', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getAllByText('0.6 miles away').length).toBeGreaterThan(0);
expect(screen.getAllByText('Mixed').length).toBeGreaterThan(0);
});
it('renders a missing figure as "Not published", never as a number', () => {
renderSection([school({ metric_value: null }), school({ urn: 100003 })]);
expect(screen.getByText('Not published')).toBeInTheDocument();
});
it('anchors each figure against this school', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getAllByText('72% at this school').length).toBe(2);
});
it('offers the compare bar once, for this school', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByTestId('compare-bar')).toHaveTextContent('bar for 100001');
});
it('keeps every card in the DOM, including the ones scrolled out of view', () => {
const six = Array.from({ length: 6 }, (_, n) =>
school({ urn: 100002 + n, school_name: `Peer ${n} School` }),
);
renderSection(six);
expect(screen.getAllByRole('link', { name: /Peer \d School/ })).toHaveLength(6);
});
it('offers no arrows when three cards fit the row', () => {
renderSection([school(), school({ urn: 100003 }), school({ urn: 100004 })]);
expect(screen.queryByRole('button', { name: /More schools/ })).not.toBeInTheDocument();
});
it('offers arrows once there is a fourth school', () => {
renderSection(Array.from({ length: 4 }, (_, n) => school({ urn: 100002 + n })));
expect(screen.getByRole('button', { name: /More schools/ })).toBeInTheDocument();
expect(screen.getByRole('button', { name: /Previous schools/ })).toBeInTheDocument();
});
it('says distances are straight-line, and offers no method panel', () => {
const { container } = renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByText(/straight-line from this school/i)).toBeInTheDocument();
expect(container.querySelector('details')).toBeNull();
});
});
@@ -172,3 +172,39 @@ describe('buildSecondaryNavItems', () => {
expect(ids).not.toContain('history');
});
});
describe('the nearby-schools nav item', () => {
const navInput = {
ofsted: null, admissions: null, admissionDistance: null,
hasLocation: true, yearlyDataLength: 1,
};
it('appears on both templates when the section renders', () => {
const primary = computeSchoolFlags(primaryFixture);
const secondary = computeSecondaryFlags(secondaryFixture);
const input = { ...navInput, hasNearbySchools: true };
expect(buildNavItems(primary, input).map((i) => i.id)).toContain('nearby');
expect(buildSecondaryNavItems(secondary, input).map((i) => i.id)).toContain('nearby');
});
it('is absent when the section does not render', () => {
const primary = computeSchoolFlags(primaryFixture);
const secondary = computeSecondaryFlags(secondaryFixture);
const input = { ...navInput, hasNearbySchools: false };
expect(buildNavItems(primary, input).map((i) => i.id)).not.toContain('nearby');
expect(buildSecondaryNavItems(secondary, input).map((i) => i.id)).not.toContain('nearby');
});
it('is absent when nothing says either way', () => {
const primary = computeSchoolFlags(primaryFixture);
expect(buildNavItems(primary, navInput).map((i) => i.id)).not.toContain('nearby');
});
it('comes last, because the section renders last', () => {
const primary = computeSchoolFlags(primaryFixture);
const ids = buildNavItems(primary, { ...navInput, hasNearbySchools: true }).map((i) => i.id);
expect(ids[ids.length - 1]).toBe('nearby');
});
});
+1 -2
View File
@@ -38,8 +38,7 @@ describe('payload mount points', () => {
});
it('isolates CMS tables in their own postgres schema', () => {
// Blog content must sit outside `public`, where the app tables, Airflow's
// metadata and scripts/migrate_csv_to_db.py --drop all live.
// Blog content must stay separate from school marts and Airflow metadata.
expect(CONFIG).toMatch(/schemaName:\s*['"]payload['"]/);
});
});
+12
View File
@@ -0,0 +1,12 @@
'use client';
export default function ErrorPage({ reset }: { error: Error & { digest?: string }; reset: () => void }) {
return (
<main style={{ maxWidth: '48rem', margin: '4rem auto', padding: '1.5rem' }}>
<h1>We couldn’t load this page</h1>
<p>School information is temporarily unavailable. Please try again.</p>
<button type="button" onClick={reset}>Try again</button>
<p><a href="/">Return to school search</a></p>
</main>
);
}
+44 -67
View File
@@ -85,73 +85,50 @@ export default async function HomePage({ searchParams }: HomePageProps) {
params.has_sixth_form
);
// Fetch data on server with error handling
try {
const [filtersData, dataInfo] = await Promise.all([fetchFilters(), fetchDataInfo().catch(() => null)]);
// Failures propagate to the retryable error boundary.
const [filtersData, dataInfo] = await Promise.all([fetchFilters(), fetchDataInfo().catch(() => null)]);
// Only fetch schools if there are search parameters
let schoolsData;
if (hasSearchParams) {
schoolsData = await fetchSchools({
search: params.search,
local_authority: params.local_authority,
school_type: params.school_type,
phase: params.phase,
postcode: params.postcode,
radius,
page,
page_size: 50,
gender: params.gender,
admissions_policy: params.admissions_policy,
has_sixth_form: params.has_sixth_form,
});
} else {
// Empty state by default
schoolsData = { schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 };
}
const resolvedFilters = filtersData || { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
// `unique_schools`, not `total_schools` — the latter is not a field this
// endpoint returns, and reading it silently yielded null on every request.
const total = dataInfo?.unique_schools ?? null;
const years = dataInfo?.years_available ?? [];
return (
<HomeView
autosuggest={autosuggest}
initialSchools={schoolsData}
filters={resolvedFilters}
totalSchools={total}
howItWorks={hasSearchParams ? null : <HowItWorksSection />}
editorial={hasSearchParams ? null : (
<EditorialSection
totalSchools={total}
localAuthorityCount={resolvedFilters.local_authorities.length}
earliestYearLabel={years.length ? formatAcademicYear(years[0]) : null}
latestYearLabel={years.length ? formatAcademicYear(years[years.length - 1]) : null}
/>
)}
/>
);
} catch (error) {
console.error('Error fetching data for home page:', error);
const emptyFilters = { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
return (
<HomeView
autosuggest={autosuggest}
initialSchools={{ schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 }}
filters={emptyFilters}
totalSchools={null}
howItWorks={hasSearchParams ? null : <HowItWorksSection />}
editorial={hasSearchParams ? null : (
<EditorialSection
totalSchools={null}
localAuthorityCount={0}
earliestYearLabel={null}
latestYearLabel={null}
/>
)}
/>
);
// Only fetch schools if there are search parameters
let schoolsData;
if (hasSearchParams) {
schoolsData = await fetchSchools({
search: params.search,
local_authority: params.local_authority,
school_type: params.school_type,
phase: params.phase,
postcode: params.postcode,
radius,
page,
page_size: 50,
gender: params.gender,
admissions_policy: params.admissions_policy,
has_sixth_form: params.has_sixth_form,
});
} else {
// Empty state by default
schoolsData = { schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 };
}
const resolvedFilters = filtersData || { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
// `unique_schools`, not `total_schools` — the latter is not a field this
// endpoint returns, and reading it silently yielded null on every request.
const total = dataInfo?.unique_schools ?? null;
const years = dataInfo?.years_available ?? [];
return (
<HomeView
autosuggest={autosuggest}
initialSchools={schoolsData}
filters={resolvedFilters}
totalSchools={total}
howItWorks={hasSearchParams ? null : <HowItWorksSection />}
editorial={hasSearchParams ? null : (
<EditorialSection
totalSchools={total}
localAuthorityCount={resolvedFilters.local_authorities.length}
earliestYearLabel={years.length ? formatAcademicYear(years[0]) : null}
latestYearLabel={years.length ? formatAcademicYear(years[years.length - 1]) : null}
/>
)}
/>
);
}
@@ -0,0 +1,17 @@
import { readFile } from 'node:fs/promises';
import path from 'node:path';
export const dynamic = 'force-dynamic';
export const runtime = 'nodejs';
export async function GET() {
try {
const frontend = JSON.parse(await readFile(path.join(process.cwd(), 'build-info.json'), 'utf8'));
const base = process.env.FASTAPI_URL || 'http://localhost:8000/api';
const res = await fetch(`${base}/release`, { cache: 'no-store', signal: AbortSignal.timeout(5000) });
if (!res.ok) throw new Error('Backend identity unavailable');
return Response.json({ frontend, backend: await res.json() }, { headers: { 'Cache-Control': 'no-store', 'X-Robots-Tag': 'noindex' } });
} catch {
return Response.json({ detail: 'Release identity unavailable' }, { status: 503, headers: { 'Cache-Control': 'no-store', 'X-Robots-Tag': 'noindex' } });
}
}
@@ -4,10 +4,11 @@
* URL format: /school/138267-school-name-here
*/
import { fetchSchoolDetails, fetchSchools, fetchNationalAverages } from '@/lib/api';
import { APIFetchError, fetchSchoolDetails, fetchSchools, fetchNationalAverages } from '@/lib/api';
import { notFound, redirect } from 'next/navigation';
import { SchoolDetailShell } from '@/components/school/SchoolDetailShell';
import { NearbyPlaces } from '@/components/school/NearbyPlaces';
import { shouldRenderNearby } from '@/components/school/NearbySchoolsSection';
import { schoolBreadcrumbJsonLd, type SchoolPlace } from '@/lib/jsonld';
import { PrimarySchoolSections } from '@/components/school/PrimarySchoolSections';
import { SecondarySchoolSections } from '@/components/school/SecondarySchoolSections';
@@ -146,8 +147,8 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
fetchNationalAverages().catch(() => null),
]);
} catch (error) {
console.error(`Failed to fetch school ${urn}:`, error);
notFound();
if (error instanceof APIFetchError && error.status === 404) notFound();
throw error;
}
const { school_info, yearly_data, absence_data, ofsted, census, admissions, admissions_history, admission_distance, deprivation, finance, destinations } = data;
@@ -155,6 +156,8 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
// nothing rather than throwing, which is how this shipped without a
// lockstep deploy of the two images.
const places: SchoolPlace[] = data.places ?? [];
// Absent on an older API build, exactly like `places` above.
const nearbySchools = data.nearby_schools ?? [];
// Redirect bare URN to canonical slug URL
const canonicalSlug = schoolUrl(urn, school_info.school_name).replace('/school/', '');
@@ -186,6 +189,7 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
admissions: admissions ?? null,
admissionDistance: admission_distance ?? null,
hasLocation: school_info.latitude != null && school_info.longitude != null,
hasNearbySchools: shouldRenderNearby(nearbySchools),
yearlyDataLength: yearly_data.length,
};
const primaryNavItems = buildNavItems(primaryFlags, navInput);
@@ -262,6 +266,7 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
finance={finance ?? null}
nationalAvg={nationalAvg}
destinations={destinations ?? null}
nearbySchools={nearbySchools}
flags={secondaryFlags}
/>
</SchoolDetailShell>
@@ -284,6 +289,7 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
deprivation={deprivation ?? null}
finance={finance ?? null}
nationalAvg={nationalAvg}
nearbySchools={nearbySchools}
flags={primaryFlags}
/>
</SchoolDetailShell>
+34 -8
View File
@@ -213,6 +213,16 @@ export function HomeView({ initialSchools, filters, totalSchools, howItWorks, ed
const [isLoadingMap, setIsLoadingMap] = useState(false);
const prevSearchParamsRef = useRef(searchParams.toString());
const mapParamsRef = useRef<string>('');
const loadMoreController = useRef<AbortController | null>(null);
// Identity changes even for A → B → A, so an old A response stays stale.
const searchScope = useRef({ key: searchParams.toString() });
if (searchScope.current.key !== searchParams.toString()) {
searchScope.current = { key: searchParams.toString() };
}
useEffect(() => {
setIsLoadingMore(false);
return () => { loadMoreController.current?.abort(); };
}, [searchParams]);
const [geoState, setGeoState] = useState<'idle' | 'requesting' | 'error'>('idle');
const [geoError, setGeoError] = useState<string | null>(null);
/*
@@ -274,17 +284,27 @@ export function HomeView({ initialSchools, filters, totalSchools, howItWorks, ed
if (resultsView !== 'map' || !isLocationSearch) return;
const paramsKey = searchParams.toString();
if (paramsKey === mapParamsRef.current) return;
mapParamsRef.current = paramsKey;
const controller = new AbortController();
const scope = searchScope.current;
const current = () => !controller.signal.aborted && searchScope.current === scope;
setIsLoadingMap(true);
const params: Record<string, any> = {};
searchParams.forEach((value, key) => { params[key] = value; });
params.page = 1;
params.page_size = 500;
fetchSchools(params, { cache: 'no-store' })
.then(r => setMapSchools(r.schools))
.catch(() => setMapSchools(initialSchools.schools))
.finally(() => setIsLoadingMap(false));
}, [resultsView, searchParams]);
fetchSchools(params, { cache: 'no-store', signal: controller.signal })
.then(r => {
if (!current()) return;
mapParamsRef.current = paramsKey;
setMapSchools(r.schools);
})
.catch(() => {
if (current()) setMapSchools(initialSchools.schools);
// No cache marker on failure: opening the map again retries.
})
.finally(() => { if (current()) setIsLoadingMap(false); });
return () => controller.abort();
}, [resultsView, searchParams, initialSchools.schools]);
// Fetch LA averages when secondary or mixed schools are visible
useEffect(() => {
@@ -305,19 +325,25 @@ export function HomeView({ initialSchools, filters, totalSchools, howItWorks, ed
if (isLoadingMore || !hasMore) return;
track('results_load_more', { next_page: currentPage + 1 });
setIsLoadingMore(true);
const scope = searchScope.current;
const controller = new AbortController();
loadMoreController.current?.abort();
loadMoreController.current = controller;
const current = () => !controller.signal.aborted && searchScope.current === scope;
try {
const params: Record<string, any> = {};
searchParams.forEach((value, key) => { params[key] = value; });
params.page = currentPage + 1;
params.page_size = initialSchools.page_size;
const response = await fetchSchools(params, { cache: 'no-store' });
const response = await fetchSchools(params, { cache: 'no-store', signal: controller.signal });
if (!current()) return;
setAllSchools(prev => [...prev, ...response.schools]);
setCurrentPage(response.page);
setHasMore(response.page < response.total_pages);
} catch {
// silently ignore
} finally {
setIsLoadingMore(false);
if (current()) setIsLoadingMore(false);
}
};
@@ -0,0 +1,45 @@
'use client';
/**
* The per-card basket toggle.
*
* The links around it are server-rendered, so the section works with JS off;
* this adds the basket interaction on top rather than being what makes the
* section function.
*/
import { useComparisonContext } from '@/context/ComparisonContext';
import type { School, NearbySchool } from '@/lib/types';
import styles from './NearbySchools.module.css';
export function AddToCompareButton({ school }: { school: NearbySchool }) {
const { addSchool, removeSchool, selectedSchools } = useComparisonContext();
const selected = selectedSchools.some((s) => s.urn === school.urn);
const toggle = () => {
if (selected) {
removeSchool(school.urn);
return;
}
// The basket only needs identity and display fields; the compare page
// fetches everything it renders by URN.
addSchool({
urn: school.urn,
school_name: school.school_name,
school_type: school.school_type,
age_range: school.age_range,
} as School);
};
return (
<button
type="button"
className={styles.add}
onClick={toggle}
aria-pressed={selected}
>
{selected ? '✓ Added to compare' : '+ Add to compare'}
<span className={styles.srOnly}> — {school.school_name}</span>
</button>
);
}
@@ -0,0 +1,64 @@
.heading { font-family: var(--font-display); font-size: 1.4rem; letter-spacing: -0.4px; margin: 0; }
.lede { margin: 0.5rem 0 1.25rem; color: var(--text-secondary); max-width: 64ch; }
.caption { margin: 1rem 0 0; font-size: 0.72rem; color: var(--text-muted); }
.top { display: flex; align-items: flex-start; justify-content: space-between; gap: 1rem; }
.arrows { display: flex; gap: 0.5rem; flex: none; }
.arrow { width: 44px; height: 44px; display: grid; place-items: center; cursor: pointer; border: 1px solid var(--border-strong); border-radius: 999px; background: var(--bg-card); color: var(--brand); }
.arrow:hover:not(:disabled) { border-color: var(--brand); background: var(--brand-bg); }
.arrow:disabled { opacity: 0.35; cursor: default; }
.arrow svg { width: 17px; height: 17px; }
/* A scroller, not a paginated view: every card is in the DOM and the arrows
only move the viewport across them. The 2px padding keeps focus rings from
being clipped — and is why the arrows' edge test needs a tolerance. */
.scroller { display: grid; grid-auto-flow: column; grid-auto-columns: calc((100% - 1.8rem) / 3); gap: 0.9rem; overflow-x: auto; scroll-snap-type: x mandatory; padding: 2px; margin: -2px; list-style: none; scrollbar-width: none; -ms-overflow-style: none; }
.scroller::-webkit-scrollbar { display: none; }
@media (max-width: 820px) { .scroller { grid-auto-columns: calc((100% - 0.9rem) / 2); } }
/* Touch widths (MOBILE.md): the arrows would take 96px from a 328px card and
crush the lede into four lines, for a control swiping already provides. They
go, and the documented right-edge fade carries the affordance — lifting at
the end of the travel, where there is nothing more to hint at. */
@media (max-width: 640px) {
.top { display: block; }
.arrows { display: none; }
.scroller { grid-auto-columns: 86%; mask-image: linear-gradient(to right, #000 calc(100% - 28px), transparent); }
.scroller[data-at-end="true"] { mask-image: none; }
}
.school { position: relative; display: flex; flex-direction: column; scroll-snap-align: start; border: 1px solid var(--border); border-radius: 8px; padding: 1rem; background: var(--bg-card); }
.school:hover { border-color: var(--border-strong); }
.distance { margin: 0 0 0.6rem; font-size: 0.75rem; color: var(--text-muted); }
.name { font-family: var(--font-display); font-size: 1rem; line-height: 1.35; margin: 0 0 0.35rem; }
.name a { color: var(--text-primary); text-decoration: none; }
/* The whole card is the link target; the button sits above it on z-index. */
.name a::after { content: ""; position: absolute; inset: 0; border-radius: 8px; }
.name a:hover { color: var(--brand); text-decoration: underline; }
.meta { margin: 0 0 0.75rem; font-size: 0.78rem; color: var(--text-muted); }
.shared { display: flex; flex-wrap: wrap; gap: 0.35rem; list-style: none; margin: 0 0 0.85rem; padding: 0; }
.chip { font-size: 0.72rem; line-height: 1.4; padding: 0.25rem 0.5rem; border-radius: 999px; background: var(--brand-bg); color: var(--brand); border: 1px solid transparent; }
.metric { margin-top: auto; padding-top: 0.8rem; border-top: 1px solid var(--border); }
/* No valence colour here, deliberately: green and terracotta mean "against the
England average" everywhere else on the site, and colouring a neighbour
against this school would read as ranking the neighbours. */
.value { font-family: var(--font-display); font-size: 1.6rem; font-weight: 700; letter-spacing: -0.6px; margin: 0; color: var(--text-primary); }
.valueAbsent { font-size: 0.95rem; font-weight: 600; margin: 0; color: var(--text-muted); }
.metricLabel { margin: 0.25rem 0 0; font-size: 0.75rem; color: var(--text-secondary); }
.metricRef { margin: 0.1rem 0 0; font-size: 0.75rem; color: var(--text-muted); }
.add { position: relative; z-index: 1; margin-top: 0.85rem; width: 100%; min-height: 44px; font: inherit; font-size: 0.82rem; font-weight: 500; cursor: pointer; border-radius: 8px; border: 1px solid var(--border-strong); background: var(--bg-card); color: var(--brand); }
.add:hover { border-color: var(--brand); background: var(--brand-bg); }
.add[aria-pressed="true"] { border-color: var(--brand); background: var(--brand-bg); font-weight: 600; }
.footer { display: flex; flex-wrap: wrap; align-items: center; justify-content: space-between; gap: 0.85rem; margin-top: 1.25rem; padding-top: 1.1rem; border-top: 1px solid var(--border); }
.footer p { margin: 0; font-size: 0.78rem; color: var(--text-muted); }
.footer strong { display: block; font-size: 0.88rem; font-weight: 600; color: var(--text-primary); }
/* Coral is the one decisive action per screen, and in this section this is it. */
.compare { min-height: 44px; padding: 0 1.25rem; border-radius: 8px; font-size: 0.88rem; font-weight: 600; background: var(--action); color: var(--action-on); border: 1px solid var(--action); text-decoration: none; display: inline-flex; align-items: center; }
.compare:hover { background: var(--action-strong); border-color: var(--action-strong); }
.compare[aria-disabled="true"] { opacity: 0.45; pointer-events: none; }
.srOnly { position: absolute; width: 1px; height: 1px; padding: 0; margin: -1px; overflow: hidden; clip: rect(0 0 0 0); white-space: nowrap; border: 0; }
@@ -0,0 +1,127 @@
'use client';
/**
* The scroller and its arrows.
*
* `children` are the server-rendered cards and `header` the server-rendered
* heading and lede: both stay server components, passed through, so this file
* owns a DOM ref and nothing else. That is what keeps all six links in the
* initial HTML — a carousel that mounted cards on click would put four of the
* six beyond a crawler and beyond a reader with no JavaScript.
*
* With JavaScript off this degrades to a horizontally scrollable row, which is
* still usable by touch and trackpad.
*/
import { useCallback, useEffect, useRef, useState, type ReactNode } from 'react';
import styles from './NearbySchools.module.css';
/** Three cards fit the row, so fewer than four has nowhere to scroll to.
* Below 640px the arrows are not rendered at all — see the stylesheet. */
const VISIBLE = 3;
/**
* Why a tolerance rather than `=== 0`.
*
* The scroller carries 2px of padding so focus rings are not clipped, and
* scroll-snap treats that padding as the first card's snap position — a row at
* rest reports scrollLeft 2, not 0. Sub-pixel rounding moves it again at other
* zoom levels. An exact test leaves the back arrow live on first paint,
* pointing nowhere.
*/
const EDGE = 8;
export function NearbySchoolsCarousel({
count,
labelledBy,
header,
children,
}: {
count: number;
labelledBy: string;
header: ReactNode;
children: ReactNode;
}) {
const scroller = useRef<HTMLUListElement>(null);
const [atStart, setAtStart] = useState(true);
const [atEnd, setAtEnd] = useState(false);
const scrollable = count > VISIBLE;
const sync = useCallback(() => {
const node = scroller.current;
if (!node) return;
const max = node.scrollWidth - node.clientWidth;
setAtStart(node.scrollLeft <= EDGE);
setAtEnd(node.scrollLeft >= max - EDGE);
}, []);
// `atEnd` is not only the forward arrow's disabled state: below 640px, where
// no arrow is rendered, it is the only thing driving the scroll-fade.
// Also on mount: the first measurement can only happen once there is layout.
useEffect(sync, [sync]);
const page = (direction: 1 | -1) => {
const node = scroller.current;
if (!node) return;
// A page is what the reader can see, so the viewport is the step.
node.scrollBy({ left: direction * node.clientWidth, behavior: 'smooth' });
};
return (
<>
<div className={styles.top}>
{header}
{scrollable && (
<div className={styles.arrows}>
<button
type="button"
className={styles.arrow}
onClick={() => page(-1)}
disabled={atStart}
aria-label="Previous schools"
>
<Chevron direction="prev" />
</button>
<button
type="button"
className={styles.arrow}
onClick={() => page(1)}
disabled={atEnd}
aria-label="More schools"
>
<Chevron direction="next" />
</button>
</div>
)}
</div>
<ul
ref={scroller}
className={styles.scroller}
onScroll={sync}
data-at-end={atEnd}
{...(scrollable
? { tabIndex: 0, role: 'group', 'aria-labelledby': labelledBy }
: {})}
>
{children}
</ul>
</>
);
}
function Chevron({ direction }: { direction: 'prev' | 'next' }) {
return (
<svg
viewBox="0 0 24 24"
fill="none"
stroke="currentColor"
strokeWidth="2"
strokeLinecap="round"
strokeLinejoin="round"
aria-hidden="true"
>
<path d={direction === 'prev' ? 'M15 18l-6-6 6-6' : 'M9 18l6-6-6-6'} />
</svg>
);
}
@@ -0,0 +1,59 @@
'use client';
/**
* The selection count and the one decisive action in the section.
*
* The CTA is a real link, not a handler: /compare already parses `urns` from
* the query string, so the hand-off needs no new compare plumbing. It counts
* this school plus whatever the reader ticked, because comparing a shortlist
* without the school they are looking at is not what they asked for.
*/
import Link from 'next/link';
import { useComparisonContext } from '@/context/ComparisonContext';
import type { NearbySchool } from '@/lib/types';
import styles from './NearbySchools.module.css';
export function NearbySchoolsCompareBar({
thisUrn,
candidates,
}: {
thisUrn: number;
candidates: NearbySchool[];
}) {
const { selectedSchools } = useComparisonContext();
// Only the schools this section offers, in the order the cards show them —
// the basket may hold schools picked up elsewhere on the site, and this bar
// speaks for this section.
const offered = candidates
.map((c) => c.urn)
.filter((urn) => selectedSchools.some((s) => s.urn === urn));
const count = offered.length;
const href = `/compare?urns=${[thisUrn, ...offered].join(',')}`;
return (
<div className={styles.footer}>
<p aria-live="polite">
<strong>
{count
? `${count} school${count === 1 ? '' : 's'} selected`
: 'Compare side by side'}
</strong>
{count
? 'This school is included automatically.'
: 'Add a school to compare it with this one.'}
</p>
{count ? (
<Link className={styles.compare} href={href}>
Compare {count + 1} schools →
</Link>
) : (
<span className={styles.compare} aria-disabled="true">
Compare →
</span>
)}
</div>
);
}
@@ -0,0 +1,152 @@
/**
* NearbySchoolsSection — the nearest eligible schools, closest first. Server
* component; only the carousel, the compare bar and the add-to-compare button
* are client-side.
*
* "Other schools nearby", not "similar" ones: the order is distance and only
* distance. The hard filters upstream still guarantee the set is comparable —
* same phase, same selectivity, mainstream never beside special — but nothing
* here ranks by how alike two schools are, so the heading does not say it does.
*
* The chips report what a school shares, and may be absent entirely. That is
* information for the reader to weigh, not a verdict this section has already
* reached on their behalf.
*
* There is deliberately no "how these are chosen" panel: the method is already
* visible in the lede, the chips and the distances. The single caption line is
* not a method note — it is the one thing a card cannot self-correct.
*/
import Link from 'next/link';
import type { NearbySchool } from '@/lib/types';
import { schoolUrl } from '@/lib/utils';
import { AddToCompareButton } from './AddToCompareButton';
import { NearbySchoolsCarousel } from './NearbySchoolsCarousel';
import { NearbySchoolsCompareBar } from './NearbySchoolsCompareBar';
import { Section } from './sectionShared';
import styles from './NearbySchools.module.css';
const MINIMUM = 2;
export function shouldRenderNearby(nearby?: NearbySchool[] | null): boolean {
return (nearby?.length ?? 0) >= MINIMUM;
}
/**
* What the lede calls the set of schools it is showing.
*
* Derived from the school's own GIAS phase rather than the template it renders
* with, because those disagree for "16 plus" (GIAS phase 6): a sixth-form
* college renders the primary template — computeSchoolFlags tests for the
* substring "secondary" — while the backend correctly matches it against the
* secondary group. Taking the noun from the template would print "Other primary
* schools near <sixth form college>" above a row of secondaries.
*
* A 16-plus school's candidates span the whole secondary group, so no single
* noun fits and it gets the honest general one.
*/
export function nearbyNoun(phase: string | null | undefined): string {
const text = (phase ?? '').trim().toLowerCase();
if (text === 'all-through') return 'all-through schools';
if (text === '16 plus') return 'schools and colleges';
if (text.includes('secondary')) return 'secondary schools';
if (text.includes('primary')) return 'primary schools';
return 'schools';
}
function metricLabel(key: string): string {
return key === 'attainment_8_score' ? 'Attainment 8' : 'Reading, writing & maths';
}
function formatMetric(value: number | null, key: string): string {
if (value == null) return 'Not published';
return key === 'attainment_8_score' ? value.toFixed(1) : `${Math.round(value)}%`;
}
export function NearbySchoolsSection({
urn,
schoolName,
phase,
thisMetricValue,
nearby,
}: {
urn: number;
schoolName: string;
/** The school's own GIAS phase, not the template it renders with. */
phase: string | null | undefined;
thisMetricValue: number | null;
nearby?: NearbySchool[] | null;
}) {
if (!shouldRenderNearby(nearby)) return null;
const schools = nearby as NearbySchool[];
// One card matched on phase alone, so the section may not claim the set
// shares an intake with this school.
const metricKey = schools[0].metric_key;
const noun = nearbyNoun(phase);
return (
<Section id="nearby">
<NearbySchoolsCarousel
count={schools.length}
labelledBy="nearby-schools-heading"
header={
<div>
<h2 id="nearby-schools-heading" className={styles.heading}>
Other schools nearby
</h2>
<p className={styles.lede}>{`Other ${noun} near ${schoolName}.`}</p>
</div>
}
>
{schools.map((school) => (
<li key={school.urn} className={styles.school}>
<p className={styles.distance}>{school.distance_miles} miles away</p>
<h3 className={styles.name}>
<Link href={schoolUrl(school.urn, school.school_name)}>
{school.school_name}
</Link>
</h3>
<p className={styles.meta}>
{[school.school_type, school.age_range && `Ages ${school.age_range}`]
.filter(Boolean)
.join(' · ')}
</p>
{school.shared.length > 0 && (
<ul className={styles.shared}>
{school.shared.map((label) => (
<li key={label} className={styles.chip}>{label}</li>
))}
</ul>
)}
<div className={styles.metric}>
<p
className={
school.metric_value == null ? styles.valueAbsent : styles.value
}
>
{formatMetric(school.metric_value, school.metric_key)}
</p>
<p className={styles.metricLabel}>{metricLabel(school.metric_key)}</p>
{thisMetricValue != null && (
<p className={styles.metricRef}>
{formatMetric(thisMetricValue, metricKey)} at this school
</p>
)}
</div>
<AddToCompareButton school={school} />
</li>
))}
</NearbySchoolsCarousel>
<NearbySchoolsCompareBar thisUrn={urn} candidates={schools} />
{/* The one caveat the cards cannot make on their own: a reader who takes
"0.6 miles away" for the walk has been misled, and nothing else here
corrects that. */}
<p className={styles.caption}>
Distances are straight-line from this school, not road distance.
</p>
</Section>
);
}
@@ -14,6 +14,7 @@
import type {
School, SchoolResult, AbsenceData, OfstedInspection, SchoolCensus,
SchoolAdmissions, SchoolAdmissionDistance, SchoolDeprivation, SchoolFinance, NationalAverages,
NearbySchool,
} from '@/lib/types';
import { ofstedLegacyAreas } from '@/lib/utils';
import type { SchoolFlags } from '@/lib/schoolSections';
@@ -26,6 +27,7 @@ import { HistorySection } from './HistorySection';
import { SchoolLifeSection } from './SchoolLifeSection';
import { LocalAreaSection } from './LocalAreaSection';
import { FinancesSection } from './FinancesSection';
import { NearbySchoolsSection } from './NearbySchoolsSection';
export interface PrimarySchoolSectionsProps {
schoolInfo: School;
@@ -39,13 +41,15 @@ export interface PrimarySchoolSectionsProps {
deprivation: SchoolDeprivation | null;
finance: SchoolFinance | null;
nationalAvg: NationalAverages | null;
/** Nearby schools of a comparable intake. Absent on an older API build. */
nearbySchools?: NearbySchool[];
flags: SchoolFlags;
}
export function PrimarySchoolSections({
schoolInfo, yearlyData, absenceData, ofsted, census,
admissions, admissionsHistory, admissionDistance,
deprivation, finance, nationalAvg, flags,
deprivation, finance, nationalAvg, nearbySchools, flags,
}: PrimarySchoolSectionsProps) {
const primaryAvg = nationalAvg?.primary ?? {};
const secondaryAvg = nationalAvg?.secondary ?? {};
@@ -146,6 +150,15 @@ export function PrimarySchoolSections({
)}
{flags.hasFinance && finance && <FinancesSection finance={finance} />}
{/* Last: it is where the reader goes next, not part of this school. */}
<NearbySchoolsSection
urn={schoolInfo.urn}
schoolName={schoolInfo.school_name}
phase={schoolInfo.phase}
thisMetricValue={flags.latestResults?.rwm_expected_pct ?? null}
nearby={nearbySchools}
/>
</>
);
}
@@ -283,6 +283,21 @@
}
/* `position: sticky` with a z-index makes .sectionNav a stacking context, so
the jump sheet's own z-index only orders it INSIDE that context. Against the
fixed bottom tab bar (Navigation.module.css, z-index 1000) what counts is
.sectionNav's 10 — which is why the sheet's last item was painted over, and
untappable, once the list grew long enough to reach the bar.
Lifted only while the sheet is open, and only to 1100: above the bar, below
the comparison toast (2000), the fullscreen map (5000) and the info popover
(9999). This rule carries the z-index and nothing else; every other
declaration belongs to .sectionNav in both states. */
.sectionNavSheetOpen {
z-index: 1100;
}
.sectionNavBack {
flex: none;
display: inline-flex;
@@ -344,7 +344,10 @@ export function SchoolDetailShell({
</header>
{/* Sticky Section Navigation — docks under the global header */}
<nav className={styles.sectionNav} aria-label="Page sections">
<nav
className={`${styles.sectionNav}${sectionsOpen ? ` ${styles.sectionNavSheetOpen}` : ''}`}
aria-label="Page sections"
>
<button onClick={scrollToTop} className={styles.sectionNavBack} aria-label="Back to top">
<span aria-hidden="true">↑</span>
<span className={styles.sectionNavBackLabel}>Top</span>
@@ -14,7 +14,7 @@
import type {
School, SchoolResult, AbsenceData, OfstedInspection, SchoolCensus,
SchoolAdmissions, SchoolAdmissionDistance, SchoolDeprivation, SchoolFinance, NationalAverages,
SchoolDestinations,
SchoolDestinations, NearbySchool,
} from '@/lib/types';
import { ofstedLegacyAreas } from '@/lib/utils';
import type { SecondaryFlags } from '@/lib/schoolSections';
@@ -27,6 +27,7 @@ import { DistanceSection } from './DistanceSection';
import { SecondaryHistorySection } from './SecondaryHistorySection';
import { WellbeingSection } from './WellbeingSection';
import { FinancesSection } from './FinancesSection';
import { NearbySchoolsSection } from './NearbySchoolsSection';
import styles from './schoolSections.module.css';
export interface SecondarySchoolSectionsProps {
@@ -47,13 +48,15 @@ export interface SecondarySchoolSectionsProps {
finance: SchoolFinance | null;
nationalAvg: NationalAverages | null;
destinations: SchoolDestinations | null;
/** Nearby schools of a comparable intake. Absent on an older API build. */
nearbySchools?: NearbySchool[];
flags: SecondaryFlags;
}
export function SecondarySchoolSections({
schoolInfo, yearlyData, ofsted, census,
admissions, admissionsHistory, admissionDistance,
deprivation, finance, nationalAvg, destinations, flags,
deprivation, finance, nationalAvg, destinations, nearbySchools, flags,
}: SecondarySchoolSectionsProps) {
const secondaryAvg = nationalAvg?.secondary ?? {};
@@ -141,6 +144,15 @@ export function SecondarySchoolSections({
{flags.hasFinance && finance && (
<FinancesSection finance={finance} showPremises />
)}
{/* Last: it is where the reader goes next, not part of this school. */}
<NearbySchoolsSection
urn={schoolInfo.urn}
schoolName={schoolInfo.school_name}
phase={schoolInfo.phase}
thisMetricValue={flags.latestResults?.attainment_8_score ?? null}
nearby={nearbySchools}
/>
</div>
);
}
+1
View File
@@ -9,6 +9,7 @@ const createJestConfig = nextJest({
const customJestConfig = {
setupFilesAfterEnv: ['<rootDir>/jest.setup.js'],
testEnvironment: 'jest-environment-jsdom',
modulePathIgnorePatterns: ['<rootDir>/.next/'],
moduleNameMapper: {
'^@/(.*)$': '<rootDir>/$1',
},
-27
View File
@@ -311,26 +311,6 @@ export async function fetchDataInfo(
return handleResponse<DataInfoResponse>(response);
}
// ============================================================================
// Client-Side Fetcher (for SWR)
// ============================================================================
/**
* Generic fetcher function for use with SWR
* @example
* ```tsx
* const { data, error } = useSWR('/api/schools', fetcher);
* ```
*/
export async function fetcher<T>(url: string): Promise<T> {
// If it's already a full URL, use it directly
// Otherwise, prepend the API_BASE_URL
const fullUrl = url.startsWith('http') ? url : `${API_BASE_URL}${url.startsWith('/') ? url : `/${url}`}`;
const response = await fetch(fullUrl);
return handleResponse<T>(response);
}
// ============================================================================
// Geocoding API
// ============================================================================
@@ -396,10 +376,3 @@ export function calculateDistance(
const c = 2 * Math.atan2(Math.sqrt(a), Math.sqrt(1 - a));
return R * c;
}
/**
* Convert kilometers to miles
*/
export function kmToMiles(km: number): number {
return km * 0.621371;
}
+5 -2
View File
@@ -1,3 +1,5 @@
import { APIFetchError } from './api';
/**
* Client for the places API.
*
@@ -69,7 +71,7 @@ const API = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL
export async function fetchPlaces(): Promise<PlaceSummary[]> {
const res = await fetch(`${API}/places`, { next: { revalidate: 604800 } });
if (!res.ok) return [];
if (!res.ok) throw new APIFetchError("Unable to load places", res.status);
return (await res.json()).places ?? [];
}
@@ -79,6 +81,7 @@ export async function fetchPlace(
const q = phase ? `?phase=${encodeURIComponent(phase)}` : '';
const res = await fetch(`${API}/places/${kind}/${slug}${q}`,
{ next: { revalidate: 604800 } });
if (!res.ok) return null;
if (res.status === 404) return null;
if (!res.ok) throw new APIFetchError("Unable to load place", res.status);
return res.json();
}
+16 -2
View File
@@ -128,6 +128,10 @@ export interface NavItemsInput {
* measure a postcode, so the nav must gate on them too or it will link to an
* anchor that was never rendered. */
hasLocation?: boolean;
/** Whether the nearby-schools section will render. Optional for the same
* reason hasLocation is: the nav must never link to an anchor that was not
* rendered, and absent has to mean "no section". */
hasNearbySchools?: boolean;
yearlyDataLength: number;
}
@@ -142,7 +146,10 @@ export interface NavItemsInput {
*/
export function buildNavItems(
flags: SchoolFlags,
{ ofsted, admissions, admissionDistance, hasLocation, yearlyDataLength }: NavItemsInput,
{
ofsted, admissions, admissionDistance, hasLocation,
hasNearbySchools, yearlyDataLength,
}: NavItemsInput,
): NavItem[] {
const navItems: NavItem[] = [];
if (ofsted) navItems.push({ id: 'ofsted', label: 'Ofsted' });
@@ -161,6 +168,8 @@ export function buildNavItems(
if (flags.hasSchoolLife) navItems.push({ id: 'school-life', label: 'School Life' });
if (flags.hasDeprivation) navItems.push({ id: 'local-area', label: 'Local Area' });
if (flags.hasFinance) navItems.push({ id: 'finances', label: 'Finances' });
// Last, because the section renders last — the scroll-spy reads this order.
if (hasNearbySchools) navItems.push({ id: 'nearby', label: 'Nearby schools' });
return navItems;
}
@@ -239,7 +248,10 @@ export function computeSecondaryFlags({
*/
export function buildSecondaryNavItems(
flags: SecondaryFlags,
{ ofsted, admissions, admissionDistance, hasLocation, yearlyDataLength }: NavItemsInput,
{
ofsted, admissions, admissionDistance, hasLocation,
hasNearbySchools, yearlyDataLength,
}: NavItemsInput,
): NavItem[] {
const navItems: NavItem[] = [];
if (ofsted) navItems.push({ id: 'ofsted', label: 'Ofsted' });
@@ -257,5 +269,7 @@ export function buildSecondaryNavItems(
if (yearlyDataLength > 1) navItems.push({ id: 'history', label: 'History' });
if (flags.hasWellbeing) navItems.push({ id: 'wellbeing', label: 'Wellbeing' });
if (flags.hasFinance) navItems.push({ id: 'finances', label: 'Finances' });
// Last, because the section renders last — the scroll-spy reads this order.
if (hasNearbySchools) navItems.push({ id: 'nearby', label: 'Nearby schools' });
return navItems;
}
+28
View File
@@ -346,6 +346,26 @@ export interface SchoolsResponse {
};
}
/**
* A nearby school, from the nearest-first set the detail page shows.
*
* `shared` is what this school genuinely has in common with the one being
* viewed, and may be empty. It is reported, never ranked on: an earlier
* version ordered by it and buried the school down the road under faith
* matches three times further away.
*/
export interface NearbySchool {
urn: number;
school_name: string;
distance_miles: number;
school_type: string | null;
age_range: string | null;
shared: string[];
metric_value: number | null;
metric_key: string;
metric_year: number | null;
}
export interface SchoolDetailsResponse {
school_info: School;
/**
@@ -357,6 +377,14 @@ export interface SchoolDetailsResponse {
* authority both fall below the publish threshold has nowhere to link.
*/
places?: SchoolPlace[];
/**
* Up to six nearby eligible schools, nearest first.
*
* Optional for the same reason as `places`: a frontend deployed ahead of the
* API that serves this must render without it. Absent and empty mean the
* same thing here — no section.
*/
nearby_schools?: NearbySchool[];
yearly_data: SchoolResult[];
absence_data: AbsenceData | null;
// Supplementary data (null until Kestra populates)
+2 -3
View File
@@ -24,9 +24,8 @@ export default buildConfig({
typescript: { outputFile: path.resolve(dirname, 'payload-types.ts') },
db: postgresAdapter({
pool: { connectionString: process.env.DATABASE_URL },
// Its own schema, so no pipeline operation on `public` can reach blog
// content. scripts/migrate_csv_to_db.py --drop lives in that blast radius,
// as does Airflow's metadata. The schema itself is created by the initial
// Its own schema separates blog content from pipeline-managed school
// tables and Airflow metadata. The schema is created by the initial
// migration: schemaName says where tables go, it does not create anything.
//
// prodMigrations runs pending migrations during server init. Without it a
+5
View File
@@ -40,4 +40,9 @@ ENV AIRFLOW_HOME=/opt/airflow
ENV AIRFLOW__CORE__DAGS_FOLDER=/opt/pipeline/dags
ENV PYTHONPATH=/opt/pipeline
ARG BUILD_SHA=development
ARG BUILD_ID=development
LABEL io.schoolcompare.build-id=$BUILD_ID
LABEL io.schoolcompare.commit=$BUILD_SHA
CMD ["airflow", "api-server"]
+69 -50
View File
@@ -11,6 +11,10 @@ Usage:
from __future__ import annotations
import argparse
import json
import logging
import re
import uuid
import os
import sys
import time
@@ -112,63 +116,78 @@ def build_document(row: dict) -> dict:
return doc
def publish_collection(client, rows: list[dict]) -> str:
"""Validate a new collection before moving the alias; keep rollback data.
Caller holds the database advisory lock across reading and publication so
overlapping school-data DAGs cannot publish or prune each other's work.
Failed drafts are left for the next successful publication to prune: a lost
alias-update response must never cause deletion of a potentially live index.
"""
if not rows or len({r["urn"] for r in rows}) != len(rows):
raise ValueError("Search source must contain nonempty, unique school URNs")
name = f"schools_{int(time.time())}_{uuid.uuid4().hex[:12]}"
try:
previous = client.aliases["schools"].retrieve()["collection_name"]
except typesense.exceptions.ObjectNotFound:
previous = None
client.collections.create({**COLLECTION_SCHEMA, "name": name})
for i in range(0, len(rows), 500):
batch = [build_document(r) for r in rows[i:i + 500]]
results = client.collections[name].documents.import_(batch, {"action": "upsert"})
if isinstance(results, str):
results = [json.loads(line) for line in results.splitlines() if line.strip()]
if len(results) != len(batch) or any(r.get("success") is not True for r in results):
raise ValueError("Search import failed; live alias unchanged")
if client.collections[name].retrieve()["num_documents"] != len(rows):
raise ValueError("Search document count mismatch; live alias unchanged")
client.aliases.upsert("schools", {"collection_name": name})
# Retention is best-effort and must not make successful publication fail.
try:
keep = {name, previous}
keep.update(a["collection_name"] for a in client.aliases.retrieve()["aliases"])
for collection in client.collections.retrieve():
old = collection["name"]
if old not in keep and re.fullmatch(r"schools_\d+(?:_[0-9a-f]+)?", old):
client.collections[old].delete()
except Exception:
logging.getLogger(__name__).exception("Search published, but old collection cleanup failed")
return name
def sync(typesense_url: str, api_key: str):
from urllib.parse import urlparse
url = urlparse(typesense_url)
client = typesense.Client({
"nodes": [{"host": typesense_url.split("//")[-1].split(":")[0],
"port": typesense_url.split(":")[-1],
"protocol": "http"}],
"nodes": [{"host": url.hostname, "port": str(url.port or 8108),
"protocol": url.scheme}],
"api_key": api_key,
"connection_timeout_seconds": 10,
})
# Create timestamped collection for zero-downtime swap
ts = int(time.time())
collection_name = f"schools_{ts}"
print(f"Creating collection: {collection_name}")
schema = {**COLLECTION_SCHEMA, "name": collection_name}
client.collections.create(schema)
# Fetch data from marts — join fact_performance if it exists
conn = get_db_connection()
with conn.cursor(cursor_factory=psycopg2.extras.RealDictCursor) as cur:
# Check whether the merged fact table exists
cur.execute("""
SELECT table_name FROM information_schema.tables
WHERE table_schema = 'marts' AND table_name = 'fact_performance'
""")
has_fact_performance = cur.fetchone() is not None
query = QUERY_BASE
if has_fact_performance:
query = query.replace(
"l.longitude as lng",
"l.longitude as lng,\n p.rwm_expected_pct,\n p.progress_8_score",
)
query += QUERY_PERFORMANCE_JOIN
cur.execute(query)
rows = cur.fetchall()
conn.close()
print(f"Indexing {len(rows)} schools...")
# Batch import
batch_size = 500
for i in range(0, len(rows), batch_size):
batch = [build_document(r) for r in rows[i : i + batch_size]]
client.collections[collection_name].documents.import_(batch, {"action": "upsert"})
print(f" Indexed {min(i + batch_size, len(rows))}/{len(rows)}")
# Swap alias
print("Swapping alias 'schools' → new collection")
try:
client.aliases.upsert("schools", {"collection_name": collection_name})
except Exception:
# If alias doesn't exist yet, create it
client.aliases.upsert("schools", {"collection_name": collection_name})
print("Done.")
with conn.cursor(cursor_factory=psycopg2.extras.RealDictCursor) as cur:
# Session-scoped lock is released even on errors when conn closes.
cur.execute("SELECT pg_advisory_lock(731042019)")
cur.execute("""
SELECT table_name FROM information_schema.tables
WHERE table_schema = 'marts' AND table_name = 'fact_performance'
""")
has_fact_performance = cur.fetchone() is not None
query = QUERY_BASE
if has_fact_performance:
query = query.replace(
"l.longitude as lng",
"l.longitude as lng, p.rwm_expected_pct, p.progress_8_score",
)
query += QUERY_PERFORMANCE_JOIN
cur.execute(query)
rows = cur.fetchall()
name = publish_collection(client, rows)
print(f"Published {len(rows)} schools in {name}")
finally:
conn.close()
def main():
+92
View File
@@ -0,0 +1,92 @@
import importlib.util
from pathlib import Path
from unittest.mock import MagicMock
import pytest
@pytest.fixture
def sync_module(monkeypatch):
folder = Path(__file__).resolve().parents[1] / 'scripts'
monkeypatch.syspath_prepend(str(folder))
spec = importlib.util.spec_from_file_location('sync_typesense', folder / 'sync_typesense.py')
module = importlib.util.module_from_spec(spec)
spec.loader.exec_module(module)
return module
@pytest.fixture
def client():
c = MagicMock()
c.aliases.__getitem__.return_value.retrieve.return_value = {'collection_name': 'schools_2'}
c.aliases.retrieve.return_value = {'aliases': [{'collection_name': 'schools_99'}]}
c.collections.__getitem__.return_value.documents.import_.return_value = [{'success': True}]
c.collections.__getitem__.return_value.retrieve.return_value = {'num_documents': 1}
c.collections.retrieve.return_value = [{'name': n} for n in ['schools_1', 'schools_2', 'schools_99', 'unrelated']]
return c
@pytest.fixture
def rows():
return [{'urn': 100001, 'school_name': 'Example', 'phase_code': 2,
'school_type_code': 1, 'local_authority': 'Testshire', 'postcode': 'TS1 1AA',
'total_pupils': 250}]
@pytest.mark.parametrize('results', [[{'success': False}], [], [{'success': True}, {'success': True}]])
def test_partial_import_never_moves_alias(sync_module, client, rows, results):
client.collections.__getitem__.return_value.documents.import_.return_value = results
with pytest.raises(ValueError):
sync_module.publish_collection(client, rows)
client.aliases.upsert.assert_not_called()
client.collections.__getitem__.return_value.delete.assert_not_called()
def test_count_mismatch_does_not_publish(sync_module, client, rows):
client.collections.__getitem__.return_value.retrieve.return_value = {'num_documents': 0}
with pytest.raises(ValueError):
sync_module.publish_collection(client, rows)
client.aliases.upsert.assert_not_called()
@pytest.mark.parametrize('empty', [True, False])
def test_invalid_source_never_creates_collection(sync_module, client, rows, empty):
with pytest.raises(ValueError):
sync_module.publish_collection(client, [] if empty else rows + rows)
client.collections.create.assert_not_called()
def test_success_retains_previous_and_other_live_aliases(sync_module, client, rows):
# Distinct mock per collection allows checking exactly which one was deleted.
collections = {}
def get(name):
if name not in collections:
c = MagicMock()
c.documents.import_.return_value = '{"success":true}\n'
c.retrieve.return_value = {'num_documents': 1}
collections[name] = c
return collections[name]
client.collections.__getitem__.side_effect = get
name = sync_module.publish_collection(client, rows)
client.aliases.upsert.assert_called_once_with('schools', {'collection_name': name})
assert set(collections) == {name, 'schools_1'}
collections['schools_1'].delete.assert_called_once()
def test_uncertain_alias_update_does_not_delete_candidate(sync_module, client, rows):
client.aliases.upsert.side_effect = RuntimeError('response lost')
with pytest.raises(RuntimeError):
sync_module.publish_collection(client, rows)
client.collections.__getitem__.return_value.delete.assert_not_called()
def test_sync_closes_database_when_publication_fails(sync_module, monkeypatch):
conn = MagicMock()
monkeypatch.setattr(sync_module, 'get_db_connection', lambda: conn)
monkeypatch.setattr(sync_module.typesense, 'Client', lambda _: MagicMock())
def fail(*args): raise ValueError('import rejected')
monkeypatch.setattr(sync_module, 'publish_collection', fail)
with pytest.raises(ValueError):
sync_module.sync('http://localhost:8108', 'dummy')
conn.close.assert_called_once()
statements = [call.args[0] for call in conn.cursor.return_value.__enter__.return_value.execute.call_args_list]
assert statements[0] == 'SELECT pg_advisory_lock(731042019)'
+174
View File
@@ -0,0 +1,174 @@
"""Release identity checks and promotion of the exact digests that passed E2E.
Uses only the standard library and Docker Buildx. No registry mutation happens
until every image in the set has been resolved and its build labels validated.
"""
import argparse
import json
import os
from pathlib import Path
import re
import subprocess
import time
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
COMPONENTS = ('BACKEND', 'FRONTEND', 'PIPELINE')
def docker(*args):
return subprocess.check_output(['docker', 'buildx', 'imagetools', *args], text=True).strip()
def check_sha(sha):
if not re.fullmatch(r'[0-9a-f]{40}', sha):
raise ValueError('Expected a full commit SHA')
return sha
def check_digest(digest):
if not re.fullmatch(r'sha256:[0-9a-f]{64}', digest):
raise ValueError('Expected an immutable image digest')
return digest
def image_identity(ref):
image = json.loads(docker('inspect', ref, '--format', '{{json .Image}}'))
configs = [image] if 'config' in image else list(image.values())
identities = set()
for config in configs:
labels = config['config']['Labels']
identities.add((labels['io.schoolcompare.commit'], labels['io.schoolcompare.build-id']))
if len(identities) != 1:
raise ValueError('Image platforms disagree about their release identity')
return next(iter(identities))
def resolve_images(sha, verified=False, expected_build_id=None):
check_sha(sha)
refs = []
build_ids = set()
for component in COMPONENTS:
image = f"{os.environ['REGISTRY']}/{os.environ[component + '_IMAGE_NAME']}"
if verified:
manifest = json.loads(docker('inspect', f'{image}:verified-{sha}', '--format', '{{json .Manifest}}'))
digest = check_digest(manifest['digest'])
else:
digest = check_digest(os.environ[component + '_DIGEST'])
ref = f'{image}@{digest}'
actual_sha, build_id = image_identity(ref)
if actual_sha != sha or not re.fullmatch(r'[0-9a-f]{32}', build_id):
raise ValueError(f'Unrecognised release identity for {component}')
if expected_build_id is not None and build_id != expected_build_id:
raise ValueError(f'Build identity mismatch for {component}')
build_ids.add(build_id)
refs.append((image, ref))
if len(build_ids) != 1:
raise ValueError('Refusing a mixed image set')
return {'sha': sha, 'build_id': build_ids.pop(), 'images': refs}
def verify(sha, build_id):
release = resolve_images(sha, expected_build_id=build_id)
for image, ref in release['images']:
docker('create', '-t', f'{image}:verified-{sha}', ref)
return release
def promote(sha):
release = resolve_images(sha, verified=True)
# Resolve all targets first; never discover a missing candidate halfway through.
for image, _ in release['images']:
try:
docker('create', '-t', f'{image}:prod-previous', f'{image}:prod')
except subprocess.CalledProcessError:
print(f'No rollback pointer saved for {image}', flush=True)
for image, ref in release['images']:
docker('create', '-t', f'{image}:prod', ref)
return release
def matches(payload, sha, build_id):
return isinstance(payload, dict) and all(
payload.get(component) == {'sha': sha, 'build_id': build_id}
for component in ('frontend', 'backend'))
def describe_identity(payload):
"""Log only release fields, never arbitrary response bodies or secret URLs."""
if not isinstance(payload, dict):
return 'Invalid release response: expected a JSON object'
identities = []
for component in ('frontend', 'backend'):
identity = payload.get(component)
if not isinstance(identity, dict):
identities.append(f'{component}=missing or invalid')
continue
values = []
for field, length in (('sha', 40), ('build_id', 32)):
value = identity.get(field)
valid = isinstance(value, str) and (
value == 'development' or re.fullmatch(r'[0-9a-f]{' + str(length) + '}', value))
values.append(f'{field}={value if valid else "missing or invalid"}')
identities.append(f'{component}: {", ".join(values)}')
return 'Release mismatch: ' + '; '.join(identities)
def wait(base_url, sha, build_id, timeout):
check_sha(sha)
if not re.fullmatch(r'[0-9a-f]{32}', build_id):
raise ValueError('Missing expected build identity')
deadline = time.monotonic() + timeout
last_observation = 'No response received'
print(f'Waiting for deployed release {sha} / {build_id}', flush=True)
while time.monotonic() < deadline:
try:
req = Request(f'{base_url.rstrip("/")}/release.json?check={time.time_ns()}',
headers={'Cache-Control': 'no-cache',
'User-Agent': 'SchoolCompare-Release-Check/1.0',
'Accept': 'application/json'})
with urlopen(req, timeout=min(10, max(.1, deadline - time.monotonic()))) as response:
payload = json.load(response)
if matches(payload, sha, build_id):
print(f'Verified deployed release {sha} / {build_id}')
return
observation = describe_identity(payload)
except HTTPError as exc:
observation = f'Release endpoint returned HTTP {exc.code}'
exc.close()
except URLError as exc:
observation = f'Release endpoint connection failed ({type(exc.reason).__name__})'
except OSError as exc:
observation = f'Release endpoint request failed ({type(exc).__name__})'
except ValueError:
observation = 'Release endpoint returned invalid JSON or request configuration'
if observation != last_observation:
print(observation, flush=True)
last_observation = observation
time.sleep(min(5, max(0, deadline - time.monotonic())))
raise RuntimeError('Deployment did not report the expected frontend/backend release '
f'{sha} / {build_id}. Last observation: {last_observation}')
def main():
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument('action', choices=['wait', 'verify', 'promote'])
parser.add_argument('--timeout', type=float, default=300)
parser.add_argument('--release', type=Path)
parser.add_argument('--output', type=Path)
args = parser.parse_args()
identity = json.loads(args.release.read_text()) if args.release else {
'sha': os.environ.get('EXPECTED_SHA', ''),
'build_id': os.environ.get('EXPECTED_BUILD_ID', ''),
}
if args.action == 'wait':
wait(os.environ['BASE_URL'], identity['sha'], identity['build_id'], args.timeout)
return
result = (verify(identity['sha'], identity['build_id']) if args.action == 'verify'
else promote(identity['sha']))
if args.output:
args.output.write_text(json.dumps(result))
if __name__ == '__main__':
main()
+130
View File
@@ -0,0 +1,130 @@
import json
from io import BytesIO
from urllib.error import HTTPError, URLError
from unittest.mock import Mock
import pytest
from scripts.ci import release
SHA = 'a' * 40
BUILD = 'b' * 32
DIGESTS = ['sha256:' + c * 64 for c in '123']
@pytest.fixture
def docker(monkeypatch):
monkeypatch.setenv('REGISTRY', 'registry.example')
refs = {}
for component, digest in zip(release.COMPONENTS, DIGESTS):
monkeypatch.setenv(component + '_IMAGE_NAME', component.lower())
monkeypatch.setenv(component + '_DIGEST', digest)
refs[f'registry.example/{component.lower()}'] = digest
def run(*args):
if args[0] == 'create': return ''
if args[-1] == '{{json .Manifest}}':
return json.dumps({'digest': refs[args[1].split(':')[0]]})
return json.dumps({'config': {'Labels': {'io.schoolcompare.commit': SHA,
'io.schoolcompare.build-id': BUILD}}})
mock = Mock(side_effect=run)
monkeypatch.setattr(release, 'docker', mock)
return mock
def test_wrong_deployed_build_is_rejected_even_at_same_commit():
assert not release.matches({'frontend': {'sha': SHA, 'build_id': BUILD},
'backend': {'sha': SHA, 'build_id': 'c' * 32}}, SHA, BUILD)
assert release.matches({c: {'sha': SHA, 'build_id': BUILD} for c in ('frontend', 'backend')}, SHA, BUILD)
def test_verification_tags_the_captured_digests(docker):
release.verify(SHA, BUILD)
creates = [c.args for c in docker.call_args_list if c.args[0] == 'create']
assert len(creates) == 3
for call, digest in zip(creates, DIGESTS):
assert call[-1].endswith('@' + digest)
assert call[2].endswith(':verified-' + SHA)
def test_promotion_resolves_all_verified_images_before_mutation(docker):
result = release.promote(SHA)
assert result['build_id'] == BUILD
calls = [c.args for c in docker.call_args_list]
first_write = next(i for i, c in enumerate(calls) if c[0] == 'create')
assert first_write == 6 # each of three candidates needs manifest + config
assert all(c[-1].endswith('@' + d) for c, d in zip(calls[-3:], DIGESTS))
def test_mixed_builds_fail_before_any_tag_is_changed(docker, monkeypatch):
identities = iter([(SHA, BUILD), (SHA, 'c' * 32), (SHA, BUILD)])
monkeypatch.setattr(release, 'image_identity', lambda _: next(identities))
with pytest.raises(ValueError, match='mixed'):
release.promote(SHA)
assert not any(c.args[0] == 'create' for c in docker.call_args_list)
def test_missing_candidate_fails_before_any_tag_is_changed(docker):
docker.side_effect = RuntimeError('missing verified tag')
with pytest.raises(RuntimeError): release.promote(SHA)
assert not any(c.args[0] == 'create' for c in docker.call_args_list)
@pytest.fixture
def poll(monkeypatch):
now = [0.0]
monkeypatch.setattr(release.time, 'monotonic', lambda: now[0])
monkeypatch.setattr(release.time, 'sleep', lambda seconds: now.__setitem__(0, now[0] + seconds))
opener = Mock()
monkeypatch.setattr(release, 'urlopen', opener)
return opener
def response(payload):
return BytesIO(json.dumps(payload).encode())
def test_wait_identifies_its_client_and_retries_until_both_services_match(poll, capsys):
poll.side_effect = [
HTTPError('https://secret.example', 503, 'unavailable', {}, None),
response({'frontend': {'sha': SHA, 'build_id': BUILD},
'backend': {'sha': SHA, 'build_id': 'c' * 32}}),
response({component: {'sha': SHA, 'build_id': BUILD}
for component in ('frontend', 'backend')}),
]
release.wait('https://secret.example/', SHA, BUILD, 15)
assert poll.call_count == 3
request = poll.call_args.args[0]
assert request.get_header('User-agent') == 'SchoolCompare-Release-Check/1.0'
assert request.get_header('Cache-control') == 'no-cache'
assert request.get_header('Accept') == 'application/json'
assert '/release.json?check=' in request.full_url
output = capsys.readouterr().out
assert 'HTTP 503' in output
assert 'backend: sha=' + SHA + ', build_id=' + 'c' * 32 in output
assert 'Verified deployed release' in output
assert 'secret.example' not in output
@pytest.mark.parametrize('failure, expected', [
(lambda: HTTPError('https://secret.example', 403, 'secret response', {}, None), 'HTTP 403'),
(lambda: URLError(OSError('secret address')), 'connection failed (OSError)'),
(lambda: TimeoutError('secret address'), 'request failed (TimeoutError)'),
(lambda: BytesIO(b'<html>secret response</html>'), 'invalid JSON'),
(lambda: response([]), 'expected a JSON object'),
(lambda: response({'frontend': {'sha': 'secret response'}}), 'missing or invalid'),
])
def test_wait_timeout_reports_last_failure_without_leaking_response_or_url(poll, capsys, failure, expected):
poll.side_effect = lambda *args, **kwargs: result_or_raise(failure())
with pytest.raises(RuntimeError) as error:
release.wait('https://secret.example', SHA, BUILD, 10)
assert expected in str(error.value)
assert SHA in str(error.value)
assert BUILD in str(error.value)
output = capsys.readouterr().out
assert sum(expected in line for line in output.splitlines()) == 1
assert 'secret' not in output + str(error.value)
assert poll.call_count == 2
def result_or_raise(result):
if isinstance(result, Exception):
raise result
return result
+34
View File
@@ -0,0 +1,34 @@
"""Check the dependency graph that ties tested digests to deployable images."""
from pathlib import Path
import yaml
ROOT = Path(__file__).resolve().parents[3]
def test_staging_verifies_identity_before_and_after_journeys():
workflow = yaml.safe_load((ROOT / '.gitea/workflows/deploy.yml').read_text())
assert workflow['concurrency'] == {'group': 'staging-release', 'cancel-in-progress': False}
jobs = workflow['jobs']
for component in ('backend', 'frontend', 'pipeline'):
job = jobs['build-' + component]
assert 'prepare' in job['needs']
assert job['outputs']['digest'] == '${{ steps.build.outputs.digest }}'
build = next(step for step in job['steps'] if step.get('id') == 'build')
assert 'BUILD_ID=${{ needs.prepare.outputs.build_id }}' in build['with']['build-args']
steps = jobs['e2e-staging']['steps']
runs = [step.get('run', '') for step in steps]
test = runs.index('npx playwright test')
assert 'release.py wait' in runs[test - 1]
assert 'release.py wait' in runs[test + 1]
assert 'release.py verify' in runs[-1]
for component in ('backend', 'frontend', 'pipeline'):
assert 'build-' + component in jobs['e2e-staging']['needs']
assert component.upper() + '_DIGEST' in steps[-1]['env']
def test_promotion_uses_verified_digest_resolver_and_build_identity_poll():
workflow = yaml.safe_load((ROOT / '.gitea/workflows/promote.yml').read_text())
steps = workflow['jobs']['promote-prod']['steps']
runs = [step.get('run', '') for step in steps]
assert 'python3 scripts/ci/release.py promote --output release.json' in runs
assert runs[-1] == 'python3 scripts/ci/release.py wait --release release.json'
+3
View File
@@ -1,5 +1,8 @@
#!/usr/bin/env python3
"""
Historical standalone CSV utility; not part of the managed Meltano/dbt pipeline.
See docs/LEGACY_CODE.md before using it for current school data.
Data Download Helper Script
This script provides instructions and utilities for downloading
+3
View File
@@ -1,5 +1,8 @@
#!/usr/bin/env python3
"""
Historical standalone CSV utility; not part of the managed Meltano/dbt pipeline.
See docs/LEGACY_CODE.md before using it for current school data.
Fetch real school performance data from UK Government sources.
This script downloads KS2 (Key Stage 2) primary school data from:
-184
View File
@@ -1,184 +0,0 @@
#!/usr/bin/env python3
"""
Geocode all school postcodes and update the database.
This script should be run as a weekly cron job to ensure all schools
have up-to-date latitude/longitude coordinates.
Usage:
python scripts/geocode_schools.py [--force]
Options:
--force Re-geocode all postcodes, even if already geocoded
Crontab example (run every Sunday at 2am):
0 2 * * 0 cd /path/to/school_compare && /path/to/venv/bin/python scripts/geocode_schools.py >> /var/log/geocode_schools.log 2>&1
"""
import argparse
import sys
from datetime import datetime
from pathlib import Path
from typing import Dict, Tuple
import requests
# Add parent directory to path for imports
sys.path.insert(0, str(Path(__file__).parent.parent))
from backend.database import SessionLocal
from backend.models import School
def geocode_postcodes_bulk(postcodes: list) -> Dict[str, Tuple[float, float]]:
"""
Geocode postcodes in bulk using postcodes.io API.
Returns dict of postcode -> (latitude, longitude).
"""
results = {}
valid_postcodes = [
p.strip().upper()
for p in postcodes
if p and isinstance(p, str) and len(p.strip()) >= 5
]
valid_postcodes = list(set(valid_postcodes))
if not valid_postcodes:
return results
batch_size = 100
total_batches = (len(valid_postcodes) + batch_size - 1) // batch_size
for i, batch_start in enumerate(range(0, len(valid_postcodes), batch_size)):
batch = valid_postcodes[batch_start : batch_start + batch_size]
print(f" Geocoding batch {i + 1}/{total_batches} ({len(batch)} postcodes)...")
try:
response = requests.post(
"https://api.postcodes.io/postcodes",
json={"postcodes": batch},
timeout=30,
)
if response.status_code == 200:
data = response.json()
for item in data.get("result", []):
if item and item.get("result"):
pc = item["query"].upper()
lat = item["result"].get("latitude")
lon = item["result"].get("longitude")
if lat and lon:
results[pc] = (lat, lon)
else:
print(f" Warning: API returned status {response.status_code}")
except Exception as e:
print(f" Warning: Geocoding batch failed: {e}")
return results
def geocode_schools(force: bool = False) -> None:
"""
Geocode all schools in the database.
Args:
force: If True, re-geocode all postcodes even if already geocoded
"""
print(f"\n{'='*60}")
print(f"School Geocoding Job - {datetime.now().isoformat()}")
print(f"{'='*60}\n")
db = SessionLocal()
try:
# Get schools that need geocoding
if force:
schools = db.query(School).filter(School.postcode.isnot(None)).all()
print(f"Force mode: Processing all {len(schools)} schools with postcodes")
else:
schools = db.query(School).filter(
School.postcode.isnot(None),
(School.latitude.is_(None)) | (School.longitude.is_(None))
).all()
print(f"Found {len(schools)} schools without coordinates")
if not schools:
print("No schools to geocode. Exiting.")
return
# Extract unique postcodes
postcodes = list(set(
s.postcode.strip().upper()
for s in schools
if s.postcode
))
print(f"Unique postcodes to geocode: {len(postcodes)}")
# Geocode in bulk
print("\nGeocoding postcodes...")
geocoded = geocode_postcodes_bulk(postcodes)
print(f"Successfully geocoded: {len(geocoded)} postcodes")
# Update database
print("\nUpdating database...")
updated_count = 0
failed_count = 0
for school in schools:
if not school.postcode:
continue
pc_upper = school.postcode.strip().upper()
coords = geocoded.get(pc_upper)
if coords:
school.latitude = coords[0]
school.longitude = coords[1]
updated_count += 1
else:
failed_count += 1
db.commit()
print(f"\nResults:")
print(f" - Updated: {updated_count} schools")
print(f" - Failed (invalid/not found): {failed_count} postcodes")
# Summary stats
total_with_coords = db.query(School).filter(
School.latitude.isnot(None),
School.longitude.isnot(None)
).count()
total_schools = db.query(School).count()
print(f"\nDatabase summary:")
print(f" - Total schools: {total_schools}")
print(f" - Schools with coordinates: {total_with_coords}")
print(f" - Coverage: {100*total_with_coords/total_schools:.1f}%")
except Exception as e:
print(f"Error during geocoding: {e}")
db.rollback()
raise
finally:
db.close()
print(f"\n{'='*60}")
print(f"Geocoding job completed - {datetime.now().isoformat()}")
print(f"{'='*60}\n")
def main():
parser = argparse.ArgumentParser(
description="Geocode school postcodes and update database"
)
parser.add_argument(
"--force",
action="store_true",
help="Re-geocode all postcodes, even if already geocoded"
)
args = parser.parse_args()
geocode_schools(force=args.force)
if __name__ == "__main__":
main()
-68
View File
@@ -1,68 +0,0 @@
#!/usr/bin/env python3
"""
CLI script for manual database migration.
Usage:
python scripts/migrate_csv_to_db.py [--drop] [--geocode]
Options:
--drop Drop existing tables before migration (full reimport)
--geocode Geocode postcodes (requires network access)
"""
import sys
from pathlib import Path
# Add parent directory to path for imports
sys.path.insert(0, str(Path(__file__).parent.parent))
import argparse
from backend.config import settings
from backend.database import Base, engine, init_db, set_db_schema_version
from backend.migration import load_csv_data, migrate_data, run_full_migration
from backend.version import SCHEMA_VERSION
def main():
parser = argparse.ArgumentParser(
description="Migrate CSV data to PostgreSQL database"
)
parser.add_argument(
"--drop", action="store_true", help="Drop existing tables before migration"
)
parser.add_argument("--geocode", action="store_true", help="Geocode postcodes")
args = parser.parse_args()
print("=" * 60)
print("School Data Migration: CSV -> PostgreSQL")
print("=" * 60)
print(f"\nDatabase: {settings.database_url.split('@')[-1]}")
print(f"Data directory: {settings.data_dir}")
print(f"Target schema version: {SCHEMA_VERSION}")
if args.drop:
print("\nRunning full migration (drop + reimport)...")
success = run_full_migration(geocode=args.geocode)
else:
print("\nCreating tables (preserving existing data)...")
init_db()
print("\nLoading CSV data...")
df = load_csv_data(settings.data_dir)
if df.empty:
print("No data found to migrate!")
return 1
migrate_data(df, geocode=args.geocode)
success = True
if success:
# Ensure schema_version table exists
init_db()
set_db_schema_version(SCHEMA_VERSION)
print(f"\nSchema version set to {SCHEMA_VERSION}")
return 0 if success else 1
if __name__ == "__main__":
sys.exit(main())