Compare commits

...
Author SHA1 Message Date
TudorandClaude Opus 5 077aca6008 test(e2e): wait for the carousel's smooth scroll to settle before measuring it
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 32s
Second failure of the same assertion, and the first fix addressed a real but
different problem. This is the one that was actually producing the number.

The arrows scroll with `behavior: 'smooth'`. The test polled for "has it moved
at all" — satisfied 50ms in, at 13px of a 1300px journey — then recorded the
offset, clicked, and recorded again. The row was still travelling throughout,
so the delta it measured was the tail of the arrow's animation, not the effect
of the selection. Hence a deterministic 659, roughly half of the 1317 this
school's row scrolls.

Measured against staging rather than reasoned about: the animation runs about
700ms, and the samples are in the helper's comment.

settledScrollLeft waits for two identical readings before trusting one. The app
was never at fault — driving staging by hand, the offset holds at 1317 across
the selection, exactly as intended.

Verified against staging both ways: main's version of this test fails there,
this version passes, along with all three mobile widths.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 15:13:32 +01:00
tudor 9dba5ff1ff Merge pull request 'test(e2e): fix the nearby-schools journey clicking a card it scrolled past' (#152) from fix/nearby-scroll-journey-clicks-offscreen-card into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m23s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 25s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m49s
Reviewed-on: #152
2026-09-22 13:50:25 +00:00
TudorandClaude Opus 5 f530a912bc test(e2e): stop the nearby-schools journey clicking a card it scrolled past
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 21s
The assertion that adding to compare does not reset the carousel failed on
staging: 537 → 2. The app was not at fault. Playwright scrolls a target into
view before clicking, and the test clicked the FIRST card's button after
paging the row to the end — so Playwright scrolled the container back to the
start, and the assertion measured that.

Reproduced on a static page with no React on it: a snap scroller at 615,
Playwright clicks the off-screen first card, scrollLeft becomes 2. Scroll-snap
was ruled out first — mandatory, proximity and no-snap all behave identically
when the button mutates in place.

Now clicks the last card's button, which is visible at the end of the travel,
and allows a few pixels for snap and sub-pixel adjustment while still failing
on a reset to the start. The property was never actually under test before;
it is now.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 14:19:00 +01:00
tudor 029fe8d8a6 Merge pull request 'fix: order nearby schools by distance, not by how alike they are' (#151) from fix/nearby-schools-order-by-distance into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m24s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 30s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m51s
Reviewed-on: #151
2026-09-22 13:09:38 +00:00
TudorandClaude Opus 5 cd1c5d1e1a docs: revise the spec to the design that survived staging
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 57s
The spec described the tier system as the design of record. It is gone, so
the document was describing something the code deliberately does not do.

The revision note and the "why not, having built it the other way first"
passage are kept rather than overwritten. The mistake is the instructive part:
treating a preference as a constraint inverted the ranking, and the stopping
rule added to prevent weak distant matches is what guaranteed six Catholic
schools and no community school down the road. A spec that quietly presents the
second design as the plan teaches nobody why the first one failed.

The mockup link is annotated as one revision behind rather than silently left
to look current.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
TudorandClaude Opus 5 cd6a45bf7d refactor: rename similar → nearby, so the code says what the section does
The section ranks on distance and is headed "Other schools nearby", but every
identifier still called it "similar" — the exact drift that leaves a later
reader trusting a name over the behaviour.

Mechanical: files, the module, the payload key, the type, the components, the
prop. No behaviour change; the suites are unchanged in count and still green.
Free to do now because #150 has not merged, so the payload key rename needs no
lockstep deploy. Uses of "similar" that are ordinary English — progress
measures compared to similar pupils, and unrelated comments — are untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
TudorandClaude Opus 5 5c0ccc693d fix: order nearby schools by distance, not by how alike they are
Reported from staging: a Catholic primary showed six Catholic primaries, none
of them close enough to be a real option, and omitted the community school
down the road.

Three causes, compounding. Ranking put tier before distance, so a faith match
at 2.9 miles outranked a community school at 0.3. The ENOUGH=3 stopping rule —
added so a cap of six would not drag in weak distant matches — filled the row
from the best tier before it ever widened, which is what made every card
Catholic. And a 3-mile tier-1 radius is sane for a secondary and most of a city
for a primary, whose catchments are routinely under a mile.

The premise was backwards. For a parent, distance is a constraint and intake is
a preference; a school beyond a primary catchment is not a weaker option, it is
not an option. So distance now decides the order and nothing else does. The
hard filters are untouched — they were always where the defensibility lived.
Similarity survives as chips on the card: reported, so a reader applies their
own weighting, rather than ranked, so we apply ours for them.

Reach is capped per phase (primary 2, secondary 6, post-16 10) as a sanity
bound, not a target: ordering already handles density, so the cap only decides
what happens where an area is sparse. A primary with nothing inside two miles
now renders no section, which is the honest answer.

Deleted: the tier system, the stopping rule, the tier-dependent lede, the
`tier` field, the tier-3 fallback chip and its style. select_similar also stops
taking is_secondary — it reads the phase from the subject's own row, so no
caller can hand it one that disagrees with the data.

The heading is now "Other schools nearby". The hard filters still guarantee a
comparable set, but nothing ranks on likeness, so the heading no longer says it
does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 13:06:32 +01:00
tudor 151cf4bc80 Merge pull request 'feat: similar schools nearby on the detail page' (#150) from feat/similar-schools-nearby into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 44s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m25s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m14s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 6s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 42s
Reviewed-on: #150
2026-09-22 05:53:06 +00:00
TudorandClaude Opus 5 83dc5ae5dc docs: correct the metric rule for "16 plus"
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m11s
The spec said the metric follows the template the page renders. That is no
longer true for GIAS phase 6: a sixth-form college renders the primary
template but is matched, correctly, against secondaries. The rule is phase
group membership, decided once in is_secondary_phase — and the section's lede
noun comes from the school's phase rather than its template for the same
reason.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 06:41:59 +01:00
TudorandClaude Opus 5 2175dccb7c fix(api): treat "16 plus" as secondary, the way PHASE_GROUPS already does
PR Checks / Frontend Typecheck + Tests (pull_request) Canceled after 9s
PR Checks / Backend Smoke (pull_request) Canceled after 0s
PR Checks / Build Backend (no push) (pull_request) Canceled after 0s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 0s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
GIAS phase 6 is "16 plus", and PHASE_GROUPS deliberately files it in the
secondary group. The payload helper decided the same question with
`"secondary" in phase_text`, which that value does not satisfy — so a
sixth-form college was handed the primary bucket and offered infant schools
as its peers, with the KS2 metric key to label them. No crash; just a page
confidently showing the wrong schools.

The decision now lives in similar_schools.is_secondary_phase, beside the
PHASE_GROUPS bucket it selects from, so the two cannot drift again. A test
pins them together.

The same binary assumption had a second output. computeSchoolFlags tests for
the substring too, so a 16-plus school renders the primary template, and the
composer was labelling the section from the template: "Other primary schools
near <sixth form college>" above a row of secondaries. The section now
derives its noun from the school's own phase, which also removes the
duplicated wording from both composers. A 16-plus school's candidates span
the whole secondary group, so no single noun fits and it gets the honest
general one.

Reported in review on #150.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-22 06:41:49 +01:00
TudorandClaude Opus 5 bd2a6c385b test(e2e): cover the similar-schools section, compare hand-off and mobile widths
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m13s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 37s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 3m1s
Two things here cannot be covered anywhere else. jsdom has no layout, so
scrollWidth and clientWidth are both 0 and the arrows' disabled state can only
be measured by a real engine. And the scroll position surviving a selection is
DOM state rather than React state, so only a real browser can prove the row
does not jump back when the footer re-renders.

MOBILE.md asks for a Playwright width check and records that it was not written
because Playwright was not in the project. It is — this suite — so the check
exists now, scoped to the page this feature touches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:43:38 +01:00
TudorandClaude Opus 5 4e0d8bcf87 feat(web): render similar schools on both detail templates
Inside SchoolDetailShell rather than after it, because the sticky nav's
scroll-spy finds sections with getElementById and can only reach one that
lives in the shell. Last in the order, and last in the nav, because the two
must agree or the nav links to an anchor that was never rendered.

hasSimilarSchools is optional on NavItemsInput, matching hasLocation beside
it: absent has to mean "no section", and making it required would have
churned ten unrelated call sites for no added safety.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:42:59 +01:00
TudorandClaude Opus 5 91314a80b5 feat(web): similar schools section, honest about what it matched
Server-rendered cards inside a client carousel that scrolls rather than
paginates, so all six links stay in the initial HTML and the row still works
with JavaScript off. Three client islands, split by what each needs: a school,
the whole selection, a DOM ref.

The lede claims a similar intake only when no card came from tier 3, chips
list what a school actually shares, a missing figure reads "Not published",
and the neighbour's number carries no valence colour — green and terracotta
mean "against England" everywhere else, and colouring it here would read as
ranking the neighbours.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:41:40 +01:00
TudorandClaude Opus 5 52b00ac752 feat(api): serve similar schools on the detail endpoint
Rides in the existing payload rather than a new endpoint: the page already
makes one server fetch for its data, and /school/[slug] regenerates weekly,
so the per-request cost is paid once per school per week.

Wrapped so a failure in selection never 500s a page that is otherwise
complete — the posture get_supplementary_data already takes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:40:35 +01:00
TudorandClaude Opus 5 b571d9c549 feat(api): decide which nearby schools a page may offer
Hard filters encode claims the section may not make — a selective school is
not an alternative to a non-selective one, a special school is not comparable
to a mainstream one, a Girls school is not an option for a Boys school's
reader — so they never relax. Soft preferences describe closeness of fit, so
they relax across three tiers, and only far enough to reach three; the
remaining slots up to six fill from the tiers already opened.

PHASE_GROUPS moves to schemas.py so this module can share it without
importing app, which would be a cycle.

_mask() exists because Series.apply on an empty Series returns a DataFrame,
and using that as a mask drops every column — so the next lookup raises
KeyError instead of yielding no rows. A special school with no special school
near it empties the frame at the provision filter, which is the ordinary case
for most special schools, so this was a crash on a common path rather than an
edge case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:40:10 +01:00
TudorandClaude Opus 5 8a23e3657d docs: meet the mobile baseline, which the carousel was not doing
Checked the section against MOBILE.md at its three reference widths instead
of assuming the breakpoints were enough. Two real failures at 360px.

The arrows sat in the heading's flex row, taking 96px from a 328px card and
crushing the lede into a four-line column — for a control that swiping
already provides. Below 640px they are now gone, the header is a single
column, one card shows at 86% so the next one peeks, and the affordance is
the right-edge scroll-fade MOBILE.md already documents for horizontal
scrollers. The fade lifts at the end of the travel, so the at-end state is
computed whether or not an arrow exists to consume it.

The arrow and add-to-compare buttons were 40px against a 44px floor. Both
are 44 now. A card title's own box is shorter, but its hit area is the whole
card through the ::after overlay, so it passes on the target that actually
receives the tap.

360, 390 and 430 now all report zero overflow, no failing tap targets and no
text under 11px. MOBILE.md wanted a Playwright width check and recorded that
Playwright was not in the project; it is, so the journey now carries one for
this page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:37:34 +01:00
TudorandClaude Opus 5 e4e8f02599 docs: six schools behind a carousel, and a rule for when to stop widening
Three schools is not a neighbourhood in inner London, so the cap is six with
three visible and arrows for the rest.

Raising the cap exposes something the old cap hid. Tiers exist to reach a
usable set, and with six slots a naive loop would keep widening to fill them
— dragging in tier-3 schools ten miles away to sit beside three good matches
that had already earned the row. So tiers now stop relaxing once three are
found, and the remaining slots are filled only from the tiers already used.
Four tier-1 matches never open tier 2.

The carousel scrolls a list rather than swapping a view: all six cards are in
the initial HTML, so every link stays crawlable and the row still scrolls with
JavaScript off. The arrows' edge test carries an 8px tolerance because the
scroller's focus-ring padding is the first snap position — a row at rest
reports scrollLeft 2, and an exact test for 0 left the back arrow live and
pointing nowhere. Caught in the mockup.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:35:11 +01:00
TudorandClaude Opus 5 b62dc17532 docs: drop the method disclosure, keep the one caveat that earns its place
The "how these schools are chosen" panel restated what the section already
shows — the phase in the lede, the shared characteristics on each card, the
distance above each name — so it cost space to say nothing new.

One line survives, and it is not a method note. A reader who sees "0.6 miles
away" and takes it for the walk has been misled by us, and no other element
on the card corrects that. The rest were claims the selection rules keep
true without narrating them.

Also records what happens past the third school: surplus matches are dropped
silently, because NearbyPlaces below already leads to the full lists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:29:07 +01:00
TudorandClaude Opus 5 4d7762d796 docs: implementation plan for similar schools nearby
Five tasks, each ending in a green test run and a commit: the pure
selection module, the endpoint key, the section and its two client
islands, the wiring into both templates, and the journey.

The selection logic gets its own module rather than another 200 lines in
app.py, which means the tier rules are testable against a synthetic frame
with no TestClient, no database and no monkeypatch. Moving PHASE_GROUPS
to schemas.py is what keeps that import acyclic.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:21:14 +01:00
TudorandClaude Opus 5 3650f7d8b7 docs: design for similar schools nearby on the detail page
A school page links outward to its places and never to another school.
This section adds that edge: three nearby schools of the same phase and a
comparable intake, each a crawlable link and each addable to the basket.

The design separates hard filters from soft preferences and never confuses
them. Selectivity, provision and opposite-sex intake are claims the section
cannot make, so they never relax, even where that means no section renders.
Gender and religious character describe closeness of fit, so they relax in
tiers — and the card states what actually survived rather than padding with
a match it did not earn.

Includes the mockup the design is drawn against, with all three tier states
live in both themes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-21 22:18:38 +01:00
tudor dfce308f1f Merge pull request 'fix(ci): make a failed release check say what it actually saw' (#149) from fix/release-check-diagnostics into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 1s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m24s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 28s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m9s
Reviewed-on: #149
2026-09-15 15:09:46 +00:00
TudorandClaude Opus 5 64ae71d7ab fix(ci): make a failed release check say what it actually saw
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m4s
The staging poller swallowed every failure identically, so a run that
timed out told us only that the expected release never appeared — not
whether the proxy refused us, the endpoint was down, or the containers
were still serving an older build. The public staging proxy also answers
403 to urllib's default user agent while the release endpoint is healthy,
which looked exactly like a deployment that never arrived.

Identify the poller, and report each distinct observation once: HTTP
status, connection failure type, invalid JSON, or the release identities
actually reported. The timeout error carries the last observation and the
identity it wanted. Responses and the base URL stay out of the logs —
only validated sha/build_id fields are echoed back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 16:02:14 +01:00
tudor 7ab084dd3a Merge pull request 'feat: verify what we actually deployed, and publish data all-or-nothing' (#148) from feat/release-identity-and-reliability-gate into main
Stage (build -> staging -> E2E gate) / prepare (push) Successful in 0s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m31s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m27s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Failing after 5m3s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Skipped
Reviewed-on: #148
2026-09-15 11:13:08 +00:00
TudorandClaude Opus 5 5ad1cbfb53 fix(api): bound the search candidate set instead of draining Typesense
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 39s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 6m44s
Fetching every match kept scoped searches correct but left the number of
round trips in the caller's hands: a one-letter query, or a deliberately
broad one, walked the whole collection a page at a time.

Cap the candidate set at 1,000 URNs — four pages — and return the
relevance-ordered prefix when the ceiling is hit. That is still far more
than one page, so the API's own authority, phase and postcode filters
keep the matches they need, while latency and upstream load stay bounded.
A capped query is logged so a genuinely truncated search is visible.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 11:21:45 +01:00
TudorandClaude Opus 5 0c901cd0d1 feat(ci): gate promotion on the image set that actually passed E2E
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m19s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 6m3s
Staging health polling asked only whether something answered HTTP 200 at
the base URL. It could not tell the new deployment from the old one, so
journeys could pass against the previous release, and concurrent merges
could move the staging tags underneath a run in flight.

Each staging run now mints a build ID and stamps all three images with
the commit and that ID, as labels and — for frontend and backend — as a
build-time JSON file that environment overrides cannot rewrite.
/release.json reports both identities uncached, and scripts/ci/release.py
polls for the expected pair before and after the journeys. Only then are
the captured build digests tagged verified-<sha>.

Promotion resolves those verified tags to immutable digests, revalidates
their labels, and refuses a mixed or incomplete set before any :prod tag
moves. The whole staging workflow shares one concurrency group with
cancellation disabled, so releases serialise.

The scripts are stdlib-only and unit-tested against mocked registry and
HTTP behaviour; PR checks now run the pipeline and CI suites too. The
runbook records what this cannot prove locally, and that the first
rollout needs a commit built by this workflow.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:17:50 +01:00
TudorandClaude Opus 5 7b41218e6e fix(web): show an outage as an outage, and drop superseded fetches
The home page caught every fetch failure and rendered its empty state,
so a backend outage looked like a site with no schools in it. School
pages turned any error into notFound(), which told visitors — and
crawlers — that a real school had ceased to exist. Place fetches did the
same by returning [] and null. Failures now reach a retryable error
boundary; only a genuine 404 still calls notFound().

"Load more" and the map fetch resolved against whatever state existed
when they returned, so results from an abandoned search appended
themselves to the new ones. Each fetch now carries an AbortController
and checks that its search scope is still current before touching state.
The map only records its cache key on success, so a failed load retries
instead of pinning the stale marker set.

Jest ignored .next/, whose build output otherwise shadowed real suites.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:17:18 +01:00
TudorandClaude Opus 5 9b75f54206 fix(api): publish a validated dataset, and stop truncating search
Reload cleared the caches first and rebuilt afterwards, so any failure
left the API serving nothing, and requests arriving mid-reload saw a
half-swapped state. It now builds and validates the replacement frames,
place registry, reverse index and sitemaps off the request loop, then
publishes them in one synchronous step under a lock. A failed reload
returns 503 and keeps the previous data. Sitemap regeneration takes the
same path rather than clearing the live registry up front.

Typesense search returned at most one page of hits and used an empty
list for both "no matches" and "search is down", so a genuine empty
result silently fell back to substring matching. It now pages through
every candidate and returns None only on failure; the fallback matches
literally, since a query containing regex metacharacters used to throw.

Empty datasets answer 503 rather than 200-with-nothing or a misleading
404, so callers can tell an outage from an absent school. Adds
/api/release, which reports the build identity baked into the image.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:16:58 +01:00
TudorandClaude Opus 5 38bc17cab3 feat(search): validate the index before the alias points at it
The old sync created a collection, imported batches without reading a
single import response, and swapped the alias regardless. A partial
import published a half-empty index, and two overlapping DAG runs could
prune each other's collections.

Publication now checks every import response and the final document
count before upserting the alias, and holds a session-scoped advisory
lock across the read and the publish so concurrent runs serialise.
Cleanup keeps the previous collection as a rollback pointer and is
best-effort: an uncertain alias response must never delete what might
still be live.

Also parses the Typesense URL properly instead of splitting on colons,
which mangled any host carrying a scheme and a default port.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-09-15 10:16:36 +01:00
tudor dc156058fe Merge pull request 'docs: describe the system that exists, and remove the importer's remains' (#147) from docs/current-truth-and-legacy-inventory into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 22s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m22s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m57s
Reviewed-on: #147
2026-09-15 08:58:54 +00:00
TudorandClaude Opus 5 1d8858fbda chore: remove the code the legacy CSV importer left behind
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m2s
`backend/migration.py` and `scripts/migrate_csv_to_db.py` import `School`,
`SchoolResult`, `init_db` and `set_db_schema_version` — names that no longer
exist. `scripts/geocode_schools.py` imports the same removed ORM model. None of
the three can be imported against the current backend, so they were not dormant
utilities anyone could fall back on; they were files that would fail on the
first line. `backend/version.py` existed only to hand `SCHEMA_VERSION` to that
importer, and the FastAPI lifespan performs no version-triggered import.

Three symbols go with them, each confirmed to have no caller: the unvectorised
`haversine_distance`, superseded by the inline NumPy calculation in search;
`fetcher`, an SWR helper for a dependency this project does not install; and
`kmToMiles`. `calculateDistance` stays — CutoffMapPanel uses it.

Two comments pointed at `migrate_csv_to_db.py --drop` to explain why Payload
owns its own schema. The reason survives the script: blog content must stay
clear of the school marts and Airflow's metadata. Reworded rather than deleted,
so the constraint keeps its justification.

docs/LEGACY_CODE.md records what was removed and where to find it in history. It
also records what was deliberately *not* removed, which is the more useful half:
unused UI components awaiting a design decision, manual data utilities whose
operators a repository search cannot see, and fallbacks that look obsolete but
are load-bearing — `data_loader.py`'s older-mart branches, the generated GIAS
dictionary copies, and the `legacy`-named dbt models that annual DAG selectors
explicitly include. A zero-import count is evidence, not a verdict.

The scripts that fetch DfE CSVs are marked historical and kept, pending
confirmation that nobody runs them by hand.

Checked: 190 backend tests, 429 frontend tests, `tsc --noEmit` clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan
2026-09-14 23:01:22 +01:00
TudorandClaude Opus 5 eaf5e5d180 docs: describe the system that exists, not the one we started with
The README still opened on "Primary School Compass", a KS2 tool for Wandsworth
and Merton served by FastAPI and vanilla JavaScript with Chart.js. Every layer
of that sentence is now wrong: coverage is England-wide across KS2, KS4,
all-through and post-16, Next.js owns the public UI, and school data comes from
dbt-built `marts.*` rather than CSVs loaded at startup. The setup instructions
walked a reader into a virtualenv and a CSV import that cannot build the current
schema, so following the docs produced an empty database and a wrong mental
model at the same time.

Replace the narrative docs with two reference documents that were checked
against the code: docs/ARCHITECTURE.md for request flow, data ownership, the
backend/frontend module boundaries and the real publication sequence, and
docs/DEVELOPMENT.md for the checks that actually run, including the container
and CI version skew that makes "just run pytest" misleading.

The env examples drifted the same way. ALLOWED_ORIGINS is a JSON array, not a
comma-separated list; the frontend needs FASTAPI_URL, DATABASE_URL and
PAYLOAD_SECRET, none of which were documented; and RATE_LIMIT_BURST,
DEFAULT_PAGE_SIZE and MAX_PAGE_SIZE were presented as tuning controls the routes
do not consult. Each is now stated as it behaves.

MIGRATION_SUMMARY.md keeps its content but gains a banner, because it reads like
setup instructions and is not.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016y2J6bs8gbuSJbH18w7Tan
2026-09-14 23:01:15 +01:00
tudor be780ebe13 Merge pull request 'fix(seo): declare the share card, which the route group stopped inheriting' (#146) from fix/og-image-route-group into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m21s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m58s
Reviewed-on: #146
2026-09-14 21:03:14 +00:00
TudorandClaude Opus 5 b6c2cd5116 fix(seo): declare the share card, which the route group stopped inheriting
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 53s
The staging E2E gate's one failure. Every link to the site pasted into
a chat has been rendering bare.

Staging serves og:title, og:description, og:url, og:site_name, og:type
and twitter:card, and no og:image at all. So metadata from the layout
reaches the page; only the file convention does not.

app/opengraph-image.tsx does work — _not-found, which lives in the app
root segment, carries an og:image from it in the build output. It does
not reach the site's pages, which live in the (frontend) route group
whose own layout.tsx is the root layout. The icon conventions are not
affected: /icon.png and /apple-icon.png are both linked correctly on
the same page, verified against staging. The asymmetry is the whole
bug, and it arrived with the route-group split that Payload required.

The file stays at the app root. Moving metadata files into a route
group is what drops /robots.txt and hashes /icon.png, which CLAUDE.md
records and which this must not undo — the build still emits all four
of /robots.txt, /icon.png, /apple-icon.png and /opengraph-image. The
root layout points at the route instead, and metadataBase makes it
absolute, which the journey needs since it calls new URL() on the value.

twitter.images is set for the same reason: the card is declared
summary_large_image, and claiming a large-image card while supplying no
image is worse than claiming a summary card.

Checked before fixing that og:image was the only broken assertion in
that journey: the test aborts at line 911, so its apple-touch-icon and
maskable-icon assertions had never run. All four of those assets return
200 image/png from staging, so this does not simply move the failure
further down the test.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 21:54:32 +01:00
tudor dc79d653e5 Merge pull request 'feat(seo): link school pages into the location layer' (#145) from feat/school-page-place-links into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 21s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m26s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m29s
Reviewed-on: #145
2026-09-14 20:29:48 +00:00
TudorandClaude Opus 5 b0d5334e06 perf(places): index the reverse lookup, and isolate the registry in tests
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m17s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m47s
Two review findings, both confirmed before fixing.

The client fixture in test_school_details.py patched load_school_data
but not _place_registry, which is a module-level cache. A probe settled
it rather than an argument: poisoning the global with a registry built
from data the fixture never saw, then issuing the fixture's own
request, returned that other dataset's places. So the new
`places == []` assertion was satisfied by a stale registry exactly as
well as by the fixture's own data, and proved nothing. Every other test
that touches place data already reset it; the fixture predates places
existing and was never updated. It resets it now.

places_for_urn walked every place in the registry and did a tuple
membership test against each, on /api/schools/{urn}, the site's
highest-traffic endpoint. It now reads a dict built once per registry.
Measured against a synthetic corpus of 27k schools in 1,650 places:
0.118ms per request becomes 0.0001ms, with the index built once in
21ms. Production carries ~5,000 places, so the scan there is larger
again. The absolute saving per request is small; the point is that it
is repeated on every school page view and costs nothing to remove.

The index is cached against the registry by identity rather than
behind a second flag. Anything that drops _place_registry — every test
that touches place data does — gets a fresh registry object, which no
longer matches what the index was built from, so the index rebuilds
with it. A separate _place_index = None would be one more thing to
forget, and a stale reverse index is precisely the first finding's bug
wearing a different hat.

That invalidation has its own test, and the test was checked by
breaking the identity check: five tests fail without it, so three
existing ones were already relying on it too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 21:22:39 +01:00
TudorandClaude Opus 5 d65eb58883 fix(seo): build the phase links the docstring already promised
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m18s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 3m48s
Review caught _places_payload documenting a `phase_url` the function
never returned. The docstring was not stray prose: the approved design
included the phase variant — "Primary schools in Beccles" was one of
its four example links — and it was dropped during implementation
without being mentioned. Deleting the sentence would have closed the
report while losing the feature, so the links are built instead.

These are the pages that most needed them. ~950 phase variants were
once reachable by nothing at all: absent from every sitemap and
unlinked from the place page. "Primary schools in brentwood" is the
query they exist to answer.

Membership is read from the registry's own `phase_urns` rather than
re-derived from the school's phase string. The registry already decides
which phases a place publishes and which schools are listed on each, so
asking it is both shorter and the only way the link cannot disagree
with the page it points at. It also means outcodes need no special
case: they carry empty `phase_urns` by design, because nobody searches
"primary schools in SW11", so they report no phase links on their own.

`phases` is a list rather than a single url. An all-through school is
listed on both the primary and secondary pages, so there is no tie to
break and no reason to invent one. Each entry renders directly after
its own place, so "22 primary schools in Brentwood" reads as part of
Brentwood rather than as an unrelated link further along the row.

The e2e journey now follows a phase link where the town publishes one
and asserts it resolves.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 20:54:59 +01:00
TudorandClaude Opus 5 7f5f0fb676 feat(seo): link school pages into the location layer
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 22s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m24s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m18s
W2 shipped ~5,000 place pages and nothing linked into them. The
location layer pointed down at school pages; school pages pointed
nowhere on the site. Their only anchor was the school's own website, so
the ~27k pages carrying most of the site's inbound authority passed it
straight off-site, and the new corpus was reachable mainly through the
sitemap.

Three things close the loop.

A reverse index over the place registry, places_for_urn, answers which
published places contain a school. Derived from the registry rather
than stored beside it, so the two cannot disagree about which places
exist: a place below the publish threshold is absent from the registry
and therefore never offered as a link. A test asserts that invariant
across every place in a built registry.

GET /api/schools/{urn} gains a `places` array carrying the name, count
and canonical path for each. It rides on the request the page already
makes, so the school page costs no extra round trip. The frontend types
it optional and defaults it to empty, because the two images deploy
separately and a frontend ahead of the API must render without it.

The page gains a "More schools near here" module and a BreadcrumbList.
The module orders narrowest first, because a reader on a school page
wants its town before its county, while the API orders widest first for
the trail. Anchors state their destination's size — "37 schools in
Brentwood" — which is worth more to a reader and a crawler than "see
more". With no published places it renders nothing rather than an empty
heading.

The trail is rooted at the homepage, not /schools. There is no /schools
index page; the location layer lives only at /schools/[place],
/schools/authority/[la] and /schools/near/[outcode]. Rooting it at the
bare path would have opened every breadcrumb with a link to a 404.
Outcodes are omitted from the trail: "schools near CM15" is a real
query and a useful link, but nobody navigates Essex to CM15 to a
school, and a breadcrumb claiming that describes a hierarchy the site
does not have.

School pages also now declare the School type rather than
EducationalOrganization, the parent type that covers universities and
nurseries alike.

The e2e journey asserts the round trip in both directions, following a
place page's own first school so the pair is genuinely related rather
than hardcoded. A one-way link is what already existed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-14 20:45:50 +01:00
tudor 47f3591ed8 Merge pull request 'feat(flags): put /about and /blog behind flags, dark by default' (#144) from feat/about-blog-flags into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m22s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m10s
Reviewed-on: #144
2026-09-08 16:47:18 +00:00
TudorandClaude Opus 5 6fc7fce948 feat(flags): put /about and /blog behind flags, dark by default
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 20s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m21s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 12s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m39s
Both features ship dark. Neither is reachable in an environment where
its flag is off, and every flag in this system starts off, so a deploy
of this commit makes both disappear until someone turns them on
deliberately.

Two independent flags rather than one, which makes blog-on-about-off a
reachable state. That state is the whole reason the change is larger
than four notFound() calls: the blog leans on the About page for its
author identity. The Person entity is anchored at /about#tudor, and
that URL 404s while about_page is dark, so a post published in that
state would claim an author resolving to nothing. Worse than having no
named author. Both bylines fall back to unlinked text and the
BlogPosting attributes to the publisher instead, so every combination
of the two flags renders something correct.

Gated: /about, /blog, /blog/[slug], the RSS feed, both footer links,
and the matching content-sitemap entries. A sitemap must never
advertise a URL that 404s. With both dark it emits a valid empty
urlset rather than a 404, because robots.txt names it unconditionally.

Not gated: /admin. Posts have to be writable before the blog is worth
switching on, so flagging the panel would make the flag unflippable.

getFlags takes a revalidate rather than always using the 300s
constant. Reading a flag pins the calling route to the lowest
revalidate among its fetches, and the footer links live in the root
layout, so a naive gate there would have dropped every school and
place page from a weekly cache to a 5-minute one. The layout passes
604800, the floor those routes already declare, and the build confirms
all four SSG route families still prerender. The cost is one-way
latency: pages follow a flip in minutes, footer links within a week.

The e2e journeys follow the existing paired shape from the
admission_distance flag: a lit journey and a dark one for each flag,
reading state from whether /about and /blog respond rather than from
/api/flags, which another journey asserts is not publicly reachable.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-08 17:12:25 +01:00
tudor eb6d918650 Merge pull request 'fix(cms): regenerate the import map so the Content field renders' (#143) from fix/payload-import-map into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 15s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m9s
Reviewed-on: #143
2026-09-02 20:12:17 +00:00
TudorandClaude Opus 5 d47ac71c47 fix(cms): regenerate the import map so the Content field renders
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m11s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m10s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 52s
Creating a post in the admin panel showed no Content editor, and saving
failed validation on the field the writer was never shown.

The admin panel does not import field components. The server hands the
client a path per field, and resolves it through the generated map at
app/(payload)/admin/importMap.js. A richText field's path is
@payloadcms/richtext-lexical/rsc#RscEntryLexicalField. The committed map
held one entry, @payloadcms/next/rsc#CollectionCards, generated before
the blog collections existed and never re-run. A path missing from the
map is not an error the panel reports: the field simply does not render,
while required is still enforced server-side on save.

next build does not regenerate the map, so the stale copy shipped in the
image and the editor was equally broken on staging and production.

Regenerated with payload generate:importmap, which adds the lexical RSC
field, cell and diff components, BlocksFeatureClient for the Callout
block, and the default toolbar features.

Two things stop it drifting again. There was no script to run, so
package.json gets generate:importmap. And a test asserts the map carries
an entry for each thing the config asks for, in the source-reading style
of the other payload suites; against the old map all five fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01DXnXQKnPpZBBP61fBQiFkq
2026-09-02 20:44:05 +01:00
tudor 17e5371e9c Merge pull request 'fix(cms): ship the initial Payload migration (recovers two commits stranded after #140 merged)' (#141) from fix/payload-initial-migration into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m13s
Reviewed-on: #141
2026-09-02 17:52:17 +00:00
TudorandClaude Opus 5 e2c63a9905 fix(cms): ship the initial migration so a container finds its tables
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m14s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m14s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m33s
Staging failed on boot with 42P01, relation "payload.users" does not
exist. The schema was empty because no migration existed, and the
adapter cannot create tables itself: db-postgres/connect.js gates push
on NODE_ENV !== 'production', so it is inert in a deployed container
regardless of config.

The generated migration is schema-qualified to "payload" throughout but
does not create that schema — schemaName says where tables go, it does
not create anything. It only worked against the throwaway database used
to generate it because the schema was created there by hand, so every
real environment would have failed on the first statement. CREATE SCHEMA
IF NOT EXISTS is hand-added at the top of up(), which makes it exactly
the kind of edit a regeneration discards silently; a test asserts it is
present and ordered before the first CREATE TABLE.

payload-types.ts is now committed rather than ignored. Ignoring it meant
CI typechecked against looser types than a developer with a generated
copy, which is how a Record<string, unknown> cast passed CI and then
failed locally the moment the file appeared. The post page uses the
generated Post and Media types instead, and narrows heroImage rather
than asserting it, since the field is an id at shallow depth.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00
TudorandClaude Opus 5 3f3c5953f6 style(copy): remove em dashes from the site's prose
The em dash is one of the clearest tells of machine-written text, which
is the exact impression this work exists to remove. Rewritten rather
than substituted: where a dash was carrying a real aside the sentence is
split or recast, not patched with a comma.

Covers the About page, the two Callout labels an editor sees in the
admin panel, and PUBLISHING.md, which defines the house style and should
follow it. The rule is now recorded in that house style and in the
spec's voice rules, so it survives this branch.

Code comments are left alone: they are not copy, and the surrounding
codebase uses the same punctuation throughout.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 18:33:27 +01:00
tudor 124c6702a9 Merge pull request 'feat: a named author, an About page and a Payload blog' (#140) from feat/about-and-blog into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 45s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 1m32s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m8s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m36s
Reviewed-on: #140
2026-09-02 16:04:58 +00:00
TudorandClaude Opus 5 e25722d9ab fix(blog): hide drafts at the access layer, and back the --drop claim
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m12s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 32s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m9s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m26s
Review findings on #140.

Drafts were reachable. Posts granted unconditional public read and the
_status filter lived only in the pages that query the collection — which
is a convenience, not a control. Payload's documentation is explicit:
"The `draft` argument alone does not restrict documents with _status:
'draft' from being returned by the API." A direct GET /cms-api/posts
would have handed every unpublished draft to any visitor. Read access
now returns a query constraint for anonymous callers, which is the
documented mechanism.

The --drop claim was asserted across four files while the spec still
listed it as an open question. Now verified rather than assumed:
run_full_migration drops exactly ["school_results", "schools"] by name,
there is no drop_all() or DROP SCHEMA anywhere in backend/, the only
other drop is schema-qualified to marts, and nothing sets search_path.
The guarantee is stronger than schema isolation alone — those two table
names do not exist in Payload — so the claim stands, but it now rests on
cited code. The spec records the evidence and closes the open item.

findPost is wrapped in React's cache(): Next calls generateMetadata and
the page separately for one request, so every post view ran the same
query against Postgres twice.

The bare .lede rule was dead — .prose p scores (0,1,1) and outranks it —
so only .prose .lede ever applied. Removed, with the specificity noted
so the surviving selector is not "simplified" back into a silent
regression.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:48:59 +01:00
TudorandClaude Opus 5 07d586d0ad docs(blog): how to publish, and why the app has two route groups
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m15s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 36s
PR Checks / Build Frontend (no push) (pull_request) Successful in 1m11s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m18s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m43s
PUBLISHING.md carries the house style with the posts, so the standard
survives without the design doc to hand — including the rule that a post
states what a metric does not show, which is the strongest signal a
human wrote it.

CLAUDE.md gains the two constraints that are invisible from the code and
expensive to rediscover: metadata file conventions break if moved into a
route group, and the build must keep succeeding with DATABASE_URL unset.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:30:09 +01:00
TudorandClaude Opus 5 b793640507 feat(blog): add the blog index, post pages, RSS and content sitemap
The rendering split is dictated by CI building with no database.

/blog, /blog/rss.xml and /content-sitemap.xml have no dynamic params, so
Next prerenders them at build time and the build fails on a missing
Payload secret — caught here, not on staging. They are force-dynamic
instead: one indexed query against Postgres on the same Docker network,
and a newly published post appears immediately rather than waiting on a
revalidation. /blog/[slug] keeps ISR, because with no
generateStaticParams there is nothing to prerender; it is generated on
first request and cached, which is exactly what the collection's
afterChange hook exists to invalidate.

RichText takes `converters`, not `blocks`, in Payload 3.88, and the
default converters must be spread or every paragraph and heading loses
its renderer and the body comes out empty.

BlogPosting references the Person and Organization by @id rather than
repeating them, so every post and the About page resolve to one author
entity instead of declaring several people with the same name.

/sitemap.xml is proxied from FastAPI, which knows nothing about Payload,
so the Next-owned URLs get their own sitemap and robots.txt lists both.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:28:58 +01:00
TudorandClaude Opus 5 21a5d18f59 feat(blog): add the posts and media collections
Drafts are on so a post can be written across sittings without saving
being publishing.

afterChange and afterDelete revalidate every path a post appears on.
Blog pages are ISR because CI builds with no database, so without these
a published post would not appear until the revalidate window expired —
up to an hour of a writer concluding that publishing is broken. Payload
runs in the same process as Next, so these are direct revalidatePath
calls with no webhook and no shared secret.

Media writes to an absolute /app/media matching the compose mount; a
mismatch would write into the container filesystem, where the next
redeploy silently discards it. Alt text is required rather than
optional.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:24:15 +01:00
TudorandClaude Opus 5 f614414070 feat(about): give the site a named author
The site had no author, no statement of why it exists and nobody
accountable for its numbers, which is most of why it reads as machine
generated.

The page states plainly that its author is not an education expert. The
credibility claim is lived experience — a parent going through primary
admissions — plus stated provenance for every figure, which is true and
cannot be undermined by someone noticing there is no teaching
qualification behind it. First name only: the Person JSON-LD carries no
familyName, worksFor or affiliation, and a test asserts it stays that
way.

The footer gains a fourth column, with a tablet breakpoint so four
columns pair up rather than crushing before the 768px collapse. The nav
is deliberately untouched — its mobile tab bar already carries four
items.

public/brand/tudor.jpg is NOT in this commit. The page references it and
will show a broken image until the photograph is supplied; a stock
portrait would defeat the entire point of the work.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:22:52 +01:00
TudorandClaude Opus 5 310b63b0cb build(cms): wire Payload into the Docker image and both stacks
Uploads go to a named volume at /app/media. The directory is created in
the image before the mount and covered by the existing chown, because
Docker seeds a fresh named volume from the image path — a missing or
root-owned directory there fails every upload with EACCES at runtime,
long after the build passed.

PAYLOAD_SECRET uses the same :? form as AIRFLOW_ADMIN_PASSWORD: refuse
to start rather than boot with an empty secret and accept forged
sessions. Staging's must differ from production's, which the header
comment now says explicitly. Portainer prefixes volume names per stack,
so payload_media isolates itself.

prodMigrations is not wired yet — generating the initial migration needs
a reachable Postgres. Follows in its own commit.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:19:27 +01:00
TudorandClaude Opus 5 c2c76c5817 feat(cms): keep the admin panel out of the index
X-Robots-Tag rather than the robots.txt Disallow alone, for the same
reason the staging rule uses one: a Disallow blocks crawling, not
indexing, so a URL found from an external link can be indexed without
ever being fetched — and blocking the crawl means the noindex is never
seen. Both mechanisms are applied to /admin and /cms-api.

The existing CSP is frame-ancestors only, which restricts who may embed
the site rather than what a page may load, so it cannot break the panel.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:17:55 +01:00
TudorandClaude Opus 5 c5a4d106da feat(cms): install Payload and serve the admin panel
Payload 3.88 runs inside the Next app against the existing Postgres, in
its own 'payload' schema so no pipeline operation on public — the app
tables, Airflow's metadata, migrate_csv_to_db.py --drop — can reach blog
content.

Its REST API is mounted at /cms-api. /api is the FastAPI proxy's
catch-all, which would swallow every admin call and forward it to the
backend with no error. The mount points live in lib/payloadRoutes.ts so
there is one definition and a test can assert it without importing
Payload: it is ESM-only, next/jest will not transform it, and appending
transformIgnorePatterns cannot un-ignore a package. Forcing it through
transpilePackages would change how the production build bundles Payload
to serve a test, so the live proof that /api still reaches FastAPI stays
where it belongs — the e2e journeys, which call /api/schools.

The package becomes ESM ("type": "module"), which Payload's CLI requires:
richtext-lexical has top-level await and the config cannot be require()d.
Only two files needed renaming, jest.config.cjs and a build script.

The build is verified to succeed with DATABASE_URL and PAYLOAD_SECRET
both unset, which is how CI builds it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:17:01 +01:00
TudorandClaude Opus 5 2437ffce42 refactor(app): move site routes into a (frontend) route group
Payload's admin panel ships its own root layout rendering html/body.
Next allows multiple root layouts only when no app/layout.tsx exists, so
the site's routes move into their own group. Route groups are invisible
to routing: every public URL is unchanged, verified against the build's
route table.

The metadata file conventions deliberately stay at the app/ root. Moving
them into the group renamed /icon.png to /icon-4usi79.png (likewise
apple-icon and opengraph-image) and dropped /robots.txt altogether,
which would have broken the /icon.png cache-control rule, the
outputFileTracingIncludes entry for the share card, and robots.txt.

darkThemeSafety reads app/globals.css off disk rather than importing it,
so it needed its own path fix — a grep for import specifiers misses it,
and it fails as an unrunnable suite rather than a failed assertion.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:11:48 +01:00
TudorandClaude Opus 5 eb648f3f76 build(next): convert the config to ESM so Payload can wrap it
withPayload() is ESM-only, so next.config.js has to become .mjs. That
file also carries the rule that keeps staging out of Google's index, so
the conversion goes in on its own, behind a test that asserts the rule
survived — along with the standalone output, the opengraph-image font
tracing and the analytics frame-ancestors CSP.

Jest resolves the .mjs config without extra configuration, so
jest.config.js is untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 16:08:58 +01:00
TudorandClaude Opus 5 74e5fffc10 docs(about-blog): implementation plan for the About page and Payload blog
Nine tasks, each ending in an independently testable deliverable.

Two structural findings that the spec did not anticipate, both recorded
in the plan. Payload's admin panel ships its own root layout rendering
html/body, and Next allows multiple root layouts only when no
app/layout.tsx exists — so every existing route moves into an
app/(frontend) route group first, on its own, with the full suite as the
gate. Route groups are invisible to routing, so no public URL changes.

The second finding corrects the spec: adding /cms-api to the FastAPI
proxy's exclusion list would be dead code, because that catch-all only
ever matches /api/*. The route remap alone is sufficient.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_017YmbBhr8s7GusjDE12hrZM
2026-09-02 15:59:44 +01:00
TudorandClaude Opus 5 748ef32180 docs(about-blog): design for a named author, an About page and a Payload blog
The site reads as synthetic because nobody is accountable for the
numbers, no editorial judgement is visible, and the voice is
institutional third person. This designs the fix: a named author
(first name, photo, explicitly not an education expert), a coded
/about page, and a blog backed by Payload CMS running inside the
existing Next app.

Also fills a hole in the SEO programme, which has eight workstreams
and no E-E-A-T or authorship signal on a YMYL corpus.

Records two collisions found while designing, both of which fail
badly if missed: Payload's default /api route fights the existing
FastAPI catch-all proxy, and withPayload() is ESM-only so
next.config.js — which carries staging's noindex header — has to
become next.config.mjs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FT1Ls4GbgLDXoQX7NAuHGT
2026-09-02 14:43:20 +01:00
tudor b0c4ea8282 Merge pull request 'fix(destinations): the DAG died on a null in a primary key' (#139) from fix/destinations-national-grain into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 14s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m36s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m6s
Reviewed-on: #139
2026-08-31 20:43:26 +00:00
TudorandClaude Opus 5 e236669fde fix(destinations): school rows and the England reference are different grains
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 52s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m27s
The annual DAG died with a BrokenPipeError from Meltano's log writer, which
is several frames from the cause: target-postgres exited first and the tap
saw its stdout close.

The tap declared primary_keys = [urn, ...] while emitting urn=None for the
national rows, and target-postgres turns primary_keys into a NOT NULL
constraint. The first national row of the run failed the insert and took
the loader with it. Every other tap in this repo keys on non-null columns.

Carrying two grains in one stream was the actual mistake, so the fix is to
separate them rather than paper over the null: four streams now, with
ees_ks4/ks5_destinations_national carrying no urn column at all — a school
identifier that is null in every row is a grain mismatch, not a column.
The staging models split the same way and the national mart reads the new
pair instead of filtering `where urn is null`.

Verified against the live API: the school stream yields 135,240 rows over
4,508 schools with no duplicate keys, no null key columns and all 31,382
suppression sentinels intact; the national streams yield 30 and 33 rows
with no urn column.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 21:36:43 +01:00
tudor fb5a0928bd Merge pull request 'fix(airflow): a fixed admin password from the environment, and a tap that missed #137' (#138) from fix/airflow-fixed-admin-password into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 55s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m36s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m14s
Reviewed-on: #138
2026-08-31 19:50:00 +00:00
TudorandClaude Opus 5 264edd2e3a fix(airflow): a login that survives a container restart
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m7s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 12s
PR Checks / Build Frontend (no push) (pull_request) Successful in 50s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 51s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m58s
The simple auth manager generates a random password on first start and
writes it to a file, so every restart of the api-server invalidated the
last one and the password had to be dug out of the container logs again.

The stack now writes that file itself from AIRFLOW_ADMIN_PASSWORD before
exec'ing the api-server. Airflow generates nothing when the file already
exists, so the login is whatever the stack environment says it is.

Written with python rather than echo, so json.dumps escapes a password
containing quotes, backslashes or non-ASCII correctly — verified against
`p@ss "wo\rd' £5`, which round-trips intact.

An unset AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits.
Falling back to a generated password would silently undo the point of the
change, and a compose-level `:?` gives the same refusal a readable reason.
This does mean the variable MUST be set in Portainer before the next
deploy of either stack.

Not affected by the two Docker gotchas in the upstream docs: this image
has no USER directive so it runs as root, and the file is rewritten from
the environment on every start rather than persisted on a volume.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 17:38:42 +01:00
TudorandClaude Opus 5 cd2cbe7be6 build(pipeline): install the destinations tap in the image
meltano install would resolve it from pip_url, but five of the six custom
taps are also installed explicitly and a new plugin failing to appear is
not something you want to debug from a deploy log.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-31 17:37:46 +01:00
tudor 73182d0c0c Merge pull request 'feat(destinations): say what happened to a school's leavers, without republishing what DfE withheld' (#137) from feat/ks4-destinations into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m6s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 2m2s
Reviewed-on: #137
2026-08-30 20:49:14 +00:00
TudorandClaude Opus 5 cbe3a9a772 fix(destinations): the table said 'withheld' for a category that just doesn't apply
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m14s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
The Share column keyed off `percentage === null`, which is true for
not_applicable as well as suppressed, so a destination that does not apply
to the school was labelled as one DfE withheld — while the Pupils column
in the same row rendered blank. Two columns, one row, disagreeing about
what the row was, and one of them making a claim about DfE that wasn't
true.

Both columns now derive from `status`, which is the distinction the mart,
the SQLAlchemy model and the serialiser all preserve deliberately:
published shows the figure, suppressed shows the withheld badge,
not_applicable shows an em-dash with a title saying so.

A published count with no published percentage now derives its share from
the cohort rather than falling through to a marker — both halves are
published, so nothing withheld is involved, and it is the same derivation
the bar widths already use.

Verified the new tests fail against the old logic before keeping them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-30 21:48:39 +01:00
TudorandClaude Opus 5 2e9b5c83c5 fix(destinations): the masking pass can no longer exit unsafely
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 7m23s
Review found _mask_for_disclosure could return with its invariant broken
and say nothing. add_companion only ever withheld a *published* cell, so a
group with one suppressed category and every other one not_applicable —
routine in special schools and AP, where few categories apply — left the
loop with the lone suppressed cell still solvable. Reproduced on a
nine-pupil cohort: one hidden cell, cohort served, residual intact.

A disclosure-control pass that fails silently is worse than none, because
everything downstream trusts it. The loop now runs until the invariant
holds and escalates when no companion exists: the pupil group is dropped
from the payload, and an empty block serialises as None so the section is
absent rather than an empty shell. disclosure_invariant_holds() is exported
so tests assert it directly instead of re-deriving it, and an exhaustive
test sweeps all 81 suppression patterns of a four-category group.

Also fixes a test that set up six measures and checked one: the loop was
`for measure in ["school_sixth_form"]`. It now checks every measure, and
against the real invariant — none hidden, or at least two, rather than
"at least two", which the five published measures would have failed.

No regression on real data: 262 mainstream secondaries, all-pupils bar
still drawable on 94%, zero invariant violations, one disadvantaged group
dropped by the new escalation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-30 21:35:03 +01:00
TudorandClaude Opus 5 102397fe69 fix(destinations): withhold at the API, not just in the chart
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m14s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 3m49s
Code review found the disclosure the whole design was meant to prevent.
R1 was written as a rendering rule and implemented as one: canRenderBar
stopped the bar being drawn, but GET /api/schools/{urn} still carried the
cohort and every published category. cohort - sum(published) returned
Whitley Bay's withheld further-education figure exactly — 18 pupils — to
any caller, and the RSC payload put it in the browser too.

app.py already stated the principle for admission_distance: this endpoint
is public and unauthenticated, so a field left in the payload is a
published field. The same reasoning applies here and did not get applied.

_mask_for_disclosure now closes both identities before serialisation —
categories sum to the cohort, and disadvantaged + other = all — by adding
secondary suppression until every row and column hides none or at least
two. My first attempt picked the smallest published cell as the companion
and a new test caught it choosing a zero, which protects nothing: the
residual still resolved to 18. The companion must carry pupils.

DfE's own aggregates are no longer served. Nothing rendered them, and one
spanning a single suppressed component names it.

Cost, measured over 262 mainstream secondaries: the all-pupils bar
survives on 94% rather than 100%. Zero lone-suppressed groups remain.

The e2e helper now tells a missing feature apart from missing data: it
fails if the API serves no destinations key at all, and skips if the key
is served but the annual DAG has not populated the marts. Failing on the
second would redden the staging gate for unrelated commits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 19:58:02 +01:00
Tudor 68a192e430 Merge remote-tracking branch 'origin/main' into feat/ks4-destinations
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m12s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 4m4s
# Conflicts:
#	nextjs-app/__tests__/components/darkThemeSafety.test.ts
2026-08-28 18:45:27 +01:00
TudorandClaude Opus 5 ccd5074c90 test(e2e): destination journeys, including the no-bar rule
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 19s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m16s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m10s
The helper throws rather than skipping when no school returns a
destinations block: a silent skip would let a real regression in the
sections ride along unnoticed, which is why the distance journeys were
changed the same way in 4f01fbd.

The disadvantaged journey computes the residual itself and asserts it
appears nowhere on the page — the one number the section must never state.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:16:25 +01:00
TudorandClaude Opus 5 2b4cf20d75 feat(destinations): the post-16 section, replacing the placeholder
The 'Post-16 destination data coming soon' note is deleted rather than
reworded: for a school with no sixth form the truthful statement is that
the question does not apply, and a placeholder there implies something is
missing. hasSixthForm and .sixthFormNote go with it — nothing else used them.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:15:21 +01:00
TudorandClaude Opus 5 cef2f77149 feat(destinations): the After Year 11 section
Question cards over one bar, with the cards acting as a lens on the bar
rather than a summary beside it — focusing a card dims everything it is
not made of, so the grouping we chose is inspectable rather than asserted.

The bar renders only when canRenderBar allows it. Where a category is
withheld the section says so and shows the table instead: the categories
sum to the cohort, so a bar drawn from the published segments leaves a
gap whose width is the withheld figure.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:14:13 +01:00
TudorandClaude Opus 5 68b6417149 feat(destinations): types and secondary section flags
Destinations are secondary-only, so the flags go on computeSecondaryFlags
rather than computeSchoolFlags. A phase counts as present only when some
pupil group carries categories — an empty block would otherwise open a nav
entry pointing at a section that never renders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:11:37 +01:00
TudorandClaude Opus 5 c5719ef362 feat(destinations): serve destinations without closing the gaps
The serialiser carries status through and computes no totals of its own.
The only aggregates in the payload are ones DfE published itself; whether
showing one is safe depends on how many of its components are suppressed,
which the frontend decides.

The batch guard grows from six tables to eight.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:10:29 +01:00
TudorandClaude Opus 5 5e5b61987a feat(destinations): marts, with R3 masking applied at the boundary
Disadvantaged and other-pupils partition the whole and the all-pupils
figure is published, so publishing both halves recovers the suppressed
one. The mask is applied in the mart rather than the API so no consumer
added later can reach an unmasked combination.

The R1 test is a warn, not an error: DfE publishes the recoverable
combination and the mart's job is to carry it faithfully. Refusing to
close the gap is the API's job and the frontend's.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:08:20 +01:00
TudorandClaude Opus 5 c564566432 feat(destinations): staging models that keep 'withheld' distinct from 'absent'
safe_numeric maps every EES sentinel to NULL, which is right for attainment
and wrong here: one of those states has to print 'withheld' and the other
has to print nothing. A status column carries the difference.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:07:17 +01:00
TudorandClaude Opus 5 9188626051 feat(destinations): a tap that preserves the suppression sentinel
EES writes 'c' where a figure is withheld and the categories sum to the
cohort, so counts and percentages are emitted as text with the sentinel
intact. safe_numeric must never be pointed at them.

School rows and the England reference need different establishment pins:
at national level selective schools, studios and UTCs are separate
populations rather than labels, so leaving establishment open multiplies
30 rows into 190. Two queries per period, each keeping its own level.

Verified against the live API for 2022/23: 135,240 school records over
4,508 schools, exactly 30 each, no duplicate keys, 31,382 sentinels kept.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 16:06:37 +01:00
TudorandClaude Opus 5 7ae9ecdc36 feat(destinations): colour tokens, with the absence hatched not coloured
Activity not captured includes independent schools and moving abroad, so a
red segment would be a factual error. The hatch doubles as the secondary
encoding that rescues the neutral/blue pair, which separates at only dE 7.6
as flat fills.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:58:25 +01:00
TudorandClaude Opus 5 1980d79eee feat(destinations): the disclosure rules, as executable guards
The destination categories sum to the cohort and DfE publishes the cohort
total, so a lone suppressed cell is recoverable by subtraction. canAggregate,
canRenderPublishedAggregate and canRenderBar are what stop a consumer doing
that; toBarSegments throws rather than leaving a readable gap.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:57:46 +01:00
TudorandClaude Opus 5 9423f11567 docs(destinations): implementation plan, ten tasks
Ordered so the disclosure guards land first and everything downstream
consumes them: lib/destinations.ts, tokens, tap, staging, marts, API,
then the two sections and the journeys.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:56:36 +01:00
TudorandClaude Opus 5 576013d627 docs(destinations): design for KS4 and post-16 destination measures
The published files suppress individual cells, not whole cohorts, and the
categories sum to the cohort — so on 22% of mainstream secondaries the
withheld figure can be recovered by subtraction. Three disclosure rules
fall out of that, and the rest of the design is downstream of them.

Verified against the EES API rather than assumed: both datasets carry
school-level rows keyed by URN, with the disadvantage split and every
destination category the display needs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BvdDKvFFSZuMVDH5fEyTob
2026-08-28 15:28:09 +01:00
tudor 7c08138fe4 Merge pull request 'fix(map): the popup never took the dark theme' (#136) from fix/dark-mode-map-popup-contrast into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m41s
Reviewed-on: #136
2026-08-27 21:57:53 +00:00
TudorandClaude Opus 5 a7829d591a fix(map): the popup never took the dark theme
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 15s
leaflet.css paints `background: white; color: #333` on the popup card and its
tip. LeafletMapInner binds themed content into it — the school name and the
headline figure are var(--text-primary) — so in dark mode #E9EEF0 landed on
#FFFFFF at 1.17:1. The two things the popup exists to say were the two least
readable things on the page.

Every other foreground in that popup failed too, from the same cause: the
muted phase line at 2.90:1, the vs-national delta at 1.94:1, the Ofsted badge
at 1.74:1. Moving the surface onto --bg-card fixes all of them at once —
13.52, 5.45, 8.14 and 9.11:1 respectively. In light mode --bg-card is #FFFFFF,
so the popup renders exactly as it did.

globals.css already pulls the rest of Leaflet's chrome onto the tokens, and
says why: "this matters most in dark mode, where Leaflet's white attribution
bar would otherwise sit on a near-black page." The popup was simply missed.

The View Details button needed its own fix. It pairs background:var(--status-
above) with a literal white label, which theming the card does not reach:
--status-above is #36743F in light but #7FCB8A in dark, taking the label from
5.63:1 to 1.94:1. --text-inverse is the token for ink on a saturated fill, and
the popup's own Ofsted badge already uses it.

darkThemeSafety already guards this defect class, but only inside .module.css.
Neither half of this one lives there — the surface is a third party's, the
text is inline in a TSX template — so it scanned clean throughout. Two rules
added for the layer it could not see. Fixing the grouped-selector blind spot
in its rules() helper was needed to write them: taking only a selector's last
line discarded every selector in a grouped rule but the final one, which makes
a safety guard fail open.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 22:34:18 +01:00
tudor 1ed4470fc2 Merge pull request 'fix(admissions): flag-off pages must not speak for the council' (#135) from fix/distance-flag-off-absence-copy into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 52s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m43s
Reviewed-on: #135
2026-08-27 20:30:25 +00:00
TudorandClaude Opus 5 7a16b1b52f fix(admissions): flag-off pages must not speak for the council
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m8s
The secondary admissions section words the absence of a cut-off distance:
"<LA> has not published a cut-off distance for this school." That sentence
is true when the authority publishes nothing. It is false when the
authority does publish and the admission_distance flag is simply off — and
off is the current state, so every secondary page with an EES admissions
row has been making a claim about a council on our behalf.

The backend already draws the distinction the copy needs. /api/schools/{urn}
omits the admission_distance key entirely while the flag is dark rather than
sending null, precisely so that "we are not publishing cut-offs" stays
distinguishable from "this school has no cut-off"; lib/types.ts says so in
as many words. The page then collapsed the two with `?? null` before the
section ever saw them.

So stop collapsing it: thread the raw field to SecondarySchoolSections and
word the absence only when the feature is on. Null still gets the sentence
naming the authority — that case is unchanged and still tested.

Primary pages are unaffected: AdmissionsSection carries no absence copy and
renders nothing when there is no figure. DistanceSection already treated
absent and null alike; only its type widens.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 21:17:16 +01:00
tudor cf9d41b476 Merge pull request 'fix(analytics): the funnel source read a referrer that never changes' (#134) from fix/navigation-source-soft-nav into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 52s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m40s
Reviewed-on: #134
2026-08-27 08:26:45 +00:00
TudorandClaude Opus 5 e820e7fecd fix(analytics): the funnel source read a referrer that never changes
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m5s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 47s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m13s
The staging E2E gate has been red since #132 merged (run 1064, and
1066 after it): "a school reached from a location page is attributed
to it, not to direct" expects `place`, receives `direct`.

#132 fixed a real bug — `/schools/` had no case and fell through to
`direct` — but the mechanism underneath it never worked.
getNavigationSource read document.referrer, which the browser writes
only when a *document* loads. Every internal navigation here is an App
Router soft navigation: history.pushState, no new document, so
document.referrer goes on naming whatever opened the tab for the whole
session.

Verified on staging: load /schools/brentwood, click a school, the URL
becomes /school/… and document.referrer is still "".

So `from` reported `direct` for essentially every in-app journey, not
just the ones through the location layer — search, rankings, compare
and detail were all being counted as "typed the URL". The unit suite
passed throughout because every case set document.referrer directly,
which only happens on a full page load.

The fix is a module-level trail written by RouteTrail, a render-nothing
client component in the root layout. Its lifetime is exactly right: it
survives soft navigation, and it dies on a real document load — which
is precisely when document.referrer becomes meaningful again, so the
two cover each other with no overlap.

Reading it skips entries equal to the current path rather than taking
the second-to-last. That makes the answer independent of whether the
layout effect or the page effect ran first — React orders those by
tree position, which is not a contract worth resting a measurement on
— and it gives the right answer both when the user returns to a page
they came from and on a hard load of a school page.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 09:22:37 +01:00
tudor 4fdeb70a93 Merge pull request 'feat(places): say what each school is, not only how it scored' (#133) from feat/place-school-attributes into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m42s
Reviewed-on: #133
2026-08-27 07:59:10 +00:00
TudorandClaude Opus 5 9a1f56c431 feat(places): say what each school is, not only how it scored
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m44s
The location tables carried one column: a percentage. A parent
shortlisting from a town page is asking a different question first —
does it take my child's age, is it a faith school, does it have a
nursery — and the page could not answer any of it.

Primary tables gain Ages, Religious character, Nursery and
Constituency; secondary tables the same minus Nursery, which is a
question about a different intake. An all-through school renders in
both groups, so its nursery shows under primary alone.

The measure moves to the second column rather than the last. Six
columns overflow a phone and .tableWrap turns that into a horizontal
swipe; with the measure last, the one number the page exists for is
the one scrolled off the screen.

Cell rules are the ones the school page already uses, so the two
surfaces cannot disagree about the same school: "Does not apply",
"None" and "Not applicable" all read as no religious character, and
the en-dash age normalisation moves into formatAgeSpan, which
formatAgeRange now delegates to.

Backend: nursery_provision and parliamentary_constituency were not in
the place response. Both are optional GIAS mart columns that
data_loader degrades to NULL, and the `in rows.columns` guard keeps a
mart the pipeline has not rebuilt working.

Also fixes a live bug on the same line: SCHOOL_COLUMNS already ends
with latitude and longitude, and the endpoint concatenated them again,
so pandas dropped one of every duplicated pair and warned "columns are
not unique" on each request. Ordered de-duplication removes the
warning and the silent drop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FuPUioHpxtaiDNagQvjxyM
2026-08-27 08:49:14 +01:00
tudor ade9dbb3ba Merge pull request 'feat(analytics): measure the location layer, and stop calling it direct' (#132) from feat/place-analytics into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m44s
2026-08-27 07:29:48 +00:00
TudorandClaude Opus 5 d1a8596208 feat(analytics): measure the location layer, and stop calling it direct
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 40s
The location pages were only half-tracked. Umami counts a pageview for
each of the ~3,900 URLs automatically, but nothing else: components/
places contained no track() call, and place_viewed was not even a
declared event name.

The part that mattered was worse than a gap. getNavigationSource mapped
a same-origin referrer to a funnel source and had no case for /schools/,
so every school view arriving through the location layer fell through to
'direct' — the bucket you read as "typed the URL, no referrer". W2's
whole purpose is funnelling search traffic onto school pages, so the one
measurement that says whether it worked was reporting the wrong answer,
and reporting it confidently. Verified live against staging: expected
"place", received "direct".

/schools/ is checked before /school/. They differ by one letter and mean
different things — the location layer versus a single school — and a
prefix test in the wrong order silently merges them.

place_viewed carries kind, slug, phase and school_count. kind is the
reason it exists: whether to keep investing in these pages turns on
which sort earns engagement, and a pageview cannot say, because all four
families share the /schools/ prefix and only the registry knows which is
which. It is a client component because PlaceView is a server component;
one line in PlaceView covers all four families, since they all render
through it.

Both E2E journeys were verified failing against staging first — one
because place_viewed does not exist there, the other on the exact
"place" vs "direct" mismatch.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-27 08:23:21 +01:00
tudor a3c09d9b67 Merge pull request 'fix(suggest): the dropdown reopened on top of the search results' (#131) from fix/suggest-reopens-over-results into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m37s
Reviewed-on: #131
2026-08-26 21:07:55 +00:00
tudor a7f4c86464 Merge pull request 'fix(search): the mobile hero search was indented by a card's padding' (#130) from fix/mobile-hero-search into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m50s
Reviewed-on: #130
2026-08-26 20:51:20 +00:00
TudorandClaude Opus 5 0804566736 fix(test): drop a committed scratch probe, and close a hole in the guard
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 9s
Code review, all three findings valid.

e2e/tests/__m.spec.ts was a throwaway probe used to measure the mobile
hero geometry. It asserts nothing, so it could never fail; it carried a
leftover `pick('form').constructor === Object ? null : null` that is
null either way and throws if no form matches; and it should never have
been committed. Deleted.

It survived because `rm -f e2e/tests/__m.spec.ts` ran with the shell
already inside e2e/, so the path resolved to e2e/e2e/tests/... — which
does not exist, and rm -f is silent about that. `git add -A` then swept
it in. I checked `git diff --stat` before committing, which lists only
tracked modifications and never shows an untracked file; `git status
--short` would have.

The scoping guard compared the last line of a rule's prelude against the
literal '.filterBar', so a regression written as a selector list —
`.filterBar, .other { padding }`, or the same split across two lines —
would have walked straight past the test meant to catch it. Selectors
are now split on commas and matched individually, and comments are
stripped first so a brace inside one cannot desynchronise the parse.

Verified against all three shapes: bare, inline comma list, and
multi-line comma list. Each is caught; each passes again once reverted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:49:48 +01:00
TudorandClaude Opus 5 55363cbd18 fix(suggest): the dropdown reopened on top of the search results
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 55s
Three staging-gate failures, two of them one real bug.

After a search, the results-page bar still holds the term in its input,
so on every render the query was >= 2 characters and the suggestion list
opened again — on top of the very results the search had just produced.
Playwright reported it as "<li role=option ...> intercepts pointer
events" while trying to click the first result; a reader would simply
have found their first result unclickable. Both the school-detail and
hero-map journeys failed on it, and neither is about autosuggest.

Suggestions now answer typing, not the mere presence of a value:
`hasTyped` gates the hook, is set on change, and is cleared when a
search is submitted or a suggestion is chosen. A pre-filled input makes
no request and shows no list.

Third failure was my test, not the product. An unphased place page
renders one table per phase, and an all-through school legitimately
appears in both — so the page's school links were never one alphabetical
run. The assertion collected them all together and only passed because
no town it picked had held an all-through school. When the data gave
Abbots Langley one, Breakspeare School appeared in the primary table and
again in the secondary, and the test failed on correct behaviour. It now
checks each table separately, and passes against the data that broke it.

Guards: a jest test that a pre-filled input neither fetches nor opens
(verified by reverting — it is the only one that fails), and an E2E
journey that submits a search and then requires the first result to be
clickable, which is the reader-facing version of the same thing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:38:16 +01:00
tudor 868eb344f5 Merge pull request 'fix(map): the hero map's fade to the header was hardcoded white' (#129) from fix/dark-map-fade into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m49s
Reviewed-on: #129
2026-08-26 20:32:45 +00:00
TudorandClaude Opus 5 0b15497c09 fix(search): the mobile hero search was indented by a card's padding
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 1m3s
Measured at 390px: the headline and lede sit at x=34, while the search
box, the hint and the location link all sat at x=48 and the field was
28px narrower than the copy above it.

The 14px came from `@media (max-width: 768px) { .filterBar { padding:
0.875rem } }`. That rule is for the results filter bar, which is a card
— background, border, shadow — and needs inner padding. The hero search
is not a card: .heroMode strips all of it, padding included.

Both selectors are specificity (0,1,0), so source order decides, and
.heroMode only wins because it is declared right after .filterBar. A
bare .filterBar rule inside a media query comes later and silently wins
instead. The two rules directly below this one in the same block were
already written as `.filterBar:not(.heroMode)`; this one was missed.

Scoping it aligns the search box, hint and location link to the same
left edge as the headline and gives the field back its 28px.

The location link also carried its own 6px of button padding, so its
label started further right than the hint even once the boxes agreed.
Pulled back with a negative margin, which keeps the tap target.

The guard is a stylesheet test: the failure is a plausible-looking
layout rather than a broken one, so nothing short of measuring or
looking would catch it. Verified by reverting: it names ".filterBar sets
padding".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:32:41 +01:00
tudor d55f6cce23 Merge pull request 'fix(suggest): let the dropdown out of the hero panel' (#128) from fix/hero-dropdown-clipping into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m50s
Reviewed-on: #128
2026-08-26 20:23:33 +00:00
TudorandClaude Opus 5 d5a6db289d fix(suggest): let the dropdown out of the hero panel
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 54s
.heroPanel had overflow: hidden to clip its artwork and scrim to the
rounded corners. It clipped the suggestion dropdown too. Measured on
staging with the flag on: the list runs 482 to 802, the panel ends at
624 — so 178px of 320 was cut off, about half the options, with nothing
on screen to say anything was missing.

The two things that actually needed clipping now round themselves:
.heroArt gets border-radius: inherit plus its own overflow, and the
::before scrim inherits the radius. Below 860px the artwork is a band
flush with the top of the panel rather than a layer covering it, so it
takes the top two corners only — inheriting all four would leave it
floating with rounded corners against the copy.

Nothing else depended on the panel clipping: .valueProps below it is
entirely static, so a positioned dropdown paints above it without a
z-index fight.

The regression test asserts the LAST option is the element actually
painted at its own coordinates. toBeVisible() would not have caught
this — it checks for a non-empty box and visibility, and an ancestor's
overflow clips neither. elementFromPoint catches clipping and occlusion
alike.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:10:32 +01:00
205 changed files with 26762 additions and 2326 deletions

No files matched your search

+13 -5
View File
@@ -20,7 +20,7 @@ PORT=80
# =============================================================================
# CORS
# =============================================================================
# Comma-separated list of allowed origins
# JSON array of allowed origins (pydantic-settings format)
# In production, only include your actual domain
ALLOWED_ORIGINS=["https://schoolcompare.co.uk"]
@@ -33,13 +33,21 @@ ADMIN_API_KEY=CHANGE_THIS_TO_A_SECURE_RANDOM_KEY
# Rate limiting (requests per minute per IP)
RATE_LIMIT_PER_MINUTE=60
RATE_LIMIT_BURST=10
GLOBAL_RATE_LIMIT_PER_MINUTE=3000
# Maximum request body size in bytes (default 1MB)
MAX_REQUEST_SIZE=1048576
# =============================================================================
# API
# SEARCH AND OPTIONAL FEATURE FLAGS
# =============================================================================
DEFAULT_PAGE_SIZE=50
MAX_PAGE_SIZE=100
TYPESENSE_URL=http://localhost:8108
TYPESENSE_API_KEY=CHANGE_THIS_TO_YOUR_TYPESENSE_KEY
# Empty URL disables Unleash-backed flags. Match the managed environment when used.
UNLEASH_URL=
UNLEASH_API_TOKEN=
# Page-size limits are currently declared by route Query parameters.
# DEFAULT_PAGE_SIZE, MAX_PAGE_SIZE and RATE_LIMIT_BURST are not reliable tuning
# controls in the current routes; see docs/LEGACY_CODE.md.
+73 -17
View File
@@ -5,6 +5,11 @@ on:
branches:
- main
# Serialise the entire build/deploy/test cycle: no other run can move staging tags.
concurrency:
group: staging-release
cancel-in-progress: false
env:
REGISTRY: privaterepo.sitaru.org
BACKEND_IMAGE_NAME: ${{ gitea.repository }}-backend
@@ -12,7 +17,18 @@ env:
PIPELINE_IMAGE_NAME: ${{ gitea.repository }}-pipeline
jobs:
prepare:
runs-on: ubuntu-latest
outputs:
build_id: ${{ steps.identity.outputs.build_id }}
steps:
- id: identity
run: python3 -c 'import uuid; print("build_id=" + uuid.uuid4().hex)' >> "$GITHUB_OUTPUT"
build-backend:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Backend (FastAPI)
runs-on: ubuntu-latest
steps:
@@ -46,17 +62,24 @@ jobs:
type=raw,value=staging
- name: Build and push Backend Docker image
id: build
uses: docker/build-push-action@v5
with:
context: .
file: ./Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-backend.outputs.tags }}
labels: ${{ steps.meta-backend.outputs.labels }}
cache-from: type=registry,ref=${{ env.REGISTRY }}/${{ env.BACKEND_IMAGE_NAME }}:buildcache
cache-to: type=registry,ref=${{ env.REGISTRY }}/${{ env.BACKEND_IMAGE_NAME }}:buildcache,mode=max
build-frontend:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Frontend (Next.js)
runs-on: ubuntu-latest
steps:
@@ -90,18 +113,23 @@ jobs:
type=raw,value=staging
- name: Build and push Frontend Docker image
id: build
uses: docker/build-push-action@v5
with:
context: ./nextjs-app
file: ./nextjs-app/Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-frontend.outputs.tags }}
labels: ${{ steps.meta-frontend.outputs.labels }}
build-args: |
FASTAPI_URL=http://backend:80/api
# Cache disabled due to registry size limits
build-pipeline:
needs: [prepare]
outputs:
digest: ${{ steps.build.outputs.digest }}
name: Build Pipeline (Meltano + dbt + Airflow)
runs-on: ubuntu-latest
steps:
@@ -135,11 +163,15 @@ jobs:
type=raw,value=staging
- name: Build and push Pipeline Docker image
id: build
uses: docker/build-push-action@v5
with:
context: ./pipeline
file: ./pipeline/Dockerfile
push: true
build-args: |
BUILD_SHA=${{ gitea.sha }}
BUILD_ID=${{ needs.prepare.outputs.build_id }}
tags: ${{ steps.meta-pipeline.outputs.tags }}
labels: ${{ steps.meta-pipeline.outputs.labels }}
cache-from: type=registry,ref=${{ env.REGISTRY }}/${{ env.PIPELINE_IMAGE_NAME }}:buildcache
@@ -148,30 +180,23 @@ jobs:
deploy-staging:
name: Deploy to Staging
runs-on: ubuntu-latest
needs: [build-backend, build-frontend, build-pipeline]
needs: [prepare, build-backend, build-frontend, build-pipeline]
steps:
- name: Trigger staging stack update
run: curl -fsSk -X POST "${{ secrets.PORTAINER_STAGING_WEBHOOK }}"
- name: Wait for staging to become healthy
run: |
echo "Polling ${STAGING_BASE_URL} for up to 5 minutes..."
for i in $(seq 1 60); do
if curl -fsS -o /dev/null --max-time 10 "${STAGING_BASE_URL}/"; then
echo "Staging is up (attempt $i)"
exit 0
fi
sleep 5
done
echo "Staging did not become healthy in time" >&2
exit 1
- uses: actions/checkout@v4
- name: Verify deployed release identity
run: python3 scripts/ci/release.py wait
env:
STAGING_BASE_URL: ${{ secrets.STAGING_BASE_URL }}
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
e2e-staging:
name: E2E Journeys against Staging
runs-on: ubuntu-latest
needs: [deploy-staging]
needs: [prepare, deploy-staging, build-backend, build-frontend, build-pipeline]
steps:
- name: Checkout repository
uses: actions/checkout@v4
@@ -187,11 +212,42 @@ jobs:
npm ci
npx playwright install --with-deps chromium
- name: Verify release before journeys
run: python3 scripts/ci/release.py wait --timeout 10
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- name: Run E2E journeys
working-directory: e2e
run: npx playwright test
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- name: Verify release after journeys
run: python3 scripts/ci/release.py wait --timeout 10
env:
BASE_URL: ${{ secrets.STAGING_BASE_URL }}
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
- uses: docker/setup-buildx-action@v3
- uses: docker/login-action@v3
with:
registry: ${{ env.REGISTRY }}
username: ${{ gitea.actor }}
password: ${{ secrets.REGISTRY_TOKEN }}
- name: Mark tested image digests as verified
run: python3 scripts/ci/release.py verify
env:
EXPECTED_SHA: ${{ gitea.sha }}
EXPECTED_BUILD_ID: ${{ needs.prepare.outputs.build_id }}
BACKEND_DIGEST: ${{ needs.build-backend.outputs.digest }}
FRONTEND_DIGEST: ${{ needs.build-frontend.outputs.digest }}
PIPELINE_DIGEST: ${{ needs.build-pipeline.outputs.digest }}
# Production deployment is a second, manual approval: see promote.yml
# ("Promote to Production (manual)") and docs/DEPLOY.md.
+2 -2
View File
@@ -68,13 +68,13 @@ jobs:
python-version: "3.12"
- name: Install dependencies
run: pip install -r requirements.txt pytest "httpx<0.28"
run: pip install -r requirements.txt pytest "httpx<0.28" pyyaml
- name: Import smoke test
run: python -c "from backend.app import app; print('backend imports OK')"
- name: Backend unit tests
run: python -m pytest backend/tests -q
run: python -m pytest backend/tests pipeline/tests scripts/ci/tests -q
build-backend:
name: Build Backend (no push)
+7 -25
View File
@@ -97,33 +97,15 @@ jobs:
username: ${{ gitea.actor }}
password: ${{ secrets.REGISTRY_TOKEN }}
- name: Retag approved images as prod (keeping rollback pointer)
run: |
SHORT_SHA="${{ steps.resolve.outputs.short }}"
for IMAGE in \
"${REGISTRY}/${BACKEND_IMAGE_NAME}" \
"${REGISTRY}/${FRONTEND_IMAGE_NAME}" \
"${REGISTRY}/${PIPELINE_IMAGE_NAME}"; do
# Keep a rollback pointer before moving :prod
docker buildx imagetools create -t "${IMAGE}:prod-previous" "${IMAGE}:prod" || true
docker buildx imagetools create -t "${IMAGE}:prod" "${IMAGE}:${SHORT_SHA}"
echo "Promoted ${IMAGE}:${SHORT_SHA} -> :prod"
done
- name: Resolve verified digests and promote the complete image set
run: python3 scripts/ci/release.py promote --output release.json
env:
EXPECTED_SHA: ${{ steps.resolve.outputs.full }}
- name: Trigger production stack update
run: curl -fsSk -X POST "${{ secrets.PORTAINER_PROD_WEBHOOK }}"
- name: Wait for production to become healthy
run: |
echo "Polling ${PROD_BASE_URL} for up to 5 minutes..."
for i in $(seq 1 60); do
if curl -fsS -o /dev/null --max-time 10 "${PROD_BASE_URL}/"; then
echo "Production is up (attempt $i)"
exit 0
fi
sleep 5
done
echo "Production did not become healthy in time" >&2
exit 1
- name: Verify production release identity
run: python3 scripts/ci/release.py wait --release release.json
env:
PROD_BASE_URL: ${{ secrets.PROD_BASE_URL }}
BASE_URL: ${{ secrets.PROD_BASE_URL }}
+13 -187
View File
@@ -1,191 +1,17 @@
# Docker Deployment Guide
# Docker deployment
## Quick Start
The maintained deployment runbook is [docs/DEPLOY.md](docs/DEPLOY.md).
Deploy the complete SchoolCompare stack (PostgreSQL + FastAPI + Next.js) with one command:
- Production: `docker-compose.portainer.yml`, using `:prod` images.
- Staging: `docker-compose.portainer.staging.yml`, using `:staging` images.
- Builds and deployment: `.gitea/workflows/deploy.yml`.
- Human-approved production promotion: `.gitea/workflows/promote.yml`.
```bash
docker-compose up -d
```
The generic `docker-compose.yml` is not a supported one-command onboarding path:
it still uses `:latest` tags that the release workflow no longer publishes and
lacks the full current CMS setup. Review the [legacy inventory](docs/LEGACY_CODE.md)
before using old compose examples. Starting an empty database does not populate
school marts.
This will start:
- **PostgreSQL** on port 5432 (database)
- **FastAPI** on port 8000 (backend API)
- **Next.js** on port 3000 (frontend)
## Service Details
### PostgreSQL Database
- **Port**: 5432
- **Container**: `schoolcompare_db`
- **Credentials**:
- User: `schoolcompare`
- Password: `schoolcompare`
- Database: `schoolcompare`
- **Volume**: `postgres_data` (persistent storage)
### FastAPI Backend
- **Port**: 8000 → 80 (container)
- **Container**: `schoolcompare_backend`
- **Built from**: Root `Dockerfile`
- **API Endpoint**: http://localhost:8000/api
- **Health Check**: http://localhost:8000/api/data-info
### Next.js Frontend
- **Port**: 3000
- **Container**: `schoolcompare_nextjs`
- **Built from**: `nextjs-app/Dockerfile`
- **URL**: http://localhost:3000
- **Connects to**: Backend via internal network
## Commands
### Start all services
```bash
docker-compose up -d
```
### View logs
```bash
# All services
docker-compose logs -f
# Specific service
docker-compose logs -f nextjs
docker-compose logs -f backend
docker-compose logs -f db
```
### Check status
```bash
docker-compose ps
```
### Stop all services
```bash
docker-compose down
```
### Rebuild after code changes
```bash
# Rebuild and restart specific service
docker-compose up -d --build nextjs
# Rebuild all services
docker-compose up -d --build
```
### Clean restart (remove volumes)
```bash
docker-compose down -v
docker-compose up -d
```
## Initial Database Setup
After first start, you may need to initialize the database:
```bash
# Enter the backend container
docker exec -it schoolcompare_backend bash
# Run migrations or data loading
python -m backend.data_loader
```
## Accessing Services
Once running:
- **Frontend**: http://localhost:3000
- **Backend API**: http://localhost:8000/api
- **API Docs**: http://localhost:8000/docs (Swagger UI)
- **Database**: localhost:5432 (use any PostgreSQL client)
## Environment Variables
Create a `.env` file in the root directory to customize:
```env
# Database
POSTGRES_USER=schoolcompare
POSTGRES_PASSWORD=your_secure_password
POSTGRES_DB=schoolcompare
# Backend
DATABASE_URL=postgresql://schoolcompare:your_secure_password@db:5432/schoolcompare
# Frontend (for client-side access)
NEXT_PUBLIC_API_URL=http://localhost:8000/api
```
Then run:
```bash
docker-compose up -d
```
## Troubleshooting
### Backend not connecting to database
```bash
# Check database health
docker-compose ps
# View backend logs
docker-compose logs backend
# Restart backend
docker-compose restart backend
```
### Frontend not connecting to backend
```bash
# Check backend health
curl http://localhost:8000/api/data-info
# Check Next.js environment variables
docker exec schoolcompare_nextjs env | grep API
```
### Port already in use
```bash
# Change ports in docker-compose.yml
# For example, change "3000:3000" to "3001:3000"
```
### Rebuild from scratch
```bash
docker-compose down -v
docker system prune -a
docker-compose up -d --build
```
## Production Deployment
For production, update the following:
1. **Use secure passwords** in `.env` file
2. **Configure reverse proxy** (Nginx) in front of Next.js
3. **Enable HTTPS** with SSL certificates
4. **Set production environment variables**:
```env
NODE_ENV=production
POSTGRES_PASSWORD=<strong-password>
```
5. **Backup database** regularly:
```bash
docker exec schoolcompare_db pg_dump -U schoolcompare schoolcompare > backup.sql
```
## Network Architecture
```
Internet
↓
Next.js (port 3000) ← User browsers
↓ (internal network)
FastAPI (port 8000) ← API calls
↓ (internal network)
PostgreSQL (port 5432) ← Data queries
```
All services communicate via the `schoolcompare-network` Docker network.
For architecture, configuration and test commands, see
[ARCHITECTURE.md](docs/ARCHITECTURE.md) and [DEVELOPMENT.md](docs/DEVELOPMENT.md).
+6
View File
@@ -24,6 +24,12 @@ RUN pip install --no-cache-dir -r requirements.txt
COPY backend/ ./backend/
COPY scripts/ ./scripts/
ARG BUILD_SHA=development
ARG BUILD_ID=development
LABEL io.schoolcompare.build-id=$BUILD_ID
LABEL io.schoolcompare.commit=$BUILD_SHA
RUN python -c 'import json,sys; open("backend/build-info.json", "w").write(json.dumps({"sha":sys.argv[1],"build_id":sys.argv[2]}))' "$BUILD_SHA" "$BUILD_ID"
# Expose the application port
EXPOSE 80
+2
View File
@@ -1,3 +1,5 @@
> Historical migration record, retained for context. Setup and architecture claims below may be obsolete. Use [README.md](README.md), [architecture](docs/ARCHITECTURE.md) and [deployment](docs/DEPLOY.md) for current guidance.
# SchoolCompare: Vanilla JS → Next.js Migration Summary
## Overview
+52 -199
View File
@@ -1,214 +1,67 @@
# Primary School Compass 🧒📚
# SchoolCompare
A modern web application for comparing **primary school (KS2)** performance data in **Wandsworth and Merton** over the last 5 years. Built with FastAPI and vanilla JavaScript with Chart.js visualizations.
SchoolCompare compares schools across England: primary (KS2), secondary (KS4),
all-through and post-16 provision, with coverage depending on the source dataset.
It provides school search, postcode maps, comparisons, rankings, place pages,
Ofsted information, admissions and destination measures. Editorial content lives
in a Payload CMS blog.
![Python](https://img.shields.io/badge/Python-3.9+-blue)
![FastAPI](https://img.shields.io/badge/FastAPI-0.109-green)
![License](https://img.shields.io/badge/License-MIT-yellow)
## Start here
## Features
- [Architecture and data flow](docs/ARCHITECTURE.md)
- [Development and validation](docs/DEVELOPMENT.md)
- [Deployment and promotion](docs/DEPLOY.md)
- [Legacy and unused-code inventory](docs/LEGACY_CODE.md)
- [Frontend conventions](nextjs-app/README.md)
- [CMS publishing](nextjs-app/docs/PUBLISHING.md)
- 📊 **Interactive Charts** - Visualize KS2 performance trends over time
- 🔍 **Smart Search** - Find primary schools by name in Wandsworth & Merton
- ⚖️ **Side-by-Side Comparison** - Compare up to 5 schools simultaneously
- 🏆 **Rankings** - View top-performing primary schools by various KS2 metrics
- 📱 **Responsive Design** - Works beautifully on desktop and mobile
## Repository map
## Key Metrics (KS2)
| Path | Responsibility |
|---|---|
| `backend/` | FastAPI routes, cached school data, read-only SQLAlchemy mappings, feature flags |
| `nextjs-app/` | Next.js App Router, React UI, Payload CMS, frontend tests |
| `pipeline/plugins/extractors/` | Custom Singer taps for GIAS, EES, Ofsted and other datasets |
| `pipeline/transform/` | dbt staging/intermediate models, marts, seeds and data tests |
| `pipeline/dags/` | Airflow extraction, transformation and publication workflows |
| `pipeline/scripts/` | Search indexing, code generation and operational diagnostics |
| `e2e/` | Playwright journeys against a running environment |
| `.gitea/workflows/` | PR checks, staging deployment and manual production promotion |
| `scripts/` | CI review tooling and historical data utilities; see the legacy inventory |
| `docs/superpowers/`, `mockups/` | Design history and prototypes, not application entry points |
The application tracks these Key Stage 2 performance indicators:
## Runtime
| Metric | Description |
|--------|-------------|
| **Reading Progress** | Progress in reading from KS1 to KS2 |
| **Writing Progress** | Progress in writing from KS1 to KS2 |
| **Maths Progress** | Progress in maths from KS1 to KS2 |
| **Reading Expected %** | Percentage meeting expected standard in reading |
| **Writing Expected %** | Percentage meeting expected standard in writing |
| **Maths Expected %** | Percentage meeting expected standard in maths |
| **Reading, Writing & Maths Combined %** | Percentage meeting expected standard in all three subjects |
The public site is **Next.js**, not the FastAPI root page. Browser `/api/*`
requests pass through a Next.js route handler to FastAPI. Server-rendered pages
call FastAPI directly using `FASTAPI_URL`, including its `/api` suffix.
## Quick Start
PostgreSQL/PostGIS stores school data. Meltano/Singer extracts source data;
dbt builds `marts.*`; FastAPI reads those tables. Typesense serves text search
and autocomplete. Payload runs inside Next.js and owns a separate `payload`
database schema and uploaded media.
### 1. Clone and Setup
There is **no automatic CSV import or sample dataset on startup**. A working
school-data environment needs populated marts from the pipeline or an approved
database snapshot. See [development](docs/DEVELOPMENT.md) before choosing a setup.
```bash
cd school_results
## Validation
# Create virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install dependencies
pip install -r requirements.txt
```sh
cd nextjs-app
npm ci
npm run typecheck
npm test -- --runInBand
```
### 2. Run the Application
Backend checks, pipeline validation, runtime versions and E2E requirements are
listed in [DEVELOPMENT.md](docs/DEVELOPMENT.md). No `npm run lint` script is
currently defined.
```bash
# Start the server
python -m uvicorn backend.app:app --reload --port 8000
```
Then open http://localhost:8000 in your browser.
The app will run with **sample data** by default, showing **110 primary schools** (66 in Wandsworth, 44 in Merton) with 5 years of KS2 performance data.
### 3. (Optional) Use Real Data
To use real UK school performance data:
1. Visit [Compare School Performance - Download Data](https://www.compare-school-performance.service.gov.uk/download-data)
2. Download **Key Stage 2** data for the years you want (2019-2024)
- Select "Key Stage 2" as the data type
3. Place the CSV files in the `data/` folder
4. Restart the server - it will automatically load and filter to Wandsworth & Merton schools
**Note:** The app only displays schools in Wandsworth and Merton. Data from other areas will be filtered out.
See the helper script for more details:
```bash
python scripts/download_data.py
```
## Project Structure
```
school_results/
├── backend/
│ └── app.py # FastAPI application with all API endpoints
├── frontend/
│ ├── index.html # Main HTML page
│ ├── styles.css # Styling (warm, editorial design)
│ └── app.js # Frontend JavaScript
├── data/
│ └── .gitkeep # Place CSV data files here
├── scripts/
│ └── download_data.py # Helper for downloading/processing data
├── requirements.txt # Python dependencies
└── README.md
```
## API Endpoints
| Endpoint | Description |
|----------|-------------|
| `GET /api/schools` | List schools with optional search/filter |
| `GET /api/schools/{urn}` | Get detailed data for a specific school |
| `GET /api/compare?urns=...` | Compare multiple schools |
| `GET /api/rankings` | Get school rankings by metric |
| `GET /api/filters` | Get available filter options |
| `GET /api/metrics` | Get available performance metrics |
### Example API Usage
```bash
# Search for schools
curl "http://localhost:8000/api/schools?search=academy"
# Get school details
curl "http://localhost:8000/api/schools/100001"
# Compare schools
curl "http://localhost:8000/api/compare?urns=100001,100002,100003"
# Get rankings
curl "http://localhost:8000/api/rankings?metric=rwm_expected_pct&year=2024"
```
## Data Format
If using your own CSV data, ensure it includes these columns (or similar):
| Column | Type | Description |
|--------|------|-------------|
| URN | Integer | Unique Reference Number |
| SCHNAME | String | School name |
| LA | String | Local Authority (must be Wandsworth or Merton) |
| READPROG | Float | Reading progress score |
| WRITPROG | Float | Writing progress score |
| MATPROG | Float | Maths progress score |
| PTRWM_EXP | Float | % meeting expected standard in reading, writing & maths |
| PTREAD_EXP | Float | % meeting expected standard in reading |
| PTWRIT_EXP | Float | % meeting expected standard in writing |
| PTMAT_EXP | Float | % meeting expected standard in maths |
The application normalizes column names automatically and filters to only show Wandsworth and Merton schools.
## Technology Stack
- **Backend**: FastAPI (Python) - High-performance async API framework
- **Frontend**: Vanilla JavaScript with Chart.js
- **Styling**: Custom CSS with CSS variables for theming
- **Data**: Pandas for CSV processing
## Design Philosophy
The UI features a warm, editorial design inspired by quality publications:
- **Typography**: DM Sans for body text, Playfair Display for headings
- **Color Palette**: Warm cream background with coral and teal accents
- **Interactions**: Smooth animations and hover effects
- **Charts**: Clean, readable data visualizations
## Development
```bash
# Run with auto-reload
python -m uvicorn backend.app:app --reload --port 8000
# Or run directly
python backend/app.py
```
## Coverage
This application is specifically designed for:
- **School Phase**: Primary schools only (Key Stage 2)
- **Geographic Area**: Wandsworth and Merton (London boroughs)
- **Time Period**: Last 5 years of data (2020-2024)
Note: 2021 data shows as unavailable because SATs were cancelled due to COVID-19.
## Data Source
Data is sourced from the UK Government's [Compare School Performance](https://www.compare-school-performance.service.gov.uk/) service, which provides official school performance data for England.
**Important**: When using real data, please comply with the [terms of use](https://www.compare-school-performance.service.gov.uk/download-data) and data protection regulations.
## Scheduled Jobs
### Geocoding Schools (Cron Job)
School postcodes are geocoded by a scheduled job, not on-demand. This improves performance and reduces API calls.
**Setup the cron job** (runs weekly on Sunday at 2am):
```bash
# Edit crontab
crontab -e
# Add this line (adjust paths as needed):
0 2 * * 0 cd /path/to/school_compare && /path/to/venv/bin/python scripts/geocode_schools.py >> /var/log/geocode_schools.log 2>&1
```
**Manual run:**
```bash
# Geocode only schools missing coordinates
python scripts/geocode_schools.py
# Force re-geocode all schools
python scripts/geocode_schools.py --force
```
## License
MIT License - feel free to use this project for educational purposes.
---
Built with ❤️ for Wandsworth & Merton families
## Deployment
Work on a feature branch and open a PR. Merging to `main` builds images and
deploys staging. Production promotion is a separate, human-triggered Gitea
workflow. Use [DEPLOY.md](docs/DEPLOY.md) and the Portainer compose files as the
operational references. The generic compose examples still reference `:latest`,
which the current release workflow does not publish.
+191 -46
View File
@@ -26,7 +26,8 @@ from starlette.middleware.base import BaseHTTPMiddleware
import asyncio
from .config import settings
from .data_loader import (
clear_cache,
build_latest_school_data,
load_school_data_as_dataframe,
compute_benchmarks,
load_school_data,
load_latest_school_data,
@@ -38,21 +39,14 @@ from .data_loader import (
)
from .data_loader import get_data_info as get_db_info
from . import flags
from .places import build_place_registry
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .places import build_place_index, build_place_registry, places_for_urn
from .schemas import METRIC_DEFINITIONS, PHASE_GROUPS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .nearby_schools import select_nearby
from .utils import clean_for_json, convert_to_native
# Values to exclude from filter dropdowns (empty strings, non-applicable labels)
EXCLUDED_FILTER_VALUES = {"", "Not applicable", "Does not apply"}
# Maps user-facing phase filter values to the GIAS PhaseOfEducation values they include.
# All-through schools appear in both primary and secondary results.
PHASE_GROUPS: dict[str, set[str]] = {
"primary": {"primary", "middle deemed primary", "all-through"},
"secondary": {"secondary", "middle deemed secondary", "all-through", "16 plus"},
"all-through": {"all-through"},
}
# Must match SITE_URL in nextjs-app/lib/site.ts. The apex 301s to www, and a
# sitemap <loc> that redirects wastes a crawl on every URL it lists.
BASE_URL = "https://www.schoolcompare.co.uk"
@@ -65,6 +59,10 @@ _sitemaps: dict[str, str] | None = None
# Built from the same DataFrame the sitemap uses, so places and sitemap can
# never describe different corpora. Reset by the same admin endpoint.
_place_registry: dict | None = None
# Cached beside the registry, and invalidated by identity against it — see
# get_place_index. Never cleared independently.
_place_index: dict | None = None
_place_index_source: dict | None = None
VALID_PLACE_KINDS = ("town", "locality", "authority", "outcode")
@@ -188,6 +186,24 @@ def get_place_registry() -> dict:
return _place_registry
def get_place_index() -> dict:
"""URN → its published places, cached against the registry it came from.
Invalidation is an identity check rather than a second flag to remember to
clear. Anything that drops `_place_registry` — the tests all do — gets a
fresh registry object here, which no longer matches the one the index was
built from, so the index rebuilds with it. A separate `_place_index = None`
would be one more thing to forget, and a stale reverse index is exactly the
bug that would put links to another dataset's places on a school page.
"""
global _place_index, _place_index_source
registry = get_place_registry()
if _place_index is None or _place_index_source is not registry:
_place_index = build_place_index(registry)
_place_index_source = registry
return _place_index
def _urlset(rows: list[str]) -> str:
return "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
@@ -211,7 +227,67 @@ def _place_url(place) -> str:
return f"/schools/{place.slug}"
def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
def _places_payload(urn: int) -> list[dict]:
"""The published places containing this school, as the school page needs
them: a name to write in the link, a count so the anchor can say what it
leads to, and the canonical path.
`phases` carries the phase variants this school actually appears on, which
is usually one and is two for an all-through school — it is listed on both
pages, so there is no tie to break.
Membership is read straight from the registry's own `phase_urns` rather
than re-derived from the school's phase string. The registry is the one
place that decides which phases a place publishes and who is on them;
computing it a second time here is how a page comes to link a school to a
phase page that does not list it, or to a route that does not exist. That
is also why outcodes need no special case: they carry empty `phase_urns`,
so they report no phase links on their own.
"""
payload = []
for place in places_for_urn(get_place_index(), int(urn)):
phases = [
{
"phase": phase,
"count": len(phase_urns),
"url": f"{_place_url(place)}/{phase}",
}
for phase, phase_urns in sorted(place.phase_urns.items())
if int(urn) in phase_urns
]
payload.append({
"kind": place.kind,
"slug": place.slug,
"name": place.name,
"count": len(place.urns),
"url": _place_url(place),
"phases": phases,
})
return payload
def _nearby_schools_payload(urn: int) -> list[dict]:
"""The nearest eligible schools this page may offer, closest first.
Phase and reach are read from the school's own row inside select_nearby,
so nothing here can hand it a phase that disagrees with the data.
Wrapped: a failure in selection must never 500 a page that is otherwise
complete, which is the posture get_supplementary_data already takes. The
section simply does not render.
"""
try:
return select_nearby(load_latest_school_data(), int(urn))
except Exception:
import logging
logging.getLogger(__name__).exception(
"Nearby schools selection failed for urn=%s", urn
)
return []
def _place_sitemap_rows(kinds: tuple[str, ...], registry=None) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
@@ -221,7 +297,9 @@ def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
linked from the place page either.
"""
rows: list[str] = []
for p in sorted(get_place_registry().values(), key=lambda p: (p.kind, p.slug)):
if registry is None:
registry = get_place_registry()
for p in sorted(registry.values(), key=lambda p: (p.kind, p.slug)):
if p.kind not in kinds:
continue
rows.append(_url_element(BASE_URL + _place_url(p)))
@@ -235,9 +313,10 @@ def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
return rows
def build_sitemaps() -> dict[str, str]:
def build_sitemaps(df=None, registry=None) -> dict[str, str]:
"""Build the sitemap index and every child, keyed by name."""
df = load_school_data()
if df is None:
df = load_school_data()
children: dict[str, str] = {
"static.xml": _urlset(
@@ -257,7 +336,7 @@ def build_sitemaps() -> dict[str, str]:
# measured apart from the school pages'.
for label, kinds in (("places", ("town", "locality", "authority")),
("outcodes", ("outcode",))):
rows = _place_sitemap_rows(kinds)
rows = _place_sitemap_rows(kinds, registry)
chunks = [rows[i:i + SITEMAP_CHUNK_SIZE]
for i in range(0, len(rows), SITEMAP_CHUNK_SIZE)] or [[]]
for n, chunk in enumerate(chunks, start=1):
@@ -639,6 +718,15 @@ async def get_config():
}
@app.get("/api/release")
async def release_identity():
import json
from pathlib import Path
path = Path(__file__).with_name("build-info.json")
identity = json.loads(path.read_text()) if path.exists() else {"sha": "development", "build_id": "development"}
return JSONResponse(identity, headers={"Cache-Control": "no-store"})
@app.get("/api/schools")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_schools(
@@ -675,7 +763,7 @@ async def get_schools(
df_latest = load_latest_school_data()
if df_latest.empty:
return {"schools": [], "total": 0, "page": page, "page_size": 0}
raise HTTPException(status_code=503, detail="School data temporarily unavailable")
# Use configured default if not specified
if page_size is None:
@@ -774,8 +862,8 @@ async def get_schools(
# Apply filters
if search:
ts_urns = search_schools_typesense(search)
if ts_urns:
ts_urns = await asyncio.to_thread(search_schools_typesense, search)
if ts_urns is not None:
urn_order = {urn: i for i, urn in enumerate(ts_urns)}
schools_df = schools_df[schools_df["urn"].isin(set(ts_urns))].copy()
schools_df["_ts_rank"] = schools_df["urn"].map(urn_order)
@@ -783,9 +871,9 @@ async def get_schools(
else:
# Fallback: Typesense unavailable, use substring match
search_lower = search.lower()
mask = schools_df["school_name"].str.lower().str.contains(search_lower, na=False)
mask = schools_df["school_name"].str.lower().str.contains(search_lower, na=False, regex=False)
if "address" in schools_df.columns:
mask = mask | schools_df["address"].str.lower().str.contains(search_lower, na=False)
mask = mask | schools_df["address"].str.lower().str.contains(search_lower, na=False, regex=False)
schools_df = schools_df[mask]
if local_authority:
@@ -844,7 +932,7 @@ async def get_school_details(request: Request, urn: int):
df = load_school_data()
if df.empty:
raise HTTPException(status_code=404, detail="No data available")
raise HTTPException(status_code=503, detail="School data temporarily unavailable")
school_data = df[df["urn"] == urn]
@@ -902,6 +990,17 @@ async def get_school_details(request: Request, urn: int):
return {
"school_info": school_info,
# Where this school sits in the location layer, for the page's link
# module and breadcrumb. Derived from the same registry the place
# pages and the sitemap use, so a link is never offered for a page
# that does not exist. Empty is a valid answer: a school whose town
# and authority both fall below the publish threshold has nowhere to
# point, and the page renders without the module.
"places": _places_payload(urn),
# The nearest eligible schools, closest first. Always present on a
# build with this code; the frontend treats absent and empty
# identically, which is what lets the two images deploy independently.
"nearby_schools": _nearby_schools_payload(urn),
"yearly_data": clean_for_json(school_data),
# Supplementary data (null if not yet populated by Kestra)
"ofsted": supplementary.get("ofsted"),
@@ -919,6 +1018,7 @@ async def get_school_details(request: Request, urn: int):
"phonics": supplementary.get("phonics"),
"deprivation": supplementary.get("deprivation"),
"finance": supplementary.get("finance"),
"destinations": supplementary.get("destinations"),
}
@@ -1337,9 +1437,22 @@ async def get_place(request: Request, kind: str, slug: str,
for m in ("rwm_expected_pct", "attainment_8_score")
}
cols = [c for c in SCHOOL_COLUMNS + ["latitude", "longitude", "phase",
"rwm_expected_pct", "attainment_8_score",
"total_pupils"]
# dict.fromkeys, not a list: SCHOOL_COLUMNS already ends with latitude and
# longitude, so concatenating them again selected each twice and pandas
# dropped one of every duplicated pair with a "columns are not unique"
# warning. Ordered de-duplication keeps the column order and the warning
# cannot come back.
#
# nursery_provision and parliamentary_constituency are not in
# SCHOOL_COLUMNS and the place table shows both. The `in rows.columns`
# guard is what keeps a mart the pipeline has not rebuilt working: those
# two are the optional GIAS columns data_loader degrades to NULL.
cols = [c for c in dict.fromkeys(
SCHOOL_COLUMNS + ["latitude", "longitude", "phase",
"nursery_provision",
"parliamentary_constituency",
"rwm_expected_pct", "attainment_8_score",
"total_pupils"])
if c in rows.columns]
return {
@@ -1460,20 +1573,51 @@ async def get_data_info(request: Request):
}
_publication_lock = asyncio.Lock()
def _prepare_publication(df):
if df.empty:
raise ValueError("Refusing to publish an empty school dataset")
if not {"urn", "year", "school_name"}.issubset(df.columns):
raise ValueError("School dataset is missing required columns")
if df["urn"].isna().any() or df.duplicated(["urn", "year"]).any():
raise ValueError("School dataset has missing URNs or duplicate school years")
latest = build_latest_school_data(df)
registry = build_place_registry(df)
index = build_place_index(registry)
sitemaps = build_sitemaps(df, registry)
return df, latest, registry, index, sitemaps
def _publish(prepared):
# Called on the event loop with no await: routes cannot observe half a swap.
# The application currently runs one worker; replicas require coordination.
from . import data_loader
global _place_registry, _place_index, _place_index_source, _sitemaps
df, latest, registry, index, sitemaps = prepared
data_loader._df_cache = df
data_loader._df_latest_cache = latest
_place_registry = registry
_place_index = index
_place_index_source = registry
_sitemaps = sitemaps
@app.post("/api/admin/reload")
@limiter.limit("5/minute")
async def reload_data(
request: Request,
_: bool = Depends(verify_admin_api_key)
):
"""
Admin endpoint to force data reload (useful after data updates).
Requires X-API-Key header with valid admin API key.
"""
clear_cache()
await asyncio.to_thread(load_school_data)
await asyncio.to_thread(load_latest_school_data)
return {"status": "reloaded"}
async def reload_data(request: Request, _: bool = Depends(verify_admin_api_key)):
"""Validate a complete replacement before publishing it; retain data on failure."""
async with _publication_lock:
try:
df = await asyncio.to_thread(load_school_data_as_dataframe)
prepared = await asyncio.to_thread(_prepare_publication, df)
except Exception as exc:
import logging
logging.getLogger(__name__).exception("Dataset reload failed")
raise HTTPException(status_code=503, detail="Dataset reload failed; previous data retained") from exc
_publish(prepared)
return {"status": "reloaded", "schools": len(prepared[1])}
@@ -1525,15 +1669,16 @@ async def regenerate_sitemap(
request: Request,
_: bool = Depends(verify_admin_api_key),
):
"""Rebuild and cache the sitemap from current school data. Called by Airflow after data updates."""
global _sitemaps, _place_registry
# Places and sitemap are rebuilt together — they read the same marts, and
# letting them drift apart would submit URLs for places that no longer
# exist.
_place_registry = None
_sitemaps = build_sitemaps()
n = sum(x.count("<url>") for x in _sitemaps.values())
return {"status": "ok", "urls": n, "sitemaps": len(_sitemaps)}
"""Rebuild derived publication data without clearing the live registry."""
async with _publication_lock:
try:
prepared = await asyncio.to_thread(_prepare_publication, load_school_data())
except Exception as exc:
raise HTTPException(status_code=503, detail="Sitemap rebuild failed; previous data retained") from exc
_publish(prepared)
n = sum(x.count("<url>") for x in prepared[4].values())
return {"status": "ok", "urls": n, "sitemaps": len(prepared[4])}
# Mount static files directly (must be after all routes to avoid catching API calls)
+301 -24
View File
@@ -20,6 +20,7 @@ from .models import (
DimSchool, DimLocation, KS2Performance,
FactOfstedInspection, FactAdmissions, FactAdmissionDistance,
FactDeprivation, FactFinance, FactPupilCharacteristics,
FactKs4Destinations, FactKs5Destinations,
)
from .ofsted_codes import ofsted_page_url, report_card_labels
from .schemas import SCHOOL_TYPE_MAP
@@ -83,21 +84,58 @@ def _get_typesense_client():
return None
def search_schools_typesense(query: str, limit: int = 250) -> List[int]:
"""Search Typesense. Returns URNs in relevance order, or [] if unavailable."""
SEARCH_PAGE_SIZE = 250
# Search results are filtered again by the API (authority, phase, postcode,
# etc.), so one page is too small for scoped searches. Keep the candidate set
# bounded, though: a broad query must not turn into an unbounded sequence of
# Typesense requests. Four pages is enough to preserve useful scoped matches
# while putting a hard ceiling on latency and upstream load.
SEARCH_MAX_CANDIDATES = 1_000
def search_schools_typesense(query: str) -> Optional[List[int]]:
"""Return a bounded set of matching URNs in relevance order.
``None`` means Typesense is unavailable; ``[]`` is a valid zero-match
result. The API applies its remaining filters after this search, so the
first few pages are fetched rather than only the first page. Once the
candidate ceiling is reached, the relevance-ordered prefix is returned on
purpose; fetching every match would make common or adversarial queries
unbounded.
"""
client = _get_typesense_client()
if client is None:
return []
return None
urns: list[int] = []
fetched = 0
try:
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority,postcode",
"per_page": min(limit, 250),
"typo_tokens_threshold": 1,
})
return [int(h["document"]["urn"]) for h in result.get("hits", [])]
page = 1
while fetched < SEARCH_MAX_CANDIDATES:
page_size = min(SEARCH_PAGE_SIZE, SEARCH_MAX_CANDIDATES - fetched)
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority,postcode",
"per_page": page_size,
"page": page,
"typo_tokens_threshold": 1,
})
hits = result.get("hits", [])
urns.extend(int(h["document"]["urn"]) for h in hits)
fetched += len(hits)
if fetched >= result.get("found", fetched):
return list(dict.fromkeys(urns))
if not hits:
raise ValueError("Search pagination ended before all matches arrived")
page += 1
logging.getLogger(__name__).info(
"Typesense search capped at %d candidates for query %r",
SEARCH_MAX_CANDIDATES,
query,
)
return list(dict.fromkeys(urns))
except Exception:
return []
logging.getLogger(__name__).exception("School search unavailable")
return None
# The most a public endpoint will return in one response.
@@ -187,16 +225,6 @@ def geocode_single_postcode(postcode: str) -> Optional[Tuple[float, float]]:
return None
def haversine_distance(lat1: float, lon1: float, lat2: float, lon2: float) -> float:
"""Calculate great-circle distance between two points (miles)."""
from math import radians, cos, sin, asin, sqrt
lat1, lon1, lat2, lon2 = map(radians, [lat1, lon1, lat2, lon2])
dlat = lat2 - lat1
dlon = lon2 - lon1
a = sin(dlat / 2) ** 2 + cos(lat1) * cos(lat2) * sin(dlon / 2) ** 2
return 2 * asin(sqrt(a)) * 3956
# =============================================================================
# MAIN DATA LOAD — joins dim_school + dim_location + fact_performance
# fact_performance is a merged KS2+KS4 table (one row per URN per year).
@@ -511,7 +539,12 @@ def load_latest_school_data() -> pd.DataFrame:
if _df_latest_cache is not None:
return _df_latest_cache
df = load_school_data()
_df_latest_cache = build_latest_school_data(load_school_data())
return _df_latest_cache
def build_latest_school_data(df: pd.DataFrame) -> pd.DataFrame:
"""Build a replacement snapshot without mutating the published caches."""
if df.empty:
return df
@@ -544,8 +577,7 @@ def load_latest_school_data() -> pd.DataFrame:
df_latest = pd.concat([df_latest, df_no_perf], ignore_index=True)
print(f"Latest-snapshot cache built: {len(df_latest)} schools")
_df_latest_cache = df_latest
return _df_latest_cache
return df_latest
def clear_cache():
@@ -816,6 +848,218 @@ def _finance_dict(f) -> dict:
}
# Destination measures that are totals DfE published itself, rather than one of
# the categories that partition the cohort.
_AGGREGATE_MEASURES = {"agg_sustained_education", "agg_sustained_all"}
def _format_cohort_year(year) -> str | None:
"""202223 -> '2022/23'.
The section has to date its own cohort. Destination measures run about two
GCSE years behind the results shown above them on the same page, so an
undated figure reads as stale data rather than as a different question.
"""
if not year:
return None
text = str(year)
if len(text) == 6:
return f"{text[:4]}/{text[4:6]}"
if len(text) == 8:
return f"{text[:4]}/{text[6:8]}"
return text
_PUPIL_GROUPS = ("disadvantaged", "other", "all")
def _lone_hidden_groups(groups: dict) -> list:
"""Pupil groups hiding exactly one category — solvable by subtraction."""
return [
key for key, group in groups.items()
if sum(1 for c in group["categories"] if c["status"] == "suppressed") == 1
]
def _lone_hidden_categories(groups: dict) -> list:
"""Categories hidden in exactly one of several pupil groups."""
lone = []
categories = {c["category"] for g in groups.values() for c in g["categories"]}
for category in categories:
found = [
c for g in groups.values() for c in g["categories"]
if c["category"] == category
]
hidden = [c for c in found if c["status"] == "suppressed"]
if len(hidden) == 1 and len(found) > 1:
lone.append(category)
return lone
def disclosure_invariant_holds(groups: dict) -> bool:
"""Every row and every column hides none, or at least two.
Public so the tests can assert it directly rather than re-deriving it.
"""
return not _lone_hidden_groups(groups) and not _lone_hidden_categories(groups)
def _mask_for_disclosure(groups: dict) -> None:
"""Withhold further cells until nothing suppressed can be solved for.
Not rendering a figure is not the same as not publishing it. This endpoint
is public and unauthenticated, so anything left in the payload is
published, whatever the UI chooses to draw — the same reasoning the
admission_distance field carries in app.py.
Two identities let a caller solve for a withheld cell:
* within a pupil group, the categories sum to the cohort, so a group with
exactly ONE suppressed category gives it away as cohort - sum(rest);
* across groups, disadvantaged + other = all for every category, so a
category suppressed in exactly ONE of the three gives itself away.
DfE's own answer is secondary suppression: withhold a second cell so the
residual spans two unknowns and identifies neither.
Where no companion can do that — a sparse cohort whose every other category
is `not_applicable`, which is common in special schools and alternative
provision — there is nothing left to withhold, so the pupil group is
DROPPED entirely. An earlier version simply gave up here and returned with
the violation intact and no signal, which is the one outcome this function
must never produce: a disclosure-control pass that fails silently is worse
than none, because everything downstream trusts it.
Mutates `groups` in place. Guaranteed to return with
disclosure_invariant_holds(groups) true.
"""
def suppress(cell):
if cell["status"] == "published":
cell["status"] = "suppressed"
cell["pupils"] = None
cell["percentage"] = None
return True
return False
def add_companion(candidates) -> bool:
"""Withhold a second cell so the residual spans two unknowns.
The companion must carry pupils. Suppressing a zero looks like
secondary suppression and protects nothing: the residual still equals
the original withheld figure exactly. Returns False when no cell can
do the job, which escalates to dropping the group.
"""
published = [c for c in candidates if c["status"] == "published"]
useful = sorted(
(c for c in published if (c["pupils"] or 0) > 0),
key=lambda c: c["pupils"],
)
if useful:
return suppress(useful[0])
# Every remaining cell is zero or not applicable: withholding any of
# them leaves the residual equal to the original figure.
return False
# Fixpoint: each new suppression can break the other identity. Terminates
# because every pass either adds a suppression, drops a group, or stops.
while not disclosure_invariant_holds(groups):
changed = False
for category in _lone_hidden_categories(groups):
siblings = [
c for g in groups.values() for c in g["categories"]
if c["category"] == category
]
if add_companion(siblings):
changed = True
for key in _lone_hidden_groups(groups):
if add_companion(groups[key]["categories"]):
changed = True
if changed:
continue
# Nothing left to withhold. Drop the groups that are still solvable,
# and any category still solvable across the groups that remain.
for key in _lone_hidden_groups(groups):
del groups[key]
changed = True
for category in _lone_hidden_categories(groups):
for group in groups.values():
for cell in group["categories"]:
if cell["category"] == category and suppress(cell):
changed = True
if not changed:
# Unreachable given the two escalations above, but a masking pass
# must never spin or exit unsafely. Withhold everything.
groups.clear()
return
def _destinations_block(rows: list) -> dict | None:
"""Shape destination rows for one phase into the API's block.
Applies secondary suppression before returning, so no caller of this public
endpoint can solve for a figure DfE withheld. See _mask_for_disclosure.
Aggregate measures are dropped entirely. DfE publishes them, and they would
be useful for a "what is published for this group" fallback, but nothing
renders them today and an aggregate spanning exactly one suppressed
component names that component. An unused field that leaks is not a
trade-off worth carrying — re-add them with their own guard if the fallback
is ever built.
Deliberately computes no residual, no "remaining pupils" figure, and no
total that would close a gap left by a suppressed category.
"""
if not rows:
return None
years = [r["year"] for r in rows if r.get("year") is not None]
if not years:
return None
latest_year = max(years)
rows = [r for r in rows if r.get("year") == latest_year]
groups: dict = {}
for row in rows:
group = groups.setdefault(
row["pupil_group"],
{"cohort": row.get("cohort_pupils"), "categories": []},
)
measure = row["destination_measure"]
published = row.get("status") == "published"
# Belt and braces: percentage is derived from the same source cell as
# pupils, but publishing one without the other would hand back the
# cohort (pupils / percentage) and with it the residual.
cell = {
"category": measure,
"pupils": row.get("pupils") if published else None,
"percentage": row.get("percentage") if published else None,
"status": row.get("status"),
}
if measure in _AGGREGATE_MEASURES:
continue
group["categories"].append(cell)
if not groups:
return None
_mask_for_disclosure(groups)
# Masking can empty the block entirely — a sparse cohort where no group
# could be made safe. Return None so the section is absent rather than
# rendering an empty shell.
if not groups:
return None
return {"cohort_year": _format_cohort_year(latest_year), "groups": groups}
def _empty_supplementary() -> dict:
return {
"ofsted": None,
@@ -827,6 +1071,7 @@ def _empty_supplementary() -> dict:
"phonics": None,
"deprivation": None,
"finance": None,
"destinations": None,
}
@@ -954,6 +1199,38 @@ def get_supplementary_data_batch(db: Session, urns: list[int]) -> dict:
result[f.urn]["finance"] = _finance_dict(f)
_safe(_finance)
# Destinations — KS4 and 16-18. Both marts are long-format, so every row
# for a URN is collected and _destinations_block picks the latest year and
# shapes the pupil groups. A phase with no rows serialises as null rather
# than an empty shell, so the frontend renders nothing rather than an empty
# section.
def _destinations():
from collections import defaultdict
def _collect(model):
per_urn = defaultdict(list)
for r in db.query(model).filter(model.urn.in_(urns)).all():
per_urn[r.urn].append({
"year": r.year,
"pupil_group": r.pupil_group,
"destination_measure": r.destination_measure,
"cohort_pupils": r.cohort_pupils,
"pupils": r.pupils,
"percentage": r.percentage,
"status": r.status,
})
return per_urn
ks4_rows = _collect(FactKs4Destinations)
ks5_rows = _collect(FactKs5Destinations)
for urn in urns:
ks4 = _destinations_block(ks4_rows.get(urn, []))
ks5 = _destinations_block(ks5_rows.get(urn, []))
result[urn]["destinations"] = (
{"ks4": ks4, "ks5": ks5} if (ks4 or ks5) else None
)
_safe(_destinations)
return result
+17
View File
@@ -57,6 +57,23 @@ REGISTRY: dict[str, Flag] = {
),
added=date(2026, 8, 26),
),
Flag(
name="about_page",
description=(
"The /about page, its footer link, its sitemap entry, and the "
"named-author byline on every blog post."
),
added=date(2026, 9, 8),
),
Flag(
name="blog",
description=(
"The /blog index, post pages, the RSS feed, their footer link "
"and their sitemap entries. Not /admin: posts must be "
"writable before the blog is readable."
),
added=date(2026, 9, 8),
),
)
}
-512
View File
@@ -1,512 +0,0 @@
"""
Database migration logic for importing CSV data.
Used by both CLI script and automatic startup migration.
"""
import re
from pathlib import Path
from typing import Dict, Optional
import numpy as np
import pandas as pd
import requests
from .config import settings
from .database import Base, engine, get_db_session
from .models import School, SchoolResult
from .schemas import (
COLUMN_MAPPINGS,
LA_CODE_TO_NAME,
NULL_VALUES,
SCHOOL_TYPE_MAP,
)
def parse_numeric(value) -> Optional[float]:
"""Parse a numeric value, handling special cases."""
if pd.isna(value):
return None
if isinstance(value, (int, float)):
return float(value) if not np.isnan(value) else None
str_val = str(value).strip().upper()
if str_val in NULL_VALUES or str_val == "":
return None
# Remove percentage signs if present
str_val = str_val.replace("%", "")
try:
return float(str_val)
except ValueError:
return None
def extract_year_from_folder(folder_name: str) -> Optional[int]:
"""Extract year from folder name like '2023-2024'."""
match = re.search(r"(\d{4})-(\d{4})", folder_name)
if match:
return int(match.group(2))
match = re.search(r"(\d{4})", folder_name)
if match:
return int(match.group(1))
return None
def geocode_postcodes_bulk(postcodes: list) -> Dict[str, tuple]:
"""
Geocode postcodes in bulk using postcodes.io API.
Returns dict of postcode -> (latitude, longitude).
"""
results = {}
valid_postcodes = [
p.strip().upper()
for p in postcodes
if p and isinstance(p, str) and len(p.strip()) >= 5
]
valid_postcodes = list(set(valid_postcodes))
if not valid_postcodes:
return results
batch_size = 100
total_batches = (len(valid_postcodes) + batch_size - 1) // batch_size
for i, batch_start in enumerate(range(0, len(valid_postcodes), batch_size)):
batch = valid_postcodes[batch_start : batch_start + batch_size]
print(
f" Geocoding batch {i + 1}/{total_batches} ({len(batch)} postcodes)..."
)
try:
response = requests.post(
"https://api.postcodes.io/postcodes",
json={"postcodes": batch},
timeout=30,
)
if response.status_code == 200:
data = response.json()
for item in data.get("result", []):
if item and item.get("result"):
pc = item["query"].upper()
lat = item["result"].get("latitude")
lon = item["result"].get("longitude")
if lat and lon:
results[pc] = (lat, lon)
except Exception as e:
print(f" Warning: Geocoding batch failed: {e}")
return results
def load_csv_data(data_dir: Path) -> pd.DataFrame:
"""Load all CSV data from data directory."""
all_data = []
for folder in sorted(data_dir.iterdir()):
if not folder.is_dir():
continue
year = extract_year_from_folder(folder.name)
if not year:
continue
# Specifically look for the KS2 results file
ks2_file = folder / "england_ks2final.csv"
if not ks2_file.exists():
continue
csv_file = ks2_file
print(f" Loading {csv_file.name} (year {year})...")
try:
df = pd.read_csv(csv_file, encoding="latin-1", low_memory=False)
except Exception as e:
print(f" Error loading {csv_file}: {e}")
continue
# Rename columns
df.rename(columns=COLUMN_MAPPINGS, inplace=True)
df["year"] = year
# Handle local authority name
la_name_cols = ["LANAME", "LA (name)", "LA_NAME", "LA NAME"]
la_name_col = next((c for c in la_name_cols if c in df.columns), None)
if la_name_col and la_name_col != "local_authority":
df["local_authority"] = df[la_name_col]
elif "LEA" in df.columns:
df["local_authority_code"] = pd.to_numeric(df["LEA"], errors="coerce")
df["local_authority"] = (
df["local_authority_code"]
.map(LA_CODE_TO_NAME)
.fillna(df["LEA"].astype(str))
)
# Store LEA code
if "LEA" in df.columns:
df["local_authority_code"] = pd.to_numeric(df["LEA"], errors="coerce")
# Map school type
if "school_type_code" in df.columns:
df["school_type"] = (
df["school_type_code"]
.map(SCHOOL_TYPE_MAP)
.fillna(df["school_type_code"])
)
# Create combined address
addr_parts = ["address1", "address2", "town", "postcode"]
for col in addr_parts:
if col not in df.columns:
df[col] = None
df["address"] = df.apply(
lambda r: ", ".join(
str(v)
for v in [
r.get("address1"),
r.get("address2"),
r.get("town"),
r.get("postcode"),
]
if pd.notna(v) and str(v).strip()
),
axis=1,
)
all_data.append(df)
print(f" Loaded {len(df)} records")
if all_data:
result = pd.concat(all_data, ignore_index=True)
print(f"\nTotal records loaded: {len(result)}")
print(f"Unique schools: {result['urn'].nunique()}")
print(f"Years: {sorted(result['year'].unique())}")
return result
return pd.DataFrame()
def migrate_data(df: pd.DataFrame, geocode: bool = False, geocode_cache: dict = None):
"""Migrate DataFrame data to database."""
if geocode_cache is None:
geocode_cache = {}
# Clean URN column - convert to integer, drop invalid values
df = df.copy()
df["urn"] = pd.to_numeric(df["urn"], errors="coerce")
df = df.dropna(subset=["urn"])
df["urn"] = df["urn"].astype(int)
# Group by URN to get unique schools (use latest year's data)
school_data = (
df.sort_values("year", ascending=False).groupby("urn").first().reset_index()
)
print(f"\nMigrating {len(school_data)} unique schools...")
# Geocode postcodes that aren't already in the cache
geocoded = dict(geocode_cache) # start with preserved coordinates
if geocode and "postcode" in df.columns:
cached_postcodes = {
str(row.get("postcode", "")).strip().upper()
for _, row in school_data.iterrows()
if int(float(str(row.get("urn", 0) or 0))) in geocode_cache
}
postcodes_needed = [
p for p in df["postcode"].dropna().unique()
if str(p).strip().upper() not in cached_postcodes
]
if postcodes_needed:
print(f"\nGeocoding {len(postcodes_needed)} postcodes ({len(geocode_cache)} restored from cache)...")
fresh = geocode_postcodes_bulk(postcodes_needed)
geocoded.update(fresh)
print(f" Successfully geocoded {len(fresh)} new postcodes")
else:
print(f"\nAll {len(geocode_cache)} postcodes restored from cache, skipping geocoding.")
with get_db_session() as db:
# Create schools
urn_to_school_id = {}
schools_created = 0
for _, row in school_data.iterrows():
# Safely parse URN - handle None, NaN, whitespace, and invalid values
urn_val = row.get("urn")
urn = None
if pd.notna(urn_val):
try:
urn_str = str(urn_val).strip()
if urn_str:
urn = int(float(urn_str)) # Handle "12345.0" format
except (ValueError, TypeError):
pass
if not urn:
continue
# Skip if we've already added this URN (handles duplicates in source data)
if urn in urn_to_school_id:
continue
# Get geocoding data
postcode = row.get("postcode")
lat, lon = None, None
if postcode and pd.notna(postcode):
coords = geocoded.get(str(postcode).strip().upper())
if coords:
lat, lon = coords
# Safely parse local_authority_code
la_code = None
la_code_val = row.get("local_authority_code")
if pd.notna(la_code_val):
try:
la_code_str = str(la_code_val).strip()
if la_code_str:
la_code = int(float(la_code_str))
except (ValueError, TypeError):
pass
school = School(
urn=urn,
school_name=row.get("school_name")
if pd.notna(row.get("school_name"))
else "Unknown",
local_authority=row.get("local_authority")
if pd.notna(row.get("local_authority"))
else None,
local_authority_code=la_code,
school_type=row.get("school_type")
if pd.notna(row.get("school_type"))
else None,
school_type_code=row.get("school_type_code")
if pd.notna(row.get("school_type_code"))
else None,
religious_denomination=row.get("religious_denomination")
if pd.notna(row.get("religious_denomination"))
else None,
age_range=row.get("age_range")
if pd.notna(row.get("age_range"))
else None,
address1=row.get("address1") if pd.notna(row.get("address1")) else None,
address2=row.get("address2") if pd.notna(row.get("address2")) else None,
town=row.get("town") if pd.notna(row.get("town")) else None,
postcode=row.get("postcode") if pd.notna(row.get("postcode")) else None,
latitude=lat,
longitude=lon,
)
db.add(school)
db.flush() # Get the ID
urn_to_school_id[urn] = school.id
schools_created += 1
if schools_created % 1000 == 0:
print(f" Created {schools_created} schools...")
print(f" Created {schools_created} schools")
# Create results
print(f"\nMigrating {len(df)} yearly results...")
results_created = 0
for _, row in df.iterrows():
# Safely parse URN
urn_val = row.get("urn")
urn = None
if pd.notna(urn_val):
try:
urn_str = str(urn_val).strip()
if urn_str:
urn = int(float(urn_str))
except (ValueError, TypeError):
pass
if not urn or urn not in urn_to_school_id:
continue
school_id = urn_to_school_id[urn]
# Safely parse year
year_val = row.get("year")
year = None
if pd.notna(year_val):
try:
year = int(float(str(year_val).strip()))
except (ValueError, TypeError):
pass
if not year:
continue
result = SchoolResult(
school_id=school_id,
year=year,
total_pupils=parse_numeric(row.get("total_pupils")),
eligible_pupils=parse_numeric(row.get("eligible_pupils")),
# Expected Standard
rwm_expected_pct=parse_numeric(row.get("rwm_expected_pct")),
reading_expected_pct=parse_numeric(row.get("reading_expected_pct")),
writing_expected_pct=parse_numeric(row.get("writing_expected_pct")),
maths_expected_pct=parse_numeric(row.get("maths_expected_pct")),
gps_expected_pct=parse_numeric(row.get("gps_expected_pct")),
science_expected_pct=parse_numeric(row.get("science_expected_pct")),
# Higher Standard
rwm_high_pct=parse_numeric(row.get("rwm_high_pct")),
reading_high_pct=parse_numeric(row.get("reading_high_pct")),
writing_high_pct=parse_numeric(row.get("writing_high_pct")),
maths_high_pct=parse_numeric(row.get("maths_high_pct")),
gps_high_pct=parse_numeric(row.get("gps_high_pct")),
# Progress
reading_progress=parse_numeric(row.get("reading_progress")),
writing_progress=parse_numeric(row.get("writing_progress")),
maths_progress=parse_numeric(row.get("maths_progress")),
# Averages
reading_avg_score=parse_numeric(row.get("reading_avg_score")),
maths_avg_score=parse_numeric(row.get("maths_avg_score")),
gps_avg_score=parse_numeric(row.get("gps_avg_score")),
# Context
disadvantaged_pct=parse_numeric(row.get("disadvantaged_pct")),
eal_pct=parse_numeric(row.get("eal_pct")),
sen_support_pct=parse_numeric(row.get("sen_support_pct")),
sen_ehcp_pct=parse_numeric(row.get("sen_ehcp_pct")),
stability_pct=parse_numeric(row.get("stability_pct")),
# Absence
reading_absence_pct=parse_numeric(row.get("reading_absence_pct")),
gps_absence_pct=parse_numeric(row.get("gps_absence_pct")),
maths_absence_pct=parse_numeric(row.get("maths_absence_pct")),
writing_absence_pct=parse_numeric(row.get("writing_absence_pct")),
science_absence_pct=parse_numeric(row.get("science_absence_pct")),
# Gender
rwm_expected_boys_pct=parse_numeric(row.get("rwm_expected_boys_pct")),
rwm_expected_girls_pct=parse_numeric(row.get("rwm_expected_girls_pct")),
rwm_high_boys_pct=parse_numeric(row.get("rwm_high_boys_pct")),
rwm_high_girls_pct=parse_numeric(row.get("rwm_high_girls_pct")),
# Disadvantaged
rwm_expected_disadvantaged_pct=parse_numeric(
row.get("rwm_expected_disadvantaged_pct")
),
rwm_expected_non_disadvantaged_pct=parse_numeric(
row.get("rwm_expected_non_disadvantaged_pct")
),
disadvantaged_gap=parse_numeric(row.get("disadvantaged_gap")),
# 3-Year
rwm_expected_3yr_pct=parse_numeric(row.get("rwm_expected_3yr_pct")),
reading_avg_3yr=parse_numeric(row.get("reading_avg_3yr")),
maths_avg_3yr=parse_numeric(row.get("maths_avg_3yr")),
)
db.add(result)
results_created += 1
if results_created % 10000 == 0:
print(f" Created {results_created} results...")
db.flush()
print(f" Created {results_created} results")
# Commit all changes
db.commit()
print("\nMigration complete!")
def _apply_schema_alterations():
"""
Add new columns to existing tables using ALTER TABLE … ADD COLUMN IF NOT EXISTS.
Safe to run on every migration — no-ops if the column already exists.
Add entries here whenever models.py gains new columns on an existing table.
"""
alterations = [
# v4: Ofsted Report Card columns
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS framework VARCHAR(20)",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_safeguarding_met BOOLEAN",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_inclusion INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_curriculum_teaching INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_achievement INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_attendance_behaviour INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_personal_development INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_leadership_governance INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_early_years INTEGER",
"ALTER TABLE ofsted_inspections ADD COLUMN IF NOT EXISTS rc_sixth_form INTEGER",
]
from sqlalchemy import text as sa_text
with engine.connect() as conn:
for stmt in alterations:
try:
conn.execute(sa_text(stmt))
except Exception as e:
print(f" Warning: alteration skipped ({e})")
conn.commit()
def _apply_schema_drops():
"""
Drop tables retired from the schema. Idempotent (DROP … IF EXISTS), so it's
safe to run on every migration. Add entries here when a model is removed.
"""
drops = [
# v6: Ofsted Parent View feature removed
"DROP TABLE IF EXISTS marts.fact_parent_view CASCADE",
]
from sqlalchemy import text as sa_text
with engine.connect() as conn:
for stmt in drops:
try:
conn.execute(sa_text(stmt))
except Exception as e:
print(f" Warning: drop skipped ({e})")
conn.commit()
def run_full_migration(geocode: bool = False) -> bool:
"""
Run a complete migration: drop all tables and reimport from CSV.
Returns True if successful, False if no data found.
Raises exception on error.
"""
# Preserve existing geocoding so a reimport doesn't throw away coordinates
# that took a long time to compute.
geocode_cache: dict[int, tuple[float, float]] = {}
inspector = __import__("sqlalchemy").inspect(engine)
if "schools" in inspector.get_table_names():
try:
with get_db_session() as db:
rows = db.execute(
__import__("sqlalchemy").text(
"SELECT urn, latitude, longitude FROM schools "
"WHERE latitude IS NOT NULL AND longitude IS NOT NULL"
)
).fetchall()
geocode_cache = {r.urn: (r.latitude, r.longitude) for r in rows}
print(f" Saved {len(geocode_cache)} existing geocoded coordinates.")
except Exception as e:
print(f" Warning: could not save geocode cache: {e}")
# Only drop the core KS2 tables — leave supplementary tables (ofsted, census,
# finance, etc.) intact so a reimport doesn't wipe integrator-populated data.
# schema_version is NOT dropped: it persists so restarts don't re-trigger migration.
ks2_tables = ["school_results", "schools"]
print(f"Dropping core tables: {ks2_tables} ...")
inspector = __import__("sqlalchemy").inspect(engine)
existing = set(inspector.get_table_names())
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
print("Creating all tables...")
Base.metadata.create_all(bind=engine)
# ALTER existing supplementary tables to add any new columns.
# create_all() only creates missing tables; it won't add columns to tables
# that already exist from an older schema version. These statements are
# idempotent (IF NOT EXISTS) so they're safe to run on every migration.
print("Applying column additions to supplementary tables...")
_apply_schema_alterations()
print("Dropping retired tables...")
_apply_schema_drops()
print("\nLoading CSV data...")
df = load_csv_data(settings.data_dir)
if df.empty:
print("Warning: No CSV data found to migrate!")
return False
migrate_data(df, geocode=geocode, geocode_cache=geocode_cache)
return True
+45
View File
@@ -321,3 +321,48 @@ class Ks2NationalAverage(Base):
gps_high_pct = Column(Float)
gps_avg_score = Column(Float)
science_expected_pct = Column(Float)
class FactKs4Destinations(Base):
"""KS4 leavers destinations — one row per URN, year, pupil group, measure.
Long format rather than wide because pupil_group is a real third dimension.
`status` is load-bearing: 'suppressed' means DfE withheld a figure it
considered disclosive and the page must print "withheld"; 'not_applicable'
means the measure does not apply and the page must print nothing. `pupils`
is null for both, so collapsing status to a null check loses the
difference — and the categories sum to the cohort, so a consumer that
treats a withheld cell as zero republishes what DfE hid.
"""
__tablename__ = "fact_ks4_destinations"
__table_args__ = (
Index("ix_ks4_dest_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
pupil_group = Column(String(20), primary_key=True)
destination_measure = Column(String(40), primary_key=True)
cohort_pupils = Column(Integer)
pupils = Column(Integer)
percentage = Column(Float)
status = Column(String(20))
class FactKs5Destinations(Base):
"""16-18 study leavers destinations — same grain as FactKs4Destinations."""
__tablename__ = "fact_ks5_destinations"
__table_args__ = (
Index("ix_ks5_dest_urn_year", "urn", "year"),
MARTS,
)
urn = Column(Integer, primary_key=True)
year = Column(Integer, primary_key=True)
pupil_group = Column(String(20), primary_key=True)
destination_measure = Column(String(40), primary_key=True)
cohort_pupils = Column(Integer)
pupils = Column(Integer)
percentage = Column(Float)
status = Column(String(20))
+259
View File
@@ -0,0 +1,259 @@
"""Which nearby schools a detail page may offer as alternatives.
HARD FILTERS decide eligibility, and encode claims the section is not allowed
to make. A selective school is not an alternative to a non-selective one, a
special school is not comparable to a mainstream one, and a Girls school is not
an option for a Boys school's reader. They never relax, at any distance, even
where that means the section does not render at all.
DISTANCE decides the order, and nothing else does.
An earlier version ranked by intake similarity first and used distance only as
a tiebreak. That put a Catholic school 2.9 miles away above the community
school 0.3 miles down the road, and — because the row filled from the best tier
before widening — filled all six slots with faith matches while omitting every
school a parent could actually walk to. For a primary, a school that far is not
a weaker option; it is not an option. Distance is a constraint and intake is a
preference, and the ranking now says so.
Similarity survives as `shared`: what a candidate genuinely has in common with
this school, reported on its card, so a reader applies their own weighting
instead of having ours applied for them.
Pure functions over a DataFrame: no I/O, no FastAPI, no database.
"""
from __future__ import annotations
import re
import numpy as np
import pandas as pd
from .schemas import PHASE_GROUPS
# Three fit the row; the rest are behind the carousel arrows.
MAX_SCHOOLS = 6
MINIMUM = 2
# How far the section will reach, in miles, when nothing closer exists.
#
# A sanity bound rather than a target: ordering by distance already handles
# density, so a school in a dense area fills all six slots inside a mile and
# never sees this. It decides one thing — what happens where the area is
# sparse — and the answer differs by phase because catchments do. Primary
# catchments are routinely under a mile; beyond two, a primary is not a weaker
# option but not an option, and no section is the honest answer.
PRIMARY_RADIUS_MILES = 2.0
SECONDARY_RADIUS_MILES = 6.0
POST16_RADIUS_MILES = 10.0
EARTH_RADIUS_MILES = 3958.8
_SPECIAL = re.compile(r"\bspecial\b|pupil referral|alternative provision", re.I)
# Values that mean "this school has no religious character".
_NO_FAITH = {"", "none", "does not apply", "not applicable"}
def is_special_provision(school_type: str | None) -> bool:
"""Mirror of isSpecialSchool() in nextjs-app/lib/utils.ts.
Special schools carry a mainstream phase, so phase alone cannot identify
them. The two implementations must agree: a school the frontend treats as
special for benchmarking but this treats as mainstream would be dropped
from its own England comparison and then offered as a peer to a mainstream
school on the next page along.
"""
return bool(_SPECIAL.search(school_type or ""))
def is_selective(admissions_policy: str | None) -> bool:
"""Strictly selective. Unknown counts as non-selective, which is the safe
direction: it can only ever exclude a pairing, never invent one."""
return (admissions_policy or "").strip().lower() == "selective"
def faith_key(denomination: str | None) -> str:
value = (denomination or "").strip().lower()
return "" if value in _NO_FAITH else value
def faith_label(denomination: str | None) -> str:
return denomination.strip() if faith_key(denomination) else "No religious character"
def genders_compatible(a: str | None, b: str | None) -> bool:
single = {"boys", "girls"}
left, right = (a or "").strip().lower(), (b or "").strip().lower()
return not (left in single and right in single and left != right)
def is_secondary_phase(phase: str | None) -> bool:
"""Whether this phase takes the secondary side: secondary group membership,
minus all-through.
Membership is read from PHASE_GROUPS rather than tested with `"secondary" in
phase`, because that substring misses "16 plus" — GIAS phase 6, which
PHASE_GROUPS deliberately files as secondary. The substring version fails
silently rather than loudly: a sixth-form college is simply handed the
primary bucket and offered infant schools as peers.
All-through is the exception. PHASE_GROUPS lists it on both sides because it
belongs on both phases' place pages, but the detail page renders it with the
primary template, and the metric follows the template.
"""
text = (phase or "").strip().lower()
return text != "all-through" and text in PHASE_GROUPS["secondary"]
def radius_miles(phase: str | None) -> float:
"""How far this phase's section will reach when nothing closer exists."""
if (phase or "").strip().lower() == "16 plus":
return POST16_RADIUS_MILES
return SECONDARY_RADIUS_MILES if is_secondary_phase(phase) else PRIMARY_RADIUS_MILES
def _phase_group(is_secondary: bool) -> set[str]:
return PHASE_GROUPS["secondary" if is_secondary else "primary"]
def _haversine_miles(lat1: float, lon1: float, lat2, lon2):
"""Vectorised, matching the postcode search in app.py."""
lat1_r, lon1_r = np.radians(lat1), np.radians(lon1)
lat2_r, lon2_r = np.radians(lat2.astype(float)), np.radians(lon2.astype(float))
dlat, dlon = lat2_r - lat1_r, lon2_r - lon1_r
a = np.sin(dlat / 2) ** 2 + np.cos(lat1_r) * np.cos(lat2_r) * np.sin(dlon / 2) ** 2
return 2 * EARTH_RADIUS_MILES * np.arcsin(np.sqrt(a))
def _native(value):
"""NaN and numpy scalars both reach JSONResponse badly; normalise here so
the caller never has to remember to."""
if value is None:
return None
if isinstance(value, np.generic):
value = value.item()
if isinstance(value, float) and np.isnan(value):
return None
return value
def _mask(series: pd.Series, predicate) -> pd.Series:
"""A boolean mask that survives an empty frame.
`Series.apply` on an empty Series returns an empty *DataFrame*, and using
that as a mask silently drops every column — so the next column lookup
raises KeyError rather than yielding no rows. This is not hypothetical: a
special school with no special school near it empties the frame at the
provision filter, which is the ordinary case for most special schools.
"""
return pd.Series([predicate(value) for value in series], index=series.index, dtype=bool)
def _shared(subject: pd.Series, candidate: pd.Series, is_secondary: bool) -> list[str]:
"""What this candidate genuinely has in common with the subject.
Empty is a real answer, and renders no chips at all. A card claiming a
shared characteristic it does not have would be worse than a bare one —
and since these no longer affect the order, an empty list costs the school
nothing but its place in the row, which distance already decided.
"""
shared: list[str] = []
gender = str(subject.get("gender") or "").strip()
if gender and str(candidate.get("gender") or "").strip().lower() == gender.lower():
shared.append(gender)
if is_secondary:
policy = str(candidate.get("admissions_policy") or "").strip()
subject_policy = str(subject.get("admissions_policy") or "").strip()
if (
policy
and policy.lower() == subject_policy.lower()
and policy.lower() not in {"not applicable", "unknown"}
):
shared.append(policy)
if faith_key(candidate.get("religious_denomination")) == faith_key(
subject.get("religious_denomination")
):
shared.append(faith_label(candidate.get("religious_denomination")))
return shared
def select_nearby(frame: pd.DataFrame, urn: int) -> list[dict]:
"""The nearest eligible schools, closest first — at most MAX_SCHOOLS, and
none at all below MINIMUM.
The phase is read from the subject's own row rather than passed in, so a
caller cannot hand this a phase that disagrees with the data it selects
from.
"""
subject_rows = frame[frame["urn"] == urn]
if subject_rows.empty:
return []
subject = subject_rows.iloc[0]
lat, lon = _native(subject.get("latitude")), _native(subject.get("longitude"))
if lat is None or lon is None:
return []
phase = subject.get("phase")
is_secondary = is_secondary_phase(phase)
reach = radius_miles(phase)
metric_key = "attainment_8_score" if is_secondary else "rwm_expected_pct"
candidates = frame[frame["urn"] != urn].copy()
for column in ("latitude", "longitude"):
candidates = candidates[candidates[column].notna()]
if candidates.empty:
return []
# ── Hard filters ────────────────────────────────────────────────────
allowed_phases = _phase_group(is_secondary)
candidates = candidates[
candidates["phase"].fillna("").str.lower().isin(allowed_phases)
]
candidates = candidates[candidates["status"].fillna("").str.lower().str.startswith("open")]
subject_special = is_special_provision(subject.get("school_type"))
special = _mask(candidates["school_type"], is_special_provision)
candidates = candidates[special if subject_special else ~special]
subject_selective = is_selective(subject.get("admissions_policy"))
selective = _mask(candidates["admissions_policy"], is_selective)
candidates = candidates[selective if subject_selective else ~selective]
subject_gender = subject.get("gender")
candidates = candidates[
_mask(candidates["gender"], lambda g: genders_compatible(subject_gender, g))
]
if candidates.empty:
return []
candidates["distance_miles"] = _haversine_miles(
lat, lon, candidates["latitude"].values, candidates["longitude"].values
).round(1)
# ── Nearest first, and nothing else has a say ───────────────────────
within = candidates[candidates["distance_miles"] <= reach]
if len(within) < MINIMUM:
return []
selected = within.sort_values(["distance_miles", "urn"]).head(MAX_SCHOOLS)
return [
{
"urn": int(row["urn"]),
"school_name": str(row.get("school_name") or ""),
"distance_miles": float(row["distance_miles"]),
"school_type": _native(row.get("school_type")),
"age_range": _native(row.get("age_range")),
"shared": _shared(subject, row, is_secondary),
"metric_value": _native(row.get(metric_key)),
"metric_key": metric_key,
"metric_year": _native(row.get("year")),
}
for _, row in selected.iterrows()
]
+42
View File
@@ -296,6 +296,48 @@ def _locality_places(df, publishable: set[int],
return out
# Ordered authority → town/locality → outcode, widest first, because that is
# the order a breadcrumb reads. The link module re-sorts for its own purposes.
_PLACE_ORDER = {"authority": 0, "town": 1, "locality": 2, "outcode": 3}
def build_place_index(registry: dict[str, Place]) -> dict[int, tuple[Place, ...]]:
"""URN → the published places containing it, built once per registry.
The reverse of the registry, and the thing school pages link out through.
Derived from the registry rather than maintained beside it, so the two
cannot disagree about which places exist: a place below the publish
threshold is absent from the registry, so it is absent from here too, and
a link is never offered for a page that does not exist.
Built as an index rather than scanned per call because /api/schools/{urn}
is the site's highest-traffic endpoint. Scanning meant walking every place
and doing a tuple membership test against each — on the order of 10^5
comparisons per request, repeated for every school page view. One pass at
registry-build time replaces all of it with a dict lookup.
"""
grouped: dict[int, list[Place]] = {}
for place in registry.values():
for urn in place.urns:
grouped.setdefault(int(urn), []).append(place)
return {
urn: tuple(sorted(places,
key=lambda p: (_PLACE_ORDER.get(p.kind, 9), p.slug)))
for urn, places in grouped.items()
}
def places_for_urn(index: dict[int, tuple[Place, ...]], urn: int) -> tuple[Place, ...]:
"""The published places containing this school, widest first.
Empty is a real answer, not a failure: a school whose town and authority
both fall below the publish threshold has nowhere to link, and the page
renders without the module.
"""
return index.get(int(urn), ())
def build_place_registry(df) -> dict[str, Place]:
"""Every place the site publishes, keyed by "<kind>:<slug>"."""
if df.empty or "urn" not in df.columns:
+12
View File
@@ -532,6 +532,18 @@ RANKING_COLUMNS = [
"gcse_grade_91_pct",
]
# Maps user-facing phase filter values to the GIAS PhaseOfEducation values they
# include. All-through schools appear in both primary and secondary results,
# which is why this is a set per phase rather than a single string comparison.
#
# Lives here rather than in app.py because nearby_schools.py needs it too, and
# importing app from there would be a cycle.
PHASE_GROUPS: dict[str, set[str]] = {
"primary": {"primary", "middle deemed primary", "all-through"},
"secondary": {"secondary", "middle deemed secondary", "all-through", "16 plus"},
"all-through": {"all-through"},
}
# School listing columns
SCHOOL_COLUMNS = [
"urn",
+269
View File
@@ -0,0 +1,269 @@
"""The destinations serialiser's contract.
Not rendering a figure is not the same as not publishing it. This endpoint is
public and unauthenticated, so whatever the payload carries is published,
whatever the UI draws. The categories sum to the cohort and the pupil groups
sum to each other, so a lone suppressed cell is solvable by subtraction — the
serialiser adds secondary suppression to prevent it.
See docs/superpowers/specs/2026-08-28-destination-measures-design.md.
"""
from backend.data_loader import (
_destinations_block, _format_cohort_year, disclosure_invariant_holds,
)
def _row(group, measure, pupils, status, cohort=180, percentage=None, year=202223):
return {
"pupil_group": group,
"destination_measure": measure,
"pupils": pupils,
"percentage": percentage,
"status": status,
"cohort_pupils": cohort,
"year": year,
}
def test_suppressed_category_serialises_as_suppressed_with_null_pupils():
rows = [
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
_row("all", "sixth_form_college", None, "suppressed"),
]
block = _destinations_block(rows)
cats = {c["category"]: c for c in block["groups"]["all"]["categories"]}
assert cats["sixth_form_college"]["status"] == "suppressed"
assert cats["sixth_form_college"]["pupils"] is None
assert cats["sixth_form_college"]["percentage"] is None
def test_published_category_keeps_its_figures():
block = _destinations_block([
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
])
cat = block["groups"]["all"]["categories"][0]
assert cat["pupils"] == 75
assert cat["percentage"] == 41.7
assert cat["status"] == "published"
def test_only_the_latest_year_is_served():
rows = [
_row("all", "school_sixth_form", 60, "published", year=202122),
_row("all", "school_sixth_form", 75, "published", year=202223),
]
block = _destinations_block(rows)
assert block["cohort_year"] == "2022/23"
assert len(block["groups"]["all"]["categories"]) == 1
assert block["groups"]["all"]["categories"][0]["pupils"] == 75
def test_all_three_pupil_groups_are_carried():
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("disadvantaged", "school_sixth_form", 17, "published", cohort=62),
_row("other", "school_sixth_form", 58, "published", cohort=118),
]
block = _destinations_block(rows)
assert set(block["groups"]) == {"all", "disadvantaged", "other"}
assert block["groups"]["disadvantaged"]["cohort"] == 62
def test_cohort_year_is_reported_so_the_page_can_date_itself():
block = _destinations_block([_row("all", "school_sixth_form", 75, "published")])
assert block["cohort_year"] == "2022/23"
def test_format_cohort_year_handles_the_six_digit_form():
assert _format_cohort_year(202223) == "2022/23"
assert _format_cohort_year(None) is None
def test_empty_rows_yield_none_not_an_empty_shell():
assert _destinations_block([]) is None
# ── Disclosure control ──────────────────────────────────────────────────────
#
# The rendering guards in lib/destinations.ts stop a withheld figure being
# DRAWN. They do nothing about it being COMPUTED: this endpoint is public and
# unauthenticated, so whatever the payload carries is published. These tests
# are the ones that matter.
def _solve_residual(group):
"""What any caller can work out: cohort minus everything published."""
published = [c["pupils"] for c in group["categories"] if c["pupils"] is not None]
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
return group["cohort"] - sum(published), len(hidden)
def test_a_lone_suppressed_category_cannot_be_solved_for():
"""Whitley Bay High School's real 2022/23 disadvantaged group: further
education withheld, everything else published, cohort 41. Before secondary
suppression the payload gave the answer away as 41 - 23 = 18."""
rows = [
_row("disadvantaged", "school_sixth_form", 15, "published", cohort=41),
_row("disadvantaged", "sixth_form_college", 0, "published", cohort=41),
_row("disadvantaged", "further_education", None, "suppressed", cohort=41),
_row("disadvantaged", "apprenticeship", 1, "published", cohort=41),
_row("disadvantaged", "employment", 2, "published", cohort=41),
_row("disadvantaged", "not_sustained", 3, "published", cohort=41),
_row("disadvantaged", "not_captured", 2, "published", cohort=41),
]
group = _destinations_block(rows)["groups"]["disadvantaged"]
residual, hidden = _solve_residual(group)
assert hidden >= 2, "a lone suppressed cell must gain a companion"
assert residual != 18, "the withheld figure is recoverable from the payload"
def test_every_group_hides_none_or_at_least_two_categories():
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("all", "sixth_form_college", None, "suppressed"),
_row("all", "further_education", 61, "published"),
_row("all", "apprenticeship", 8, "published"),
_row("all", "employment", 6, "published"),
_row("all", "not_sustained", 5, "published"),
_row("all", "not_captured", 4, "published"),
]
group = _destinations_block(rows)["groups"]["all"]
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
assert len(hidden) >= 2
def test_a_category_hidden_in_one_group_is_hidden_in_a_second():
"""disadvantaged + other = all for every category, so a category withheld
in exactly one of the three is recoverable from the other two."""
rows = []
for measure, a, d, o in [
("school_sixth_form", 75, None, 58),
("further_education", 61, 27, 34),
("apprenticeship", 8, 4, 4),
("employment", 6, 1, 5),
("not_sustained", 5, 3, 2),
("not_captured", 4, 2, 2),
]:
rows.append(_row("all", measure, a, "published", cohort=159))
rows.append(_row("disadvantaged", measure, d,
"published" if d is not None else "suppressed", cohort=37))
rows.append(_row("other", measure, o, "published", cohort=122))
groups = _destinations_block(rows)["groups"]
measures = {c["category"] for g in groups.values() for c in g["categories"]}
assert len(measures) == 6, "the fixture's six measures must all be checked"
for measure in sorted(measures):
hidden = sum(
1 for g in groups.values() for c in g["categories"]
if c["category"] == measure and c["status"] == "suppressed"
)
# The invariant is "none, or at least two" — not "at least two".
assert hidden != 1, f"{measure} is solvable across the pupil groups"
def test_a_suppressed_cell_never_keeps_its_percentage():
"""percentage / pupils would hand back the cohort, and with it the residual."""
rows = [
_row("all", "school_sixth_form", 75, "published", percentage=41.7),
_row("all", "sixth_form_college", None, "suppressed", percentage=11.7),
_row("all", "further_education", 61, "published", percentage=33.9),
]
group = _destinations_block(rows)["groups"]["all"]
for cell in group["categories"]:
if cell["status"] != "published":
assert cell["pupils"] is None
assert cell["percentage"] is None
def test_aggregates_are_not_served():
"""An aggregate spanning exactly one suppressed component names it, and
nothing renders them today."""
rows = [
_row("all", "school_sixth_form", 75, "published"),
_row("all", "agg_sustained_all", 171, "published"),
]
group = _destinations_block(rows)["groups"]["all"]
assert [c["category"] for c in group["categories"]] == ["school_sixth_form"]
assert "aggregates" not in group
def test_a_fully_published_group_is_left_alone():
"""Secondary suppression must not cost anything where nothing is withheld —
this is the all-pupils view on every mainstream secondary."""
rows = [
_row("all", m, p, "published")
for m, p in [("school_sixth_form", 75), ("sixth_form_college", 21),
("further_education", 61), ("apprenticeship", 8),
("employment", 6), ("not_sustained", 5), ("not_captured", 4)]
]
group = _destinations_block(rows)["groups"]["all"]
assert all(c["status"] == "published" for c in group["categories"])
assert len(group["categories"]) == 7
def test_the_invariant_is_asserted_directly_not_re_derived():
"""A group with one suppressed category and nothing else to withhold."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=9),
_row("all", "sixth_form_college", None, "not_applicable", cohort=9),
_row("all", "further_education", None, "not_applicable", cohort=9),
]
block = _destinations_block(rows)
assert block is None or disclosure_invariant_holds(block["groups"])
def test_a_sparse_cohort_with_no_companion_drops_the_group():
"""Special schools and AP routinely have one suppressed category and every
other one not applicable. There is nothing left to withhold, so the group
goes — an earlier version returned here with the violation intact."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=9),
_row("all", "sixth_form_college", None, "not_applicable", cohort=9),
_row("all", "further_education", None, "not_applicable", cohort=9),
_row("all", "apprenticeship", None, "not_applicable", cohort=9),
_row("all", "employment", None, "not_applicable", cohort=9),
_row("all", "not_sustained", None, "not_applicable", cohort=9),
_row("all", "not_captured", None, "not_applicable", cohort=9),
]
block = _destinations_block(rows)
assert block is None or "all" not in block["groups"], (
"a group that cannot be made safe must not be served"
)
def test_zeros_are_not_treated_as_a_usable_companion():
"""Suppressing a zero protects nothing — the residual is unchanged. With
only zeros available the group must be dropped, not falsely 'fixed'."""
rows = [
_row("all", "school_sixth_form", None, "suppressed", cohort=5),
_row("all", "sixth_form_college", 0, "published", cohort=5),
_row("all", "further_education", 0, "published", cohort=5),
]
block = _destinations_block(rows)
if block and "all" in block["groups"]:
group = block["groups"]["all"]
published = sum(c["pupils"] for c in group["categories"]
if c["pupils"] is not None)
hidden = [c for c in group["categories"] if c["status"] == "suppressed"]
assert len(hidden) != 1, "a zero companion leaves the figure solvable"
assert group["cohort"] - published != 5
def test_masking_always_terminates_in_a_safe_state():
"""Exhaustive over every suppression pattern of a four-category group."""
from itertools import product
MEASURES = ["school_sixth_form", "sixth_form_college",
"further_education", "apprenticeship"]
for statuses in product(["published", "suppressed", "not_applicable"],
repeat=len(MEASURES)):
rows = [
_row("all", m, 3 if st == "published" else None, st, cohort=12)
for m, st in zip(MEASURES, statuses)
]
block = _destinations_block(rows)
if block is None:
continue
assert disclosure_invariant_holds(block["groups"]), (
f"invariant broken for {statuses}"
)
+383
View File
@@ -0,0 +1,383 @@
"""Selection rules for the nearby-schools section.
Hard filters encode claims the section is not allowed to make — that a
selective school is an alternative to a non-selective one, that a special
school is comparable to a mainstream one, or that a Girls school is an option
for a Boys school's reader. They decide who is eligible.
Distance decides the order, and nothing else does. An earlier version ranked by
intake similarity first, which put a Catholic school 2.9 miles away above the
community school 0.3 miles down the road — for a primary, a school that far is
not a weaker option, it is not an option. Similarity is now reported on the
card and never reorders the row.
"""
import numpy as np
import pandas as pd
from backend.nearby_schools import (
is_secondary_phase,
radius_miles,
select_nearby,
)
BASE_LAT, BASE_LON = 51.5000, -0.1000
def _row(urn, name, **overrides):
base = {
"urn": urn,
"school_name": name,
"local_authority": "Testshire",
"school_type": "Community school",
"phase": "Primary",
"age_range": "4-11",
"status": "Open",
"gender": "Mixed",
"religious_denomination": "None",
"admissions_policy": "Not applicable",
"latitude": BASE_LAT,
"longitude": BASE_LON,
"year": 202425,
"rwm_expected_pct": 70.0,
"attainment_8_score": np.nan,
}
base.update(overrides)
return base
def _frame(*rows):
return pd.DataFrame(list(rows))
def _at(miles):
"""A latitude `miles` north of BASE_LAT."""
return BASE_LAT + miles / 69.0
# ---------------------------------------------------------------------------
# Order: distance, and only distance
# ---------------------------------------------------------------------------
def test_returns_nearest_first():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Mid", latitude=_at(1.0)),
_row(100003, "Near", latitude=_at(0.4)),
_row(100004, "Far", latitude=_at(1.8)),
)
result = select_nearby(frame, 100001)
assert [s["urn"] for s in result] == [100003, 100002, 100004]
assert result[0]["distance_miles"] == 0.4
def test_a_faith_match_never_outranks_a_closer_school():
"""The reported defect. A Catholic primary surrounded by Catholic primaries
showed six of them and omitted the community school down the road."""
frame = _frame(
_row(100001, "St Jude's RC Primary", religious_denomination="Roman Catholic"),
_row(100002, "Elm Grove Primary", religious_denomination="None", latitude=_at(0.3)),
_row(100003, "Holy Cross RC", religious_denomination="Roman Catholic", latitude=_at(0.8)),
_row(100004, "Sacred Heart RC", religious_denomination="Roman Catholic", latitude=_at(1.2)),
_row(100005, "St Peter's RC", religious_denomination="Roman Catholic", latitude=_at(1.6)),
)
result = select_nearby(frame, 100001)
assert result[0]["urn"] == 100002, "the nearest school leads, whatever its intake"
assert [s["distance_miles"] for s in result] == sorted(s["distance_miles"] for s in result)
def test_the_nearest_eligible_school_is_always_shown():
"""Whatever else changes, a section titled "nearby" cannot omit the nearest
school while listing one four times further away."""
frame = _frame(
_row(100001, "Subject", gender="Boys", religious_denomination="Roman Catholic"),
_row(100002, "Nearest", gender="Mixed", religious_denomination="None", latitude=_at(0.2)),
*[
_row(100010 + n, f"Match {n}", gender="Boys",
religious_denomination="Roman Catholic", latitude=_at(0.9 + n * 0.1))
for n in range(6)
],
)
assert select_nearby(frame, 100001)[0]["urn"] == 100002
def test_caps_at_six_taking_the_nearest():
frame = _frame(
_row(100001, "Subject"),
*[_row(100010 + n, f"Peer {n}", latitude=_at(0.1 * (n + 1))) for n in range(7)],
)
result = select_nearby(frame, 100001)
assert len(result) == 6
assert 100016 not in {s["urn"] for s in result}, "the seventh-nearest is the one dropped"
def test_fewer_than_two_matches_returns_empty():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Only neighbour", latitude=_at(0.5)),
)
assert select_nearby(frame, 100001) == []
def test_excludes_the_subject_school():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
assert 100001 not in {s["urn"] for s in select_nearby(frame, 100001)}
def test_a_school_is_never_listed_twice():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
result = select_nearby(frame, 100001)
assert len(result) == len({s["urn"] for s in result})
# ---------------------------------------------------------------------------
# Reach: a sanity bound, not a target
# ---------------------------------------------------------------------------
def test_primary_does_not_reach_past_two_miles():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Just inside", latitude=_at(1.9)),
_row(100003, "Just outside", latitude=_at(2.4)),
_row(100004, "Miles away", latitude=_at(4.0)),
)
# One inside the cap is below the minimum, so nothing renders at all —
# a primary with nothing within two miles has no nearby schools.
assert select_nearby(frame, 100001) == []
def test_secondary_reaches_further_than_primary():
frame = _frame(
_row(100001, "Subject", phase="Secondary"),
_row(100002, "A", phase="Secondary", latitude=_at(3.0)),
_row(100003, "B", phase="Secondary", latitude=_at(5.5)),
)
assert {s["urn"] for s in select_nearby(frame, 100001)} == {100002, 100003}
def test_the_cap_follows_the_phase():
assert radius_miles("Primary") == 2.0
assert radius_miles("Middle deemed primary") == 2.0
assert radius_miles("All-through") == 2.0
assert radius_miles("Secondary") == 6.0
assert radius_miles("Middle deemed secondary") == 6.0
# Post-16 is the phase people travel furthest for.
assert radius_miles("16 plus") == 10.0
# ---------------------------------------------------------------------------
# Hard filters: eligibility, never order
# ---------------------------------------------------------------------------
def test_selective_never_meets_non_selective():
frame = _frame(
_row(100001, "Grammar", phase="Secondary", admissions_policy="Selective"),
_row(100002, "Comp A", phase="Secondary", admissions_policy="Non-selective", latitude=_at(0.5)),
_row(100003, "Comp B", phase="Secondary", admissions_policy="Non-selective", latitude=_at(0.6)),
)
assert select_nearby(frame, 100001) == []
assert 100001 not in {s["urn"] for s in select_nearby(frame, 100002)}
def test_special_schools_match_only_each_other():
frame = _frame(
_row(100001, "Special", school_type="Community special school"),
_row(100002, "Mainstream A", latitude=_at(0.5)),
_row(100003, "Mainstream B", latitude=_at(0.6)),
)
assert select_nearby(frame, 100001) == []
assert select_nearby(frame, 100002) == []
def test_boys_never_meets_girls():
frame = _frame(
_row(100001, "Boys School", gender="Boys"),
_row(100002, "Girls School", gender="Girls", latitude=_at(0.5)),
_row(100003, "Mixed School", gender="Mixed", latitude=_at(0.6)),
_row(100004, "Another Mixed", gender="Mixed", latitude=_at(0.7)),
)
urns = {s["urn"] for s in select_nearby(frame, 100001)}
assert 100002 not in urns
assert urns == {100003, 100004}
def test_closed_schools_and_missing_coordinates_are_dropped():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "Closed", status="Closed", latitude=_at(0.5)),
_row(100003, "No coords", latitude=np.nan, longitude=np.nan),
_row(100004, "Good A", latitude=_at(0.6)),
_row(100005, "Good B", latitude=_at(0.7)),
)
assert {s["urn"] for s in select_nearby(frame, 100001)} == {100004, 100005}
def test_all_through_is_offered_on_both_phase_sides():
frame = _frame(
_row(100001, "Primary subject", phase="Primary"),
_row(100002, "All through", phase="All-through", latitude=_at(0.5)),
_row(100003, "Primary peer", phase="Primary", latitude=_at(0.6)),
)
assert 100002 in {s["urn"] for s in select_nearby(frame, 100001)}
secondary = _frame(
_row(100010, "Secondary subject", phase="Secondary"),
_row(100002, "All through", phase="All-through", latitude=_at(0.5)),
_row(100011, "Secondary peer", phase="Secondary", latitude=_at(0.6)),
)
assert 100002 in {s["urn"] for s in select_nearby(secondary, 100010)}
def test_sixteen_plus_is_matched_against_secondary_not_primary():
"""GIAS phase 6 is "16 plus", and PHASE_GROUPS puts it in the secondary
group — a sixth-form college's peers are secondaries and other colleges,
never primary schools. A substring test for "secondary" misses it silently:
no crash, just a page offering infant schools to a sixth form."""
frame = _frame(
_row(100001, "Sixth Form College", phase="16 plus", age_range="16-19"),
_row(100002, "Nearby Secondary", phase="Secondary", latitude=_at(0.5),
attainment_8_score=52.0),
_row(100003, "Nearby College", phase="16 plus", latitude=_at(0.6)),
_row(100004, "Nearby Primary", phase="Primary", latitude=_at(0.1)),
)
result = select_nearby(frame, 100001)
urns = {s["urn"] for s in result}
assert 100004 not in urns, "a primary school is not a peer for a sixth form"
assert urns == {100002, 100003}
assert all(s["metric_key"] == "attainment_8_score" for s in result)
def test_is_secondary_phase_agrees_with_the_phase_groups_it_selects_from():
for phase in ("Secondary", "Middle deemed secondary", "16 plus"):
assert is_secondary_phase(phase) is True, phase
for phase in ("Primary", "Middle deemed primary", "Nursery", "", None):
assert is_secondary_phase(phase) is False, phase
# In PHASE_GROUPS an all-through school is on both sides, but it renders
# with the primary template, and the metric follows the phase side.
assert is_secondary_phase("All-through") is False
# ---------------------------------------------------------------------------
# What the card reports
# ---------------------------------------------------------------------------
def test_shared_lists_only_what_is_actually_shared():
frame = _frame(
_row(100001, "Subject", phase="Secondary", gender="Mixed",
religious_denomination="None", admissions_policy="Non-selective"),
_row(100002, "Full match", phase="Secondary", gender="Mixed",
religious_denomination="None", admissions_policy="Non-selective", latitude=_at(0.5)),
_row(100003, "Faith differs", phase="Secondary", gender="Mixed",
religious_denomination="Church of England", admissions_policy="Non-selective", latitude=_at(0.6)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert by_urn[100002]["shared"] == ["Mixed", "Non-selective", "No religious character"]
assert by_urn[100003]["shared"] == ["Mixed", "Non-selective"]
def test_a_shared_faith_is_named():
frame = _frame(
_row(100001, "Subject", religious_denomination="Roman Catholic"),
_row(100002, "Also RC", religious_denomination="Roman Catholic", latitude=_at(0.4)),
_row(100003, "Secular", religious_denomination="None", latitude=_at(0.5)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert "Roman Catholic" in by_urn[100002]["shared"]
assert by_urn[100003]["shared"] == ["Mixed"]
def test_shared_is_empty_when_nothing_is_shared():
frame = _frame(
_row(100001, "Subject", gender="Boys", religious_denomination="Roman Catholic"),
_row(100002, "A", gender="Mixed", religious_denomination="None", latitude=_at(0.4)),
_row(100003, "B", gender="Mixed", religious_denomination="Church of England", latitude=_at(0.5)),
)
assert all(s["shared"] == [] for s in select_nearby(frame, 100001))
def test_no_tier_is_reported_because_there_are_no_tiers():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.4)),
_row(100003, "B", latitude=_at(0.5)),
)
assert all("tier" not in s for s in select_nearby(frame, 100001))
def test_metric_follows_the_phase_side_not_the_neighbour():
frame = _frame(
_row(100001, "Subject", phase="Secondary", attainment_8_score=50.0),
_row(100002, "A", phase="Secondary", attainment_8_score=52.8, latitude=_at(0.5)),
_row(100003, "B", phase="Secondary", attainment_8_score=np.nan, latitude=_at(0.6)),
)
by_urn = {s["urn"]: s for s in select_nearby(frame, 100001)}
assert by_urn[100002]["metric_key"] == "attainment_8_score"
assert by_urn[100002]["metric_value"] == 52.8
assert by_urn[100002]["metric_year"] == 202425
assert by_urn[100003]["metric_value"] is None
def test_values_are_json_safe_native_types():
frame = _frame(
_row(100001, "Subject"),
_row(100002, "A", latitude=_at(0.5)),
_row(100003, "B", latitude=_at(0.6)),
)
for school in select_nearby(frame, 100001):
assert isinstance(school["urn"], int)
assert isinstance(school["distance_miles"], float)
assert not isinstance(school["metric_value"], np.generic)
# ---------------------------------------------------------------------------
# The endpoint
# ---------------------------------------------------------------------------
import pytest
from fastapi.testclient import TestClient
def _endpoint_frame():
return _frame(
_row(100001, "Subject Primary"),
_row(100002, "Neighbour A", latitude=_at(0.5)),
_row(100003, "Neighbour B", latitude=_at(0.6)),
)
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_latest_school_data", _endpoint_frame)
monkeypatch.setattr(app_module, "load_school_data", _endpoint_frame)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
return TestClient(app_module.app, raise_server_exceptions=False)
def test_detail_payload_carries_nearby_schools(client):
resp = client.get("/api/schools/100001")
assert resp.status_code == 200, resp.text
similar = resp.json()["nearby_schools"]
assert [s["school_name"] for s in similar] == ["Neighbour A", "Neighbour B"]
assert similar[0]["metric_key"] == "rwm_expected_pct"
def test_a_failure_in_selection_does_not_break_the_page(client, monkeypatch):
from backend import app as app_module
def _explode(*args, **kwargs):
raise ValueError("selection blew up")
monkeypatch.setattr(app_module, "select_nearby", _explode)
resp = client.get("/api/schools/100001")
assert resp.status_code == 200, resp.text
assert resp.json()["nearby_schools"] == []
+78 -1
View File
@@ -8,7 +8,8 @@ import numpy as np
import pandas as pd
import pytest
from backend.places import MIN_SCHOOLS, build_place_registry
from backend.places import (MIN_SCHOOLS, build_place_index,
build_place_registry, places_for_urn)
def _df(rows: list[dict]) -> pd.DataFrame:
@@ -418,3 +419,79 @@ def test_an_authority_still_publishes_phase_variants():
and /schools/authority/[la]/[phase] is the route that serves it."""
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Maidstone", "Kent")))
assert reg["authority:kent"].publishes_phase("primary")
# ── The reverse index: which published places contain a school ──────────────
#
# School pages link out to the location layer through this. It is the whole
# point of the index: before it, ~27k school pages linked to nothing on the
# site and stranded whatever authority they held.
def test_a_school_resolves_to_every_published_place_containing_it():
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
places = places_for_urn(build_place_index(reg), 100000)
kinds = {p.kind for p in places}
assert "town" in kinds
assert "authority" in kinds
def test_a_school_in_an_unpublished_town_still_resolves_to_its_authority():
# A town below the threshold has no page, so there is no link to offer —
# but the authority above it clears the threshold on the same schools and
# is where that reader should be sent.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS - 1, "Tinytown", "Essex")
+ _town(MIN_SCHOOLS, "Brentwood", "Essex", start=200000)
))
places = places_for_urn(build_place_index(reg), 100000)
# The town is below the threshold, so it has no page and must not be
# offered as a link. The authority above it does, and is the right target.
assert all(p.slug != "tinytown" for p in places)
assert "authority" in {p.kind for p in places}
def test_an_unknown_urn_resolves_to_nothing_rather_than_raising():
# A school page renders for any URN the API knows; the link module is not
# entitled to take the page down when it has nothing to say.
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
assert places_for_urn(build_place_index(reg), 999999) == ()
def test_the_index_is_consistent_with_the_registry_it_was_built_from():
# The invariant that matters: a link module must never offer a place whose
# page does not exist, and never omit one that does.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS, "Brentwood", "Essex")
+ _town(MIN_SCHOOLS, "Bedford", "Bedford", start=300000)
))
index = build_place_index(reg)
for key, place in reg.items():
for urn in place.urns:
assert place in places_for_urn(index, urn), (
f"{urn} is in {key} but the index does not say so")
def test_the_index_holds_no_school_the_registry_does_not():
# The reverse direction of the invariant above. An index entry for a URN
# no published place contains would put a link on a page for a place that
# does not list that school.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS, "Brentwood", "Essex")
+ _town(MIN_SCHOOLS - 1, "Tinytown", "Essex", start=400000)
))
index = build_place_index(reg)
for urn, places in index.items():
for place in places:
assert urn in place.urns
assert place.key in reg
def test_the_index_preserves_the_widest_first_order():
# The breadcrumb reads authority then town, and takes this order as given.
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
kinds = [p.kind for p in places_for_urn(build_place_index(reg), 100000)]
assert kinds.index("authority") < kinds.index("town")
+46
View File
@@ -137,3 +137,49 @@ def test_an_authority_without_a_page_is_named_but_carries_no_slug(straddling_cli
by_name = {a["name"]: a for a in body["place"]["authorities"]}
assert by_name["Essex"]["slug"] == "essex"
assert by_name["Isles Of Scilly"]["slug"] is None
def _attributed_df() -> pd.DataFrame:
"""The same town, with the four attributes the place table now shows."""
df = _schools_df()
df["age_range"] = "4-11"
df["religious_denomination"] = "Church of England"
df["nursery_provision"] = True
df["parliamentary_constituency"] = "Brentwood and Ongar"
return df
@pytest.fixture()
def attributed_client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _attributed_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _attributed_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_place_detail_carries_the_attributes_the_table_shows(attributed_client):
"""age_range and religious_denomination ride in on SCHOOL_COLUMNS.
nursery_provision and parliamentary_constituency do not, and the place
table needs all four — a column the response cannot fill is a column of
dashes on ~3,900 pages.
"""
body = attributed_client.get("/api/places/town/brentwood").json()
school = body["schools"][0]
assert school["age_range"] == "4-11"
assert school["religious_denomination"] == "Church of England"
assert school["nursery_provision"] is True
assert school["parliamentary_constituency"] == "Brentwood and Ongar"
def test_place_detail_survives_a_mart_without_the_optional_columns(client):
"""The base fixture has neither column, as an unrebuilt mart does not.
data_loader degrades those to NULL rather than failing the load, so the
endpoint must not assume they are present.
"""
res = client.get("/api/places/town/brentwood")
assert res.status_code == 200
assert "nursery_provision" not in res.json()["schools"][0]
+68
View File
@@ -0,0 +1,68 @@
"""Publication must preserve the current dataset until every replacement is ready."""
import asyncio
import pandas as pd
import pytest
from fastapi.testclient import TestClient
from backend import app as api, data_loader
from backend.tests.test_sixth_form_flag import _schools_df
@pytest.fixture
def client(monkeypatch):
old = _schools_df()
monkeypatch.setattr(data_loader, '_df_cache', old)
monkeypatch.setattr(data_loader, '_df_latest_cache', old)
monkeypatch.setattr(api, '_place_registry', {'old': 'registry'})
monkeypatch.setattr(api, '_place_index', {'old': 'index'})
monkeypatch.setattr(api, '_place_index_source', api._place_registry)
monkeypatch.setattr(api, '_sitemaps', {'old.xml': 'old sitemap'})
monkeypatch.setattr(api, '_publication_lock', asyncio.Lock())
monkeypatch.setattr(api.limiter, 'enabled', False)
api.app.dependency_overrides[api.verify_admin_api_key] = lambda: True
yield TestClient(api.app, raise_server_exceptions=False)
api.app.dependency_overrides.clear()
def state():
return (data_loader._df_cache, data_loader._df_latest_cache, api._place_registry,
api._place_index, api._place_index_source, api._sitemaps)
@pytest.mark.parametrize('failure', ['empty', 'database', 'sitemap', 'duplicate'])
def test_failed_reload_preserves_every_published_object(client, monkeypatch, failure):
before = state()
df = _schools_df()
if failure == 'empty':
df = pd.DataFrame()
if failure == 'duplicate':
df = pd.concat([df, df.iloc[:1]], ignore_index=True)
def load():
if failure == 'database':
raise RuntimeError('database unavailable')
return df
monkeypatch.setattr(api, 'load_school_data_as_dataframe', load)
if failure == 'sitemap':
monkeypatch.setattr(api, 'build_sitemaps', lambda *args: (_ for _ in ()).throw(RuntimeError('bad XML')))
response = client.post('/api/admin/reload')
assert response.status_code == 503
assert all(a is b for a, b in zip(before, state()))
def test_success_publishes_school_data_places_and_sitemaps(client, monkeypatch):
df = _schools_df()
df.loc[0, 'school_name'] = 'Replacement School'
monkeypatch.setattr(api, 'load_school_data_as_dataframe', lambda: df)
response = client.post('/api/admin/reload')
assert response.status_code == 200
assert data_loader.load_school_data() is df
assert data_loader.load_latest_school_data().iloc[0].school_name == 'Replacement School'
assert api._place_index_source is api._place_registry
assert 'old.xml' not in api._sitemaps
assert 'replacement-school' in api._sitemaps['schools-1.xml']
def test_failed_sitemap_regeneration_keeps_existing_publication(client, monkeypatch):
before = state()
monkeypatch.setattr(api, 'build_sitemaps', lambda *args: (_ for _ in ()).throw(RuntimeError('bad XML')))
assert client.post('/api/admin/regenerate-sitemap').status_code == 503
assert all(a is b for a, b in zip(before, state()))
+176
View File
@@ -56,6 +56,11 @@ def client(monkeypatch):
monkeypatch.setattr(
app_module, "get_supplementary_data", lambda db, urn: {}
)
# The place registry is a module-level cache, so without this the endpoint
# answers from whatever registry an earlier test happened to leave behind
# — and a `places == []` assertion is satisfied by a stale registry just
# as well as by this fixture's own data, which makes it prove nothing.
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
@@ -69,3 +74,174 @@ def test_nan_gias_fields_serialize_as_null(client):
assert info["capacity"] is None
assert info["total_pupils"] is None
assert info["school_name"] == "West London Performing Arts Academy"
# ── Links out to the location layer ─────────────────────────────────────────
#
# School pages carried no link into the site at all: the only anchor on the
# template pointed at the school's own website, so ~27k pages received
# whatever authority the site had and sent it off-site. `places` is what the
# link module and the breadcrumb are built from.
def test_places_is_present_even_when_the_school_belongs_to_none(client):
# This fixture's single school cannot clear any publish threshold, so the
# honest answer is an empty list. The key must still be there: a missing
# key and "no places" are different things to the page rendering it.
body = client.get("/api/schools/150275").json()
assert body["places"] == []
def test_places_names_only_pages_that_exist(monkeypatch):
from backend import app as app_module
from backend.places import MIN_SCHOOLS
def _df():
return pd.DataFrame([
{
"urn": 100000 + i,
"school_name": f"Brentwood School {i}",
"town": "Brentwood",
"local_authority": "Essex",
"postcode": "CM15 8AA",
"phase": "Primary",
"year": 202425,
"rwm_expected_pct": 60.0,
"attainment_8_score": np.nan,
"ofsted_grade": 2.0,
"ofsted_date": None,
}
for i in range(MIN_SCHOOLS)
])
monkeypatch.setattr(app_module, "load_school_data", _df)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
client = TestClient(app_module.app, raise_server_exceptions=False)
places = client.get("/api/schools/100000").json()["places"]
assert places, "a school in a published town must offer links"
by_kind = {p["kind"]: p for p in places}
assert by_kind["town"]["url"] == "/schools/brentwood"
assert by_kind["authority"]["url"] == "/schools/authority/essex"
# Every entry carries what the link text needs, and a count, so the anchor
# can say what it leads to rather than "click here".
for place in places:
assert place["name"]
assert place["count"] >= 1
assert place["url"].startswith("/schools/")
def _brentwood_df(phase: str = "Primary", n: int = None):
from backend.places import MIN_SCHOOLS
n = n if n is not None else MIN_SCHOOLS
return lambda: pd.DataFrame([
{
"urn": 100000 + i,
"school_name": f"Brentwood School {i}",
"town": "Brentwood", "local_authority": "Essex",
"postcode": "CM15 8AA", "phase": phase, "year": 202425,
"rwm_expected_pct": 60.0, "attainment_8_score": 50.0,
"ofsted_grade": 2.0, "ofsted_date": None,
}
for i in range(n)
])
def _places_for(monkeypatch, df_factory, urn: int):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", df_factory)
monkeypatch.setattr(app_module, "get_supplementary_data", lambda db, urn: {})
monkeypatch.setattr(app_module, "_place_registry", None)
client = TestClient(app_module.app, raise_server_exceptions=False)
return client.get(f"/api/schools/{urn}").json()["places"]
def test_a_place_offers_the_phase_page_this_school_appears_on(monkeypatch):
# "primary schools in brentwood" is the query the phase pages exist for,
# and ~950 of them were once reachable by nothing at all.
places = _places_for(monkeypatch, _brentwood_df("Primary"), 100000)
town = next(p for p in places if p["kind"] == "town")
assert town["phases"], "a primary school in a published primary town has a link"
assert town["phases"][0]["url"] == "/schools/brentwood/primary"
assert town["phases"][0]["count"] >= 1
def test_an_all_through_school_offers_both_phase_pages(monkeypatch):
# It genuinely appears on both, so there is no tie to break.
places = _places_for(monkeypatch, _brentwood_df("All-through"), 100000)
town = next(p for p in places if p["kind"] == "town")
assert {p["phase"] for p in town["phases"]} == {"primary", "secondary"}
def test_outcodes_never_offer_a_phase_page(monkeypatch):
# The registry gives outcodes no phase route — nobody searches "primary
# schools in SW11" — and computing them anyway once put a link to a
# nonexistent route on all 1,720 outcode pages.
places = _places_for(monkeypatch, _brentwood_df("Primary"), 100000)
outcode = next((p for p in places if p["kind"] == "outcode"), None)
if outcode is not None:
assert outcode["phases"] == []
def test_a_school_absent_from_the_phase_page_is_not_linked_to_it(monkeypatch):
# The check is URN membership in the registry's own phase list, not a
# re-derivation of the phase mapping. A secondary school must not be sent
# to a primary phase page that does not list it.
from backend.places import MIN_SCHOOLS
def df():
rows = [
{"urn": 100000 + i, "school_name": f"P{i}", "town": "Brentwood",
"local_authority": "Essex", "postcode": "CM15 8AA",
"phase": "Primary", "year": 202425, "rwm_expected_pct": 60.0,
"attainment_8_score": np.nan, "ofsted_grade": 2.0,
"ofsted_date": None}
for i in range(MIN_SCHOOLS)
]
rows.append({
"urn": 900000, "school_name": "Lone Secondary", "town": "Brentwood",
"local_authority": "Essex", "postcode": "CM15 8AA",
"phase": "Secondary", "year": 202425, "rwm_expected_pct": np.nan,
"attainment_8_score": 50.0, "ofsted_grade": 2.0, "ofsted_date": None,
})
return pd.DataFrame(rows)
places = _places_for(monkeypatch, df, 900000)
town = next(p for p in places if p["kind"] == "town")
# The town publishes a primary page, but this secondary school is not on
# it, and there are too few secondaries for a secondary page.
assert town["phases"] == []
def test_the_place_index_rebuilds_when_the_registry_is_replaced(monkeypatch):
"""The reverse index is cached; a stale one would put another dataset's
places on a school page. Invalidation is an identity check against the
registry rather than a second flag, so this asserts the check works."""
from backend import app as app_module
monkeypatch.setattr(app_module, "_place_registry", None)
monkeypatch.setattr(app_module, "_place_index", None)
monkeypatch.setattr(app_module, "_place_index_source", None)
monkeypatch.setattr(app_module, "load_school_data", _brentwood_df("Primary"))
first = app_module.get_place_index()
assert 100000 in first
# Same registry object, so the index is reused rather than rebuilt.
assert app_module.get_place_index() is first
# Drop the registry the way every test that touches place data does. The
# index must follow it, not survive it.
app_module._place_registry = None
monkeypatch.setattr(app_module, "load_school_data",
_brentwood_df("Primary", n=0))
rebuilt = app_module.get_place_index()
assert rebuilt is not first
assert 100000 not in rebuilt, "the index outlived the registry it came from"
+114
View File
@@ -0,0 +1,114 @@
from types import SimpleNamespace
import pytest
from fastapi.testclient import TestClient
from backend import app as api, data_loader
from backend.tests.test_sixth_form_flag import _schools_df
def client_for(monkeypatch, search):
client = SimpleNamespace(collections={'schools': SimpleNamespace(documents=SimpleNamespace(search=search))})
monkeypatch.setattr(data_loader, '_get_typesense_client', lambda: client)
def test_search_returns_matches_beyond_first_page(monkeypatch):
pages = []
def search(params):
pages.append(params['page'])
urns = range(100000, 100250) if params['page'] == 1 else [100999]
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in urns]}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('academy')
assert len(result) == 251
assert result[-1] == 100999
assert pages == [1, 2]
def test_search_caps_broad_queries_at_a_bounded_number_of_pages(monkeypatch):
requests = []
def search(params):
requests.append(params)
return {
'found': 10_000,
'hits': [
{'document': {'urn': 100000 + params['page'] * 1000 + i}}
for i in range(params['per_page'])
],
}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('school')
assert len(result) == data_loader.SEARCH_MAX_CANDIDATES
assert len(requests) == data_loader.SEARCH_MAX_CANDIDATES // data_loader.SEARCH_PAGE_SIZE
assert all(request['per_page'] == data_loader.SEARCH_PAGE_SIZE for request in requests)
assert requests[-1]['page'] == len(requests)
def test_search_uses_a_smaller_final_page_when_the_cap_is_not_a_page_multiple(monkeypatch):
monkeypatch.setattr(data_loader, 'SEARCH_MAX_CANDIDATES', 251)
requests = []
def search(params):
requests.append(params)
return {
'found': 10_000,
'hits': [{'document': {'urn': 100000 + len(requests) * 1000 + i}}
for i in range(params['per_page'])],
}
client_for(monkeypatch, search)
result = data_loader.search_schools_typesense('school')
assert len(result) == 251
assert [request['per_page'] for request in requests] == [250, 1]
def test_later_page_failure_does_not_return_partial_results(monkeypatch):
def search(params):
if params['page'] == 2:
raise RuntimeError('timeout')
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in range(100000, 100250)]}
client_for(monkeypatch, search)
assert data_loader.search_schools_typesense('academy') is None
def test_zero_matches_are_distinct_from_unavailable(monkeypatch):
client_for(monkeypatch, lambda _: {'found': 0, 'hits': []})
assert data_loader.search_schools_typesense('academy') == []
monkeypatch.setattr(data_loader, '_get_typesense_client', lambda: None)
assert data_loader.search_schools_typesense('academy') is None
@pytest.mark.parametrize('matches, expected', [([], []), (None, [100001])])
def test_fallback_only_on_dependency_failure(monkeypatch, matches, expected):
monkeypatch.setattr(api.limiter, 'enabled', False)
monkeypatch.setattr(api, 'load_latest_school_data', _schools_df)
monkeypatch.setattr(api, 'search_schools_typesense', lambda _: matches)
response = TestClient(api.app).get('/api/schools?search=Alpha')
assert response.status_code == 200
assert [s['urn'] for s in response.json()['schools']] == expected
def test_filtered_api_keeps_match_from_second_search_page(monkeypatch):
monkeypatch.setattr(api.limiter, 'enabled', False)
df = _schools_df()
monkeypatch.setattr(api, 'load_latest_school_data', lambda: df)
def search(params):
urns = range(200000, 200250) if params['page'] == 1 else [100001]
return {'found': 251, 'hits': [{'document': {'urn': u}} for u in urns]}
client_for(monkeypatch, search)
response = TestClient(api.app).get('/api/schools?search=Alpha&local_authority=Testshire')
assert response.status_code == 200
assert response.json()['total'] == 1
assert response.json()['schools'][0]['urn'] == 100001
def test_unavailable_dataset_is_not_a_missing_school_or_empty_search(monkeypatch):
import pandas as pd
monkeypatch.setattr(api.limiter, 'enabled', False)
monkeypatch.setattr(api, 'load_school_data', lambda: pd.DataFrame())
monkeypatch.setattr(api, 'load_latest_school_data', lambda: pd.DataFrame())
client = TestClient(api.app)
assert client.get('/api/schools/100001').status_code == 503
assert client.get('/api/schools?search=school').status_code == 503
+9 -2
View File
@@ -132,16 +132,23 @@ def test_one_query_per_table_and_latest_row_per_urn():
"FactPupilCharacteristics": [],
"FactDeprivation": [],
"FactFinance": [],
"FactKs4Destinations": [],
"FactKs5Destinations": [],
}
session = _FakeSession(rows)
out = get_supplementary_data_batch(session, [1, 2])
# Exactly one query per table — six total, regardless of two URNs.
# Exactly one query per table — eight total, regardless of two URNs.
assert sorted(session.queries) == [
"FactAdmissionDistance", "FactAdmissions", "FactDeprivation",
"FactFinance", "FactOfstedInspection", "FactPupilCharacteristics",
"FactFinance", "FactKs4Destinations", "FactKs5Destinations",
"FactOfstedInspection", "FactPupilCharacteristics",
]
# A school with no destination rows gets null, not an empty shell — the
# frontend renders the section from the block's presence.
assert out[1]["destinations"] is None
# Latest Ofsted kept per URN
assert out[1]["ofsted"]["overall_effectiveness"] == 2
assert out[2]["ofsted"]["overall_effectiveness"] == 1
-26
View File
@@ -1,26 +0,0 @@
"""
Schema versioning for database migrations.
HOW TO USE:
- Bump SCHEMA_VERSION when making changes to database models
- This triggers an automatic full data reimport on next app startup
WHEN TO BUMP:
- Adding/removing columns in models.py
- Changing column types or constraints
- Modifying CSV column mappings in schemas.py
- Any change that requires fresh data import
"""
# Current schema version - increment when models change
SCHEMA_VERSION = 6
# Changelog for documentation
SCHEMA_CHANGELOG = {
1: "Initial schema with School and SchoolResult tables",
2: "Added pupil absence fields (reading, maths, gps, writing, science)",
3: "Added supplementary data tables: ofsted, parent_view, census, admissions, sen_detail, phonics, deprivation, finance; GIAS columns on schools",
4: "Added Ofsted Report Card columns to ofsted_inspections (new framework from Nov 2025)",
5: "Apply ALTER TABLE additions for RC columns missed by create_all on existing tables",
6: "Removed the Ofsted Parent View feature: dropped fact_parent_view table and model",
}
+29 -126
View File
@@ -1,134 +1,37 @@
# SchoolCompare.co.uk - Project Context
# SchoolCompare project context
## Overview
## Maintained documentation
SchoolCompare is a web application for comparing UK primary school (KS2) performance data. It allows users to:
- Search and browse schools by name, location (postcode), or local authority
- Compare multiple schools side-by-side with charts and tables
- View school rankings by various KS2 metrics
- See historical performance trends across years
Read [README.md](README.md), [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) and
[docs/DEVELOPMENT.md](docs/DEVELOPMENT.md) for the current implementation.
[docs/LEGACY_CODE.md](docs/LEGACY_CODE.md) records obsolete paths and deliberate
compatibility code. Historical design documents are not current setup instructions.
## Architecture
## Architecture constraints
### Backend (Python/FastAPI)
- **Framework**: FastAPI with uvicorn
- **Database**: PostgreSQL with SQLAlchemy ORM
- **Data Source**: UK Government "Compare School Performance" CSV downloads
Key files:
- `backend/app.py` - Main FastAPI application, API routes
- `backend/config.py` - Configuration via pydantic-settings (env vars, .env file)
- `backend/database.py` - SQLAlchemy engine, session management
- `backend/models.py` - Database models (School, SchoolResult)
- `backend/data_loader.py` - Data queries, geocoding, legacy DataFrame compatibility
- `backend/schemas.py` - Column mappings, metric definitions, LA code mappings
### Frontend (Vanilla JS)
- Single-page application with hash-based routing
- Chart.js for data visualization
- No build step required
Key files:
- `frontend/index.html` - Main HTML structure
- `frontend/app.js` - All application logic, API calls, rendering
- `frontend/styles.css` - Styling (CSS variables, responsive design)
### Database Schema
```
schools school_results
├── id (PK) ├── id (PK)
├── urn (unique, indexed) ├── school_id (FK → schools.id)
├── school_name ├── year (indexed)
├── local_authority ├── rwm_expected_pct
├── school_type ├── reading_expected_pct
├── postcode ├── ... (all KS2 metrics)
├── latitude, longitude └── unique(school_id, year)
└── results → SchoolResult[]
```
## Configuration
Environment variables (or `.env` file):
- `DATABASE_URL` - PostgreSQL connection string (default: `postgresql://schoolcompare:schoolcompare@localhost:5432/schoolcompare`)
- `HOST`, `PORT` - Server binding (default: `0.0.0.0:80`)
- `ALLOWED_ORIGINS` - CORS origins
## Running Locally
1. Start PostgreSQL:
```bash
docker compose up -d db
```
2. Run migration to import CSV data:
```bash
python scripts/migrate_csv_to_db.py --drop
# Add --geocode to geocode postcodes (slower, adds lat/long)
```
3. Start the app:
```bash
uvicorn backend.app:app --host 0.0.0.0 --port 8000
```
## Docker Deployment
```bash
docker compose up -d
```
This starts:
- `db` - PostgreSQL 16 with persistent volume
- `app` - FastAPI application on port 80
## Data
- Source: UK Government Compare School Performance downloads
- Location: `data/` directory with year folders (e.g., `2023-2024/england_ks2final.csv`)
- The `scripts/download_data.py` can fetch data from the government website
## Key Features
- **Location Search**: Enter postcode to find nearby schools (uses postcodes.io API)
- **Multi-school Comparison**: Select multiple schools, view metrics across years
- **Rankings**: Top schools by any KS2 metric, filterable by local authority
- **Variability Analysis**: Shows standard deviation of scores across years
## API Endpoints
- `GET /api/schools` - List/search schools (supports pagination, location search)
- `GET /api/schools/{urn}` - School details with all yearly data
- `GET /api/compare?urns=123,456` - Compare multiple schools
- `GET /api/rankings` - School rankings by metric
- `GET /api/filters` - Available filter options (LAs, types, years)
- `GET /api/metrics` - Metric definitions (single source of truth)
- `GET /api/data-info` - Database stats
- Next.js serves the public UI. FastAPI serves school data from dbt-built `marts.*`.
The backend does not create school tables or import CSVs at startup.
- School coverage spans England and multiple phases, not only primary schools in
Wandsworth and Merton.
- `/api/*` belongs to the FastAPI proxy. Payload uses `/cms-api` and `/admin`.
- Payload runs inside Next.js, with its own `payload` schema and persistent media.
Keep CMS migrations independent of school-data transformations.
- Public and Payload route groups have separate root layouts. Do not introduce
`app/layout.tsx`. Keep site-wide metadata files at the `app/` root.
- Builds must succeed with `DATABASE_URL` unset. Do not call `getCachedPayload()`
at module scope or add DB-backed `generateStaticParams`.
- After changing CMS fields/editors, run `npm run generate:importmap` and commit
the generated import map. See `nextjs-app/docs/PUBLISHING.md`.
- The backend and pipeline GIAS dictionary copies are generated together; preserve
their parity. Tests enforce it.
## SDLC
Full details in `docs/DEPLOY.md`. The short version:
Follow [docs/DEPLOY.md](docs/DEPLOY.md).
- **Never push to `main` directly.** Work on a feature branch and open a PR;
branch protection requires the PR checks (typecheck, tests, builds, AI review)
to pass before merge.
- Merging to `main` deploys automatically **to staging only**: images are
built once, deployed to the staging Portainer stack, and verified by the
Playwright journeys in `e2e/`. Production is a second, manual approval:
the "Promote to Production (manual)" workflow in Gitea Actions, run after
testing the feature on staging. It refuses commits whose staging E2E gate
isn't green. Never trigger it yourself — promotion is the human's call.
- If you change user-facing behaviour, update or extend the `e2e/` journey
tests in the same PR — they gate whether staging is fit for human testing
and whether a commit is promotable.
## Recent Changes
- Added staging environment + automated staging→prod pipeline (Gitea Actions)
- Migrated from CSV file storage to PostgreSQL database
- Added location-based search using postcode geocoding
- Added local authority filter to rankings
- Improved frontend with featured schools, loading states, API caching
# Important
- Do not attempt to start a local server to test the application, it does not work
- Never push directly to `main`. Use a feature branch and a PR with passing checks.
- Merges deploy staging only. Production promotion is a separate human decision;
do not trigger the promotion workflow yourself.
- Update E2E journeys in the same PR when changing user-facing behaviour.
- Do not attempt to start a local server to test the application; use unit checks
and the configured integration environment.
+39 -2
View File
@@ -18,7 +18,14 @@
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: development
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs)
# PAYLOAD_SECRET — Payload CMS encryption secret. REQUIRED: long and
# random, and DIFFERENT from production's. Sharing
# it would let a staging session authenticate
# against production.
# AIRFLOW_ADMIN_USER — Airflow admin username (default: admin)
# AIRFLOW_ADMIN_PASSWORD — Airflow admin password. REQUIRED: the api-server
# refuses to start without it, rather than falling
# back to a generated one that changes on restart.
# STAGING_DB_IP — macvlan IP for staging Postgres (default 10.0.1.190)
# STAGING_FRONTEND_IP — macvlan IP for staging frontend (default 10.0.1.151)
@@ -86,9 +93,20 @@ services:
- FASTAPI_URL=http://backend:80/api
- TYPESENSE_URL=http://typesense:8108
- TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme}
# Payload CMS runs inside this container, in the `payload` schema of the
# staging database. Staging has its own stack, its own Postgres and its
# own admin account — never production's.
- DATABASE_URL=postgresql://${DB_USERNAME}:${DB_PASSWORD}@sc_database:5432/${DB_DATABASE_NAME}
- PAYLOAD_SECRET=${PAYLOAD_SECRET:?set PAYLOAD_SECRET in the staging Portainer stack environment}
volumes:
# Portainer prefixes volume names with the stack name, so this is
# automatically isolated from production's media.
- payload_media:/app/media
depends_on:
backend:
condition: service_healthy
sc_database:
condition: service_healthy
networks:
backend: {}
macvlan:
@@ -124,7 +142,23 @@ services:
airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:staging
container_name: sc_staging_airflow_api
command: airflow api-server --port 8080
# The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports:
- "8081:8080"
environment:
@@ -136,6 +170,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-staging-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "${AIRFLOW_ADMIN_USER:-admin}:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:?set AIRFLOW_ADMIN_PASSWORD in the Portainer stack environment}
AIRFLOW__LOGGING__BASE_LOG_FOLDER: /opt/airflow/logs
PG_HOST: sc_database
PG_PORT: "5432"
@@ -221,3 +257,4 @@ volumes:
typesense_data:
airflow_logs:
unleash_cache:
payload_media:
+39 -2
View File
@@ -9,7 +9,13 @@
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: production
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs)
# PAYLOAD_SECRET — Payload CMS encryption secret. REQUIRED: long and
# random. Changing it invalidates every admin
# session. Staging MUST use a different value.
# AIRFLOW_ADMIN_USER — Airflow admin username (default: admin)
# AIRFLOW_ADMIN_PASSWORD — Airflow admin password. REQUIRED: the api-server
# refuses to start without it, rather than falling
# back to a generated one that changes on restart.
services:
@@ -75,9 +81,21 @@ services:
- FASTAPI_URL=http://backend:80/api
- TYPESENSE_URL=http://typesense:8108
- TYPESENSE_API_KEY=${TYPESENSE_SEARCH_KEY:-changeme}
# Payload CMS runs inside this container. It reaches Postgres over the
# `backend` network and keeps its tables in the `payload` schema, so no
# pipeline operation on `public` can touch blog content.
- DATABASE_URL=postgresql://${DB_USERNAME}:${DB_PASSWORD}@sc_database:5432/${DB_DATABASE_NAME}
# Same :? form as AIRFLOW_ADMIN_PASSWORD: refuse to start rather than
# boot with an empty secret and silently accept forged sessions.
- PAYLOAD_SECRET=${PAYLOAD_SECRET:?set PAYLOAD_SECRET in the Portainer stack environment}
volumes:
# Blog images. Not reproducible from the pipeline — must be backed up.
- payload_media:/app/media
depends_on:
backend:
condition: service_healthy
sc_database:
condition: service_healthy
networks:
backend: {}
macvlan:
@@ -113,7 +131,23 @@ services:
airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:prod
container_name: schoolcompare_airflow_api
command: airflow api-server --port 8080
# The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports:
- "8080:8080"
environment:
@@ -125,6 +159,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "${AIRFLOW_ADMIN_USER:-admin}:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:?set AIRFLOW_ADMIN_PASSWORD in the Portainer stack environment}
AIRFLOW__LOGGING__BASE_LOG_FOLDER: /opt/airflow/logs
PG_HOST: sc_database
PG_PORT: "5432"
@@ -210,3 +246,4 @@ volumes:
typesense_data:
airflow_logs:
unleash_cache:
payload_media:
+19 -1
View File
@@ -105,7 +105,23 @@ services:
airflow-api-server:
image: privaterepo.sitaru.org/tudor/school_compare-pipeline:latest
container_name: schoolcompare_airflow_api
command: airflow api-server --port 8080
# The simple auth manager generates a random password on first start and
# writes it to a file, so every container restart invalidates the last one.
# Writing the file ourselves from an environment variable makes the login
# deterministic. Airflow does not generate anything when the file exists.
#
# Built with python rather than echo/printf so a password containing quotes,
# backslashes or spaces is escaped correctly by json.dumps. An unset
# AIRFLOW_ADMIN_PASSWORD raises KeyError and the container exits: falling
# back to a generated password would silently undo the point of this.
command:
- bash
- -c
- |
set -euo pipefail
mkdir -p /opt/airflow
python -c "import json, os, pathlib; pathlib.Path('/opt/airflow/simple_auth_manager_passwords.json').write_text(json.dumps({os.environ.get('AIRFLOW_ADMIN_USER', 'admin'): os.environ['AIRFLOW_ADMIN_PASSWORD']}))"
exec airflow api-server --port 8080
ports:
- "8080:8080"
environment: &airflow-env
@@ -117,6 +133,8 @@ services:
AIRFLOW__API_AUTH__JWT_SECRET: "school-compare-airflow-jwt-secret-key-long-enough-for-sha512"
AIRFLOW__API_AUTH__JWT_ISSUER: airflow
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_USERS: "admin:admin"
AIRFLOW__CORE__SIMPLE_AUTH_MANAGER_PASSWORDS_FILE: /opt/airflow/simple_auth_manager_passwords.json
AIRFLOW_ADMIN_PASSWORD: ${AIRFLOW_ADMIN_PASSWORD:-admin}
PG_HOST: db
PG_PORT: "5432"
PG_USER: schoolcompare
+125
View File
@@ -0,0 +1,125 @@
# Architecture
This describes the implementation as reviewed on 2026-09-14. It distinguishes
current behaviour from improvements still to be implemented.
## Request flow
```text
Browser → Next.js public routes
├─ /api/* proxy → FastAPI → cached DataFrames / PostgreSQL marts
│ ├─ Typesense (search and suggestions)
│ └─ postcodes.io (postcode lookup)
└─ /admin, /cms-api, /blog → Payload → payload schema + media volume
Next.js server rendering → FastAPI directly through FASTAPI_URL
```
`nextjs-app/lib/api.ts` contains typed fetch wrappers and revalidation defaults.
The proxy is `nextjs-app/app/(frontend)/api/[...path]/route.ts`. Payload uses
`/cms-api` so its routes do not collide with the FastAPI proxy. The proxy denies
`/api/flags`; server-side rendering reads flags directly from FastAPI.
## Data ownership
| Layer | Owner and role |
|---|---|
| Source data | GIAS, DfE EES, Ofsted, finance, deprivation and council admission-distance sources |
| `raw` | Singer taps and the PostgreSQL target configured in `pipeline/meltano.yml` |
| Staging/intermediate/marts | dbt models in `pipeline/transform`; marts are materialized tables |
| `marts.dim_school`, `marts.dim_location` | School identity and location, filtered to supported England establishments |
| `marts.fact_*` | Performance and supplementary datasets; coverage and years vary |
| Typesense `schools` alias | Search documents built by `pipeline/scripts/sync_typesense.py` |
| `payload` | CMS collections and migrations in `nextjs-app/`; independent of dbt |
| Media volume | Uploaded blog media; requires backup and cannot be regenerated from school datasets |
`backend/models.py` maps existing marts for reading. It does not create the school
schema. There is no startup schema-version migration or CSV reimport. Payload's
`nextjs-app/migrations/` is active and must not be confused with the removed
legacy backend migration code.
Coordinates normally come from GIAS British National Grid coordinates transformed
by PostGIS in `dim_location.sql`. `pipeline/scripts/geocode_postcodes.py` is a
manual fallback utility, not a task wired into the current school-data DAG.
Backend postcode searches also use postcodes.io; that lookup does not populate
school coordinates in the database.
## Backend boundaries
- `app.py`: routes, middleware, search filtering, sitemap/place publication and response assembly.
- `data_loader.py`: SQL loading, process-local DataFrame caches, Typesense calls,
postcode lookups, supplementary queries and benchmark calculation.
- `database.py`: synchronous SQLAlchemy engine and sessions.
- `schemas.py`: metric definitions, column mappings and display metadata; despite
its name this is not a collection of Pydantic API response models.
- `places.py` and `localities.py`: place registry and curated locality information.
- `flags.py`: Unleash-backed feature flags, disabled when no server is configured.
- `gias_codes.py` / `ofsted_codes.py`: source-code translation and display rules.
Search starts from a cached latest-row-per-school snapshot. Detail pages read
history from the full DataFrame and supplementary data from marts. Comparisons
batch supplementary queries across selected URNs. Async routes still contain
synchronous dependency calls; a fully asynchronous database layer is not present.
## Frontend boundaries
`app/(frontend)` owns the public root layout and pages. `app/(payload)` owns the
CMS root layout. Do not add a shared `app/layout.tsx`: these groups deliberately
have separate root layouts. Root metadata files remain in `app/`.
Server pages fetch initial data and pass it to client views. Client state uses
React hooks, URL search parameters and the comparison context/localStorage.
There is no SWR dependency. Leaflet maps are loaded through dynamic wrappers;
Chart.js renders performance and comparison charts.
`components/school/` contains detail sections, with section decisions and data
preparation in `lib/schoolSections.ts`. The nearby-schools section is selected in
`backend/nearby_schools.py` — hard filters decide eligibility (phase, provision,
selectivity, gender) and distance alone decides the order, capped per phase —
and served on `/api/schools/{urn}`. Its rules are presentation logic,
deliberately kept out of `marts.*` so they can be tuned by deploy rather than by
pipeline run. `lib/types.ts` contains manually maintained
API types. `payload-types.ts` and the Payload import map are generated artifacts.
## Publication and caching today
1. Airflow DAGs extract and validate source data, then run selected dbt builds.
2. Relevant DAGs rebuild Typesense and swap the `schools` alias.
3. They call `POST /api/admin/reload` with `X-API-Key`. It builds and validates
replacement DataFrames, places, reverse membership and sitemaps off the request
loop, then publishes them together. Failure returns 503 and preserves live data.
4. A separate weekly sitemap DAG can regenerate the derived publication from the
current DataFrame without clearing the live registry first.
GIAS is scheduled daily, Ofsted monthly, and annual datasets are manually
triggered. The DAG definitions are authoritative for selectors and dependencies.
Caches exist in several independent layers: backend DataFrames and registries,
backend HTTP Cache-Control/ETags, Next.js fetch/page revalidation, and browser or
shared HTTP caches where configured. Place fetches request a one-week revalidation
interval. HTTP ETags are computed after route execution, not before database work.
Typesense publication validates every import response and the final document
count before switching aliases. A session-scoped PostgreSQL advisory lock
serialises index reads/publication across DAGs. The previous collection remains
available for rollback; old unaliased collections are pruned after success.
Failed drafts are retained until a later successful cleanup, because an uncertain
alias-update response must never cause deletion of a potentially live index.
The backend snapshot swap is process-local and assumes the current single-worker
deployment. It is not an atomic transaction spanning PostgreSQL marts, Typesense
and Next.js caches. Next.js caches are not explicitly purged by the pipeline.
School search retrieves a relevance-ordered candidate prefix (currently capped at
1,000 URNs) before applying API filters. This keeps scoped searches useful while
putting a hard ceiling on Typesense round trips; only a dependency failure invokes
substring fallback, not a valid empty match set.
## Deployment references
See [DEPLOY.md](DEPLOY.md). PR checks include frontend typechecking/tests, backend
unit tests, image builds and AI review. Staging journeys run after merging.
Staging runs are serialised across builds, deployment and E2E. Build-stamped
frontend/backend identities are checked before and after journeys. Only then are
the captured image digests marked verified. Promotion resolves and validates the
complete verified image set before retagging production. See the runbook for
first-rollout requirements and remaining integration checks.
+59 -4
View File
@@ -19,8 +19,9 @@ PR checks (.gitea/workflows/pr-checks.yml)
▼
Stage pipeline (.gitea/workflows/deploy.yml) — automatic
1. build & push images → tags sha-<sha>, staging
2. staging Portainer webhook → wait for staging health
2. staging Portainer webhook → verify frontend/backend SHA + build ID
3. Playwright E2E journeys against staging ← gate before human testing
4. verify identity again; tag tested digests verified-<full-sha>
▼
Manual testing on staging (stx.schoolcompare.co.uk)
│ Actions → "Promote to Production (manual)" ← approval #2
@@ -28,14 +29,15 @@ Manual testing on staging (stx.schoolcompare.co.uk)
Promote pipeline (.gitea/workflows/promote.yml) — manual dispatch
1. resolve target sha (input, or latest main if empty)
2. REFUSE unless that commit's "E2E Journeys against Staging" status is green
3. retag sha-<sha> → :prod (same bytes — build once, promote the image)
3. resolve verified-<full-sha> digests, validate labels, retag digests → :prod
previous :prod saved as :prod-previous
4. prod Portainer webhook → wait for prod health
4. prod Portainer webhook → verify expected SHA + build ID
```
Key principle: **build once, promote the exact image**. Production pins `:prod`,
which only moves when a human runs the promote workflow — and the workflow
only accepts commits that passed the staging E2E gate. Nothing tags `:latest`
only accepts commits that passed the staging E2E gate and have a complete verified
image set. Nothing tags `:latest`
anymore.
## Branch & PR workflow
@@ -98,6 +100,12 @@ fail the E2E gate. That's the point: staging absorbs the risk.
pr-checks status checks (frontend, backend, builds, ai-review) to pass.
5. **Bootstrap staging data via Airflow** (no prod dump — staging populates
itself from source, exercising the pipeline image end-to-end):
- Set `AIRFLOW_ADMIN_PASSWORD` in the stack environment first. The
api-server refuses to start without it. Airflow's simple auth manager
otherwise generates a password on first start and writes it to a file, so
the login changes every time the container restarts; the stack writes that
file itself from this variable instead. `AIRFLOW_ADMIN_USER` defaults to
`admin`.
- Open the staging Airflow UI (`http://<host>:8081`) and trigger, in order:
`school_data_daily`, `school_data_monthly_ofsted`, then the manual-schedule
`school_data_annual_ees` and `school_data_annual_idaci`.
@@ -250,3 +258,50 @@ how long any feature is exposed to this.
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
attempted. That is the correct behaviour for local development and CI, and it
means the test suites need no flag server.
## Release identity and the P1 reliability gate
Every staging run creates a random build ID before building its three images.
Each image carries the commit and build ID as labels. Frontend/backend images
also contain a build-time JSON file; environment overrides cannot rewrite it.
`/release.json` returns both identities with `Cache-Control: no-store`. It fails
with 503 when either identity cannot be read. FastAPI's internal endpoint is
`/api/release`.
The entire staging workflow shares one concurrency group, with cancellation
disabled. This needs Gitea 1.26 or newer, where workflow concurrency is supported
([release notes](https://blog.gitea.com/release-of-1.26.0/)); the configured server
reported 1.27.3 during this change. Do not run the workflow on an older server
that ignores the concurrency key. Manual deployments outside this workflow must
also avoid changing staging during journeys.
The gate checks both identities before and after Playwright. It then validates
labels on the captured build output digests and tags them `verified-<full-sha>`.
The manual promotion script resolves all three verified tags to immutable digests
and confirms one matching commit/build ID before moving any `:prod` tag. It polls
production for that same identity using a locally saved release manifest.
A registry error can still interrupt the three tag writes; the Portainer webhook
only runs after successful promotion, and rerunning promotion resolves the full
verified set again. There is no cross-registry atomic tag transaction.
**First rollout:** old green commits without verified tags/build identities are
not promotable through this gate. Build and test a commit containing the new
workflow first. The release route must be reachable through the configured
`STAGING_BASE_URL`/`PROD_BASE_URL`; it deliberately avoids the public staging
`/api` proxy limitation. No new deployment secret is required.
`scripts/ci/release.py` implements identity polling and digest verification.
The poller identifies itself as `SchoolCompare-Release-Check/1.0`: the public
staging proxy has returned HTTP 403 to Python's default urllib user agent even
while the release endpoint was healthy. It logs changes in HTTP/connection
failures or observed release identities, and includes the last observation in
the timeout error. If verification fails, use that observation to distinguish
proxy rejection (403), an unavailable release endpoint (503), and containers
still reporting an older SHA/build ID. Check the configured base URL from the
CI runner; a successful request from another machine does not establish runner
connectivity. Do not bypass identity verification to unblock a deployment.
Its mocked tests run in PR checks alongside backend and index-publication tests.
The new Playwright journeys also check deployed identity and stale pagination.
Local unit checks do not validate registry credentials, Portainer behaviour,
proxy routing or a deployed image; those require the staging run. Production
promotion remains a separate human action.
+94
View File
@@ -0,0 +1,94 @@
# Development and validation
## Prerequisites and environment boundaries
Use a feature branch. The deployed stack is the integration environment; do not
assume a local server can run from a fresh checkout. This cleanup did not start
local servers or provision databases. Unit tests use fixtures and mocks.
The current versions are not yet aligned:
| Component | Container | PR checks |
|---|---|---|
| Backend | Python 3.11 | Python 3.12 |
| Frontend | Node 24 | Node 22 |
| Pipeline | Python 3.13 | Pipeline image build |
Use the component's container version when reproducing deployment behaviour.
The backend dependency pins predate Python 3.14; do not assume the system Python
can install or run them. Version alignment is a separate maintenance task.
## Frontend checks
```sh
cd nextjs-app
npm ci
npm run typecheck
npm test -- --runInBand
```
`npm run build` is the production build check. There is no `lint` script.
Tests live in `__tests__/` and use Jest/React Testing Library. These checks do not
prove that live PostgreSQL queries, Typesense or a deployed proxy work.
The frontend `.env.example` documents runtime variables. Browser traffic normally
uses `/api`; `FASTAPI_URL` is an absolute server-side URL ending in `/api`.
Payload additionally needs `DATABASE_URL` and `PAYLOAD_SECRET` when used at runtime.
Never commit credentials or real `.env` files.
## Backend checks
From the repository root, using an available Python 3.11 or 3.12 interpreter:
```sh
python3.11 -m venv /tmp/schoolcompare-backend-venv
/tmp/schoolcompare-backend-venv/bin/python -m pip install -r requirements.txt pytest 'httpx<0.28' pyyaml
/tmp/schoolcompare-backend-venv/bin/python -m pytest backend/tests pipeline/tests scripts/ci/tests -q
```
Substitute `python3.12` if matching PR CI. The test dependencies above match the
current workflow; they are not yet captured in a dedicated development lockfile.
Backend configuration is defined in `backend/config.py`; `.env.example` documents
commonly used values. `ALLOWED_ORIGINS` uses a JSON array, not a comma-separated string.
## Data and pipeline work
The app needs populated `marts.*` tables. A new Postgres instance alone is not a
working school-data environment. Use the existing managed pipeline or an approved
snapshot; the removed CSV importer cannot build the current schema.
The pipeline container includes Meltano, dbt/Postgres, Airflow and the custom taps.
Airflow commands/selectors live in `pipeline/dags/`. Schema tests live in
`pipeline/transform/tests/` and model YAML files. Run the relevant `dbt build`
selector in an isolated data environment for model changes; it writes tables and
is not a read-only smoke test. Prefer `python -m dbt.cli.main` as the DAGs do.
GIAS dictionaries are generated together by
`pipeline/scripts/generate_gias_codes.py`. The backend and pipeline copies are
intentional; `backend/tests/test_gias_codes.py` checks that they stay identical.
For Payload collection/editor changes, run `npm run generate:importmap` in
`nextjs-app/` and include the generated map. Preserve CMS migrations and the
separate `payload` schema. See [publishing](../nextjs-app/docs/PUBLISHING.md).
## End-to-end checks
Against an existing, authorised test environment:
```sh
cd e2e
npm ci
npx playwright install chromium
BASE_URL=https://your-test-environment.example npx playwright test
```
The suite does not start a web server. CI installs Chromium with system dependencies
and runs against staging. Use the configured staging target: `docs/DEPLOY.md`
records the public staging proxy limitation. User-visible behaviour changes should
update the corresponding journeys.
## Before requesting review
Run checks relevant to the change, inspect `git diff --check`, and report checks
that could not run. Do not publish or promote as part of local validation.
[DEPLOY.md](DEPLOY.md) documents the PR and human promotion gates.
+89
View File
@@ -0,0 +1,89 @@
# Legacy and unused-code inventory
Reviewed 2026-09-14. This inventory records source evidence, not production usage
telemetry. A command with no repository caller may still be run manually or from
an external scheduler. Historical specs and prototypes are not runtime imports.
## Method and scope
Searched backend imports, tests, CLI scripts, Airflow DAGs, Meltano configuration,
Gitea workflows, Dockerfiles and documentation. For frontend candidates, inspected
TypeScript imports, re-exports, literal dynamic imports and `require` calls,
resolving relative and `@/` paths while excluding tests, dependencies and build
output. Checked candidates again with text searches including tests.
Next.js route files, generated Payload import-map entries and plugin discovery
are entry points even without ordinary imports. This is why a zero-import count
alone is not sufficient grounds for deletion. Computed imports and external
operators are outside this static audit.
## Removed in this cleanup
These names are recorded for Git-history lookup; they are no longer file links.
| Removed path or symbol | Evidence and replacement |
|---|---|
| `backend/migration.py` | Imported `School` and `SchoolResult`, which no longer exist in `backend/models.py`. Only the legacy CSV CLI imported it. Current tables are built by dbt. |
| `backend/version.py` | Only the legacy importer consumed `SCHEMA_VERSION`. FastAPI lifespan does not perform version-triggered imports. This is unrelated to active Payload migrations. |
| `scripts/migrate_csv_to_db.py` | Imported removed `init_db`/`set_db_schema_version` helpers and the obsolete models indirectly. No runtime, DAG or workflow calls it. Use the managed pipeline for current marts. |
| `scripts/geocode_schools.py` | Imported the removed `School` ORM model. No pipeline/workflow calls it. Coordinates now come from GIAS/PostGIS; a separate mart-aware manual utility remains under `pipeline/scripts/`. |
| `backend.data_loader.haversine_distance` | No callers. Search uses its inline vectorised NumPy calculation. |
| `nextjs-app/lib/api.ts: fetcher` | No callers; SWR is not installed. Application fetches use the named API wrappers. |
| `nextjs-app/lib/api.ts: kmToMiles` | No callers. `calculateDistance` remains because `CutoffMapPanel` uses it. |
The removed command files could not import successfully against the current
backend. This cleanup does not run replacements, migrate data or modify databases.
Their previous implementations remain recoverable from Git history.
## Unused candidates retained for a separate cleanup
| Candidate | Evidence | Recommended next step |
|---|---|---|
| `nextjs-app/components/LoadingSkeleton.tsx` and its CSS | No application or test imports found. | Remove together after confirming no planned use. |
| `nextjs-app/components/Pagination.tsx` and its CSS | No application or test imports found; HomeView implements load-more behaviour. | Remove as a pair if numbered pagination will not return. |
| `nextjs-app/components/SchoolCard.tsx` and its CSS | Imported by its own tests, not application code. HomeView uses SchoolRow/SecondarySchoolRow. | Decide whether to retire the card design; if removed, remove its dedicated tests as well. Passing tests do not establish runtime use. |
| `backend/database.py: get_db`, `get_db_session` | No remaining callers after removing the importer. Current code creates SessionLocal directly. | Either adopt these helpers during session-lifecycle cleanup or remove them; do not rewrite active sessions in a documentation change. |
| `backend/schemas.py: COLUMN_MAPPINGS`, `NULL_VALUES`, `LA_CODE_TO_NAME` | No remaining Python consumers found after importer removal. Other constants in this module are active. | Remove individual constants after checking external data utilities; retain the module. |
| `backend/config.py: data_dir`, `max_page_size`, `rate_limit_burst` | No active consumers found. `default_page_size` appears only in a branch that expects None, although the route supplies a concrete default. | Reconcile settings with route validation in a focused API change. |
## Legacy/manual paths requiring operational verification
| Path | Status and reason to retain for now |
|---|---|
| FastAPI `/`, `/compare`, `/rankings`, `/favicon.svg`, `/robots.txt`, and conditional `/static` | Old frontend-serving routes reference a `frontend/` directory absent from the checkout and backend image. Next.js owns these public surfaces. Removal changes externally callable routes, so first check proxy/operator usage and define replacement responses. |
| `scripts/fetch_real_data.py`, `scripts/download_data.py` | Historical standalone CSV utilities. The fetch script targets Wandsworth/Merton; neither is wired into the managed pipeline. Marked historical, retained pending confirmation of manual use. |
| `pipeline/scripts/geocode_postcodes.py` | Mart-aware postcode fallback, not called by the current DAGs. Do not confuse it with the removed legacy ORM geocoder. Verify the target schema before manual use. |
| `docker-compose.yml` | Uses unpublished `:latest` release tags and lacks frontend Payload DB/secret/media configuration. Retained as an old development topology, not recommended onboarding. |
| `nextjs-app/docker-compose.yml` | Standalone legacy recipe with old backend port assumptions and no CMS persistence setup. Retained until its consumers are checked. |
| `MIGRATION_SUMMARY.md`, `docs/superpowers/`, `mockups/` | Historical designs and prototypes. Retain as history; do not follow as current deployment instructions. |
| `scripts/sql/drop_fact_parent_view.sql` | One-off maintenance SQL. Not an application entry point; repository call-site searches cannot establish whether it is still needed operationally. |
## Active code that can look obsolete
- `backend/data_loader.py` older-mart query fallbacks are covered by backend tests
and support databases at different migration stages. Remove only after verifying
the deployed schemas in every supported environment.
- `backend/gias_codes.py` and `pipeline/scripts/gias_codes.py` are intentionally
generated copies for separate runtime images. Their parity is tested.
- `nextjs-app/migrations/`, `payload-types.ts` and the Payload import map are active
CMS artifacts, not remnants of the removed school importer.
- `get_available_years`, `get_available_local_authorities` and `get_schools_count`
in `data_loader.py` are called through `get_data_info`, which serves the backend
data-info endpoint. They are not dead functions.
- `get_supplementary_data` is an intentional single-school wrapper around the
batch implementation.
- `pipeline/transform` models named `legacy` can be active data sources: annual
DAG selectors explicitly include legacy KS2/KS4 lineage. Names alone do not
establish obsolescence.
## Suggested next passes
1. Decide the fate of the three unused UI components and remove paired assets/tests.
2. Consolidate backend session usage and remove abandoned settings/constants.
3. Verify external consumers, then retire static-serving API routes and old compose recipes.
4. Audit manual data utilities with pipeline operators before deleting them.
5. Revisit compatibility fallbacks only after documenting supported schema versions.
Validation for this cleanup should include frontend typechecking/tests, Python
syntax checks, reference searches and documentation link checks. Live database,
external scheduler and deployed route usage require separate integration evidence.
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,399 @@
# Destination Measures — Design
**Date:** 2026-08-28
**Status:** awaiting review
**Scope:** secondary school detail pages only
## Goal
Say what happened to a school's leavers after they left. Two sections on the
secondary template:
- **After Year 11** — every secondary, from the KS4 destination measures
- **After the sixth form** — sixth-form schools only, from the 16-18 measures
This replaces the "Post-16 destination data coming soon" placeholder standing in
`nextjs-app/components/school/SecondaryAdmissionsSection.tsx:117` since the exam
phase taxonomy work, and fills the `ks5_destinations_pct` slot specified but
never built in `2026-07-07-exam-phase-taxonomy-design.md:201`.
Mockup, with all three data states live:
<https://claude.ai/code/artifact/5be149d6-252f-473c-9a4f-4c36b05161b0>
## The finding that shapes everything
**Suppression is per cell, and the cells sum to the cohort.**
DfE withholds a figure it considers disclosive by writing `c`. It does this at
the level of an individual destination category, not the whole school, and it
publishes the cohort total alongside. The categories form a clean partition. So
where exactly one category is suppressed, subtracting the published ones from the
cohort recovers it exactly.
Verified against three real schools in the 2022/23 file:
| School | URN | Withheld | Recovers to |
|---|---|---|---|
| North East Futures UTC | 145900 | School sixth form | **3 pupils** |
| Whitley Bay High School | 108638 | Further education | **18 pupils** |
| St Matthew's RC High School | 148389 | School sixth form | **4 pupils** |
Those are the precise numbers the `c` exists to hide, and in a random 400-school
sample **22% of mainstream secondaries** have exactly one suppressed category in
their disadvantaged group. This is the normal case, not an edge case.
Three rules follow, and everything else in this document is downstream of them.
**R1 — Never *publish* enough to derive a remainder.**
An earlier draft of this rule said "never *render* a derived remainder", and
that was the defect code review caught in PR #137. Not drawing a number does
nothing to stop it being computed: `GET /api/schools/{urn}` is public and
unauthenticated, so anything in the payload is published whatever the UI
chooses to draw. The rendering guards shipped; the payload still carried the
cohort and every published category, and `cohort - sum(published)` returned
Whitley Bay's withheld figure exactly.
The rule is therefore about the serialiser, and the UI guards are a second line
of defence behind it. Two identities have to be closed:
- within a pupil group the categories sum to the cohort, so a group with
exactly **one** suppressed category gives it away;
- across groups, disadvantaged + other = all for every category, so a category
suppressed in exactly **one** of the three gives itself away.
`_mask_for_disclosure` applies DfE's own answer — secondary suppression —
withholding a companion cell until every row and every column hides either none
or at least two. It iterates, because each new suppression can break the other
identity, and terminates because cells are only ever added.
The companion must carry pupils. Suppressing a zero looks like secondary
suppression and protects nothing: the residual still equals the original
withheld figure.
Where no companion can do the job — a sparse cohort whose every other category
is `not_applicable`, routine in special schools and alternative provision — the
pupil group is **dropped from the payload entirely**. A first version simply
returned at that point with the violation intact and no signal, which review
caught: a disclosure-control pass that fails silently is worse than none,
because everything downstream trusts it. The function now cannot terminate
except in a state where `disclosure_invariant_holds()` is true, and an
exhaustive test sweeps all 81 suppression patterns of a four-category group to
prove it.
Measured cost on the 400-school sample: the all-pupils bar survives on **94%**
of mainstream secondaries rather than 100%. That is the price of not
republishing what DfE withheld.
**R2 — Never aggregate across a suppression boundary.** Summing published
components to fill a gap is R1 with extra steps.
DfE's own aggregates (`Sustained education destination`, `Sustained education,
employment & apprenticeships`) are ingested but **not served**. An aggregate
spanning exactly one suppressed component names it, and nothing renders them
today — an unused field that leaks is not a trade-off worth carrying. They can
be re-added with their own guard if the fallback ladder is ever built.
**R3 — The three pupil groups are one disclosure surface, not three.**
Disadvantaged and Not-known-to-be-disadvantaged partition All pupils, so
rendering any *two* of them recovers the third. Where a category is suppressed in
the disadvantaged group, it must therefore also be withheld from **all other
pupils** — the all-pupils view is the primary one and keeps it.
This costs almost nothing, because DfE already applies the same masking: across
the sample, 493 of 498 suppressed disadvantaged cells were suppressed in the
other group too. The mart enforces the remaining 5, which fell on 2 schools of
262. **The all-pupils bar is unaffected** — masking the whole page wherever the
disadvantaged group is thin would remove the bar from 80% of schools, and is not
what this rule says.
R1 and R2 both hold within a group and still leak across the switch, which is why
R3 is stated separately.
### The convention that would break this quietly
`macros/safe_numeric.sql` coerces every EES sentinel — `z`, `c`, `x`, `q`, `u` —
to `NULL`, deliberately and correctly for attainment, where "suppressed" and "no
data" are equally unrenderable. Here they are not the same thing: one must print
*withheld*, the other must print nothing at all, and the difference is what keeps
R1 enforceable.
**`safe_numeric` must not be used on destination counts.** The staging model
keeps the sentinel in a companion status column. This is the single most likely
way for this feature to regress into a disclosure, so it gets its own dbt test.
## What is actually available
Measured against the EES public API (open, no key). Both datasets carry
`geographicLevel: School` with `urn` on every location option, so the join to
`dim_school` is direct.
| | KS4 | 16-18 |
|---|---|---|
| Dataset id | `019d4f41-22d1-71b2-a1a7-f3b91026815b` | `019d4e73-6440-7523-b60c-bfab1ad4a30d` |
| Rows | 1,871,739 | 3,862,658 |
| Institutions | 4,946 | 3,065 |
| Time periods | 2009/10–2022/23 | 2016/17–2022/23 |
**Destination categories (KS4).** School sixth form · Sixth form college ·
Further education · Other education destination · Sustained apprenticeships (with
level breakdown) · Sustained employment destination · Not recorded as a sustained
destination · Activity not captured. Plus the aggregates `Sustained education
destination` and `Sustained education, employment & apprenticeships`.
**16-18 adds** UK higher education institution and FE split by level, which is
what makes the post-16 section worth having.
**Breakdowns.** `Disadvantage Status` gives Disadvantaged / Not known to be
disadvantaged / Total — exactly the three-way switch. Sex, ethnicity, FSM status,
prior attainment and SEN provision also travel in the same table; we ingest none
of them.
**Indicators.** Both counts and percentages, plus the cohort size. Bar widths use
the counts — the published percentages do not sum to 100.
### Coverage, and what degrades
Random 400-school sample, 2022/23, mainstream secondaries (n=262):
| View | As published by DfE | After R1–R3 masking | Consequence |
|---|---|---|---|
| All pupils, all categories | 100% | **94%** | Bar works nearly everywhere |
| Disadvantaged, headline rate | 95% | 95% | Gap panel works |
| Disadvantaged, three grouped cards | 68% | 68% | Degrades card by card |
| Disadvantaged, all six categories | 20% | **20%** | Bar unusable for this group |
The middle column is what the site actually serves. Masking costs the
all-pupils bar on 6% of mainstream secondaries — those are schools where a
category was suppressed in exactly one pupil group and no non-zero companion
existed below the all-pupils row.
Special schools and alternative provision are far worse: 13% and 41% respectively
have the whole cohort suppressed even for all pupils. The empty state is
load-bearing, not defensive.
## The display
Question-led. Three cards over one bar, with the cards acting as a lens on the
bar rather than a summary beside it — hovering a card dims the bar, table and
England reference to the categories that card is built from. The full mockup is
linked above; what matters for implementation:
**The headline is not the sustained rate.** That figure sits between 92% and 97%
for nearly every school in England. The mix is what varies, so the mix leads.
**The grouping is ours, not DfE's.** "Academic route" = school sixth form +
sixth-form college; "College" = FE and other colleges; "Work" = apprenticeship +
employment. This is the most arguable thing on the page, so it lives in one place
in `lib/destinations.ts`, is explained in a tooltip, and is reversible in one
edit.
**The absence is hatched neutral, never a colour.** "Activity not captured" means
no record in the sources DfE holds — it includes independent schools, moving
abroad and private training. Colouring it as a bad outcome would be a factual
error rendered in CSS. The hatch also fixes a real contrast problem: neutral
against the employment blue failed CVD separation at ΔE 7.6, and texture is the
secondary encoding that rescues it. Every other adjacent pair clears ΔE 10.9
under protanopia.
**Colour tokens.** Education is one hue in three steps (school-like to
college-like); apprenticeship and employment are separate hues. Six new tokens in
`globals.css`, defined in both themes, per the existing token discipline.
**The disadvantage split rides the same control.** One visualisation serving
three cohorts, with the England reference repointing to the matching national
group. The gap statement stays visible below the bar whatever is selected,
because a gap nobody clicks on is a gap nobody sees.
## Data model
### Extraction
A new `tap-uk-ees-destinations` extractor, separate from `tap-uk-ees`. The
existing tap downloads a release ZIP and reads a CSV inside it; the destinations
files are far larger than we need and the query API filters server-side, so this
one POSTs to `/v1/data-sets/{id}/query` and pages through results.
With every dimension pinned — destination measures, disadvantage status, sex
Total, characteristic topic Total — one year returns **252,610 rows** across all
geographic levels. Three school-level years is comfortably tractable.
Pinning is mandatory, not an optimisation: leaving the characteristic dimensions
unconstrained returned 45 rows where 9 were wanted, because every breakdown
shares one table.
The tap emits the raw value as text. **It does not coerce `c`.**
### Staging
`stg_ees_ks4_destinations` / `stg_ees_ks5_destinations`. Each raw value becomes
two columns:
```sql
case when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then raw::numeric end as pupils,
case
when raw ~ '^-?[0-9]+(\.[0-9]+)?$' then 'published'
when lower(trim(raw)) = 'c' then 'suppressed'
else 'not_applicable'
end as status
```
### Marts
`fact_ks4_destinations` and `fact_ks5_destinations`, **long format**:
```
urn, year, pupil_group, destination_category, cohort_pupils, pupils, percentage, status
```
This departs from the wide house pattern (`fact_ks4_performance` and friends) on
purpose. `pupil_group` is a genuine third dimension; going wide would need three
sets of every column, and R2 is far easier to test on rows than on columns.
Roughly 8 categories × 3 groups × 4,946 schools × 3 years ≈ 356k rows.
`fact_destination_national` carries the same grain for England, so the page's
England reference repoints with the switch.
### dbt tests
- `assert_destinations_no_derived_remainder` — for every (urn, year,
pupil_group) with exactly one suppressed category, assert no aggregate row
exists that would let the residual be recovered. **This is the R1 guard.**
- `assert_destinations_group_masking` — for every (urn, year, category), if the
disadvantaged group carries `suppressed`, so does the other-pupils group.
**This is the R3 guard**, applied in the mart so no consumer can reach an
unmasked combination.
- `assert_destination_status_null_agreement` — `pupils is null` wherever
`status != 'published'`, and never null where it is.
- `assert_destinations_join_dim_school` — no orphaned URNs, matching the
existing `assert_no_orphaned_facts` pattern.
## API
`GET /api/schools/{urn}` gains a `destinations` block:
```json
{
"ks4": {
"cohort_year": "2022/23",
"published": "2026-04",
"groups": {
"all": { "cohort": 180, "categories": [ … ], "aggregates": { … } },
"disadvantaged": { … },
"other": { … }
}
},
"ks5": { … }
}
```
Each category carries `pupils`, `percentage` and `status`. **The serialiser never
emits a computed remainder**, and a backend test asserts that a group containing a
suppressed category serialises no total that closes the gap.
`null` for the whole block where nothing is published — the frontend renders the
empty state from its absence, not from a sentinel.
## Frontend
| File | Kind | Job |
|---|---|---|
| `lib/destinations.ts` | pure | Category list, the academic/college/work grouping, `canAggregate()` enforcing R2, percentage derivation from counts |
| `components/school/DestinationsSection.tsx` | server | Section shell, renders **all pupils** into the HTML |
| `components/school/DestinationsView.tsx` | client | Cohort switch, card↔bar linkage |
| `components/school/Post16DestinationsSection.tsx` | server | Year 13 section, sixth-form schools only |
| `app/globals.css` | tokens | Six destination colours, both themes |
Server-first matches the directory's existing discipline — every component in
`components/school/` is a server component except `AdmissionsViewToggle`, which
is the precedent this follows. All-pupils figures are in the HTML for crawlers
and for no-JS; only the switch and the hover linkage need the client.
`lib/schoolSections.ts` gains `hasKs4Destinations` / `hasKs5Destinations` flags
and the nav items, following the existing `computeSchoolFlags` pattern.
**Placement** on the secondary template: GCSE results → After Year 11 → After the
sixth form → admissions. Destinations follow attainment because they answer "and
then what happened".
**Dating.** The latest destination year is 2022/23, published April 2026, while
the site's newest KS4 year is 2024/25. The section header states its own cohort
year, or it reads as stale data next to the GCSE section above it.
## Edge states
| State | Frequency | Behaviour |
|---|---|---|
| Whole cohort suppressed | 13% of special, 41% of AP | Section renders the explanation, no chart |
| Some categories withheld | 80% of disadvantaged views | Cards degrade individually; **no bar**; table marks withheld rows |
| Disadvantaged group suppressed entirely | 5% | Switch drops to two options, gap panel not rendered |
| No sixth form | — | Post-16 section not rendered at all — absence is correct, a "no data" placeholder would imply something is missing |
| School too new | — | "First figures expected in 2026", not a bare no |
## Testing
Per CLAUDE.md, user-facing behaviour extends `e2e/` in the same PR.
**Unit** — `lib/destinations.ts` is where R1 and R2 live, so it carries the
heaviest tests: `canAggregate()` refuses a group containing one suppressed cell,
allows one spanning two, and the bar builder refuses to emit segments for any
group with suppression. These are the tests that must fail loudly if someone
later "fixes" a gap in the chart.
**dbt** — the three tests above.
**Backend** — the serialiser emits no closing total for a partially suppressed
group.
**E2E** — a school with full data renders three cards and a bar; a school with a
partially suppressed disadvantaged group renders the withheld state and **no bar
element**; a suppressed school renders the explanation; a school with no sixth
form renders no post-16 section.
Note the staging caveat: mart changes are inert until the Airflow pipeline runs,
and the staging E2E gate runs post-merge.
## Out of scope
- **Compare view and rankings.** The long mart shape supports both; neither is
built here. Flagged because "% to a school sixth form" is a plausible rankings
metric and the mart shape should not have to change to allow it.
- **Ethnicity, sex, SEN and prior-attainment breakdowns.** Available in the same
file, ingested deliberately not at all — each is a separate editorial decision
about what a school page should assert.
- **Longer term destinations** (3 and 5 years out) and **Progression to higher
education** — separate publications, worth a later look for sixth forms.
- **Primary schools.** No KS2 destination measures publication exists; DfE
tracking starts at KS4. Naming the secondaries a primary's leavers go to needs
the National Pupil Database, which is not publishable at that grain.
## Risks
**A later change reintroduces the disclosure.** The likeliest routes are
applying `safe_numeric` to a destination column for consistency, adding a
`coalesce` in a mart, or — as happened in review — enforcing a disclosure rule
at the rendering layer instead of the publishing layer. Mitigation is the dbt
tests plus `backend/tests/test_destinations_api.py`, which reconstructs the
residual the way an attacker would and asserts it no longer resolves.
**The two-year lag reads as staleness.** Mitigated by dating the cohort in the
section header rather than only in a tooltip.
**Sixth-form retention will be misread.** "41% went to a school sixth form" says
nothing about *which* school. The published file reports destination type, never
destination institution. Copy must never imply "stayed on here", and the tooltip
should say so.
**Section length.** The secondary template is already long and this adds two
sections. If it becomes a problem the post-16 section is the one to collapse
behind a disclosure, not the Year 11 one.
## Open questions
1. Is the disadvantage split its own section or a sub-block inside the
destinations section? Modelled as a sub-block; it is the most differentiating
figure on the page and the most easily misread on a small cohort.
2. Do we ingest the apprenticeship level breakdown (intermediate / advanced /
higher) now, or collapse to one apprenticeship figure and revisit? Collapsed
in this design.
@@ -0,0 +1,368 @@
# Giving schoolcompare a human author: an About page and a blog
**Date:** 2026-09-02
**Status:** Design — awaiting review
**Scope:** A named author for the site, an `/about` page, and a Payload-CMS-backed
blog at `/blog`.
## Why
The site reads as synthetic. Not because of its tone, but because of three
specific absences:
1. **Nobody is accountable for the numbers.** There is no author, no statement
of why the site exists, and no one who can be wrong. The only human trace on
the entire site is `contact@schoolcompare.co.uk` in the footer.
2. **No visible judgement.** Every figure is presented as though it fell out of
a machine. Hundreds of editorial decisions went into this codebase — which
metrics to show, when a benchmark is invalid, what to suppress — and not one
of them is visible to a reader. `isSpecialSchool()` silently drops the
England comparison for special schools and PRUs because that comparison is
meaningless; nowhere does the site *say* so.
3. **The voice is institutional third person.** "schoolcompare brings it all
into one place." "Built for parents, governors, journalists." That is
brochure register, and it is precisely the register that machine-generated
content defaults to.
There is a second, independent reason. The SEO programme
(`2026-08-20-seo-programme-design.md`) defines eight workstreams and none of
them address E-E-A-T or authorship. School performance data is YMYL territory;
an anonymous site republishing DfE figures has no authorship signal at all. This
work fills that hole, and the blog gives W6 (explainer content) somewhere to
live.
### The failure mode to avoid
The standard fix — a stock photo and "Hi, I'm Tudor, and I'm passionate about
education!" — reads as *more* synthetic than the current coldness. Manufactured
warmth is a stronger machine-tell than plain institutional voice. Everything
here has to be specific, occasionally awkward, and willing to be unflattering,
or it makes the problem worse.
## Positioning
The author is **Tudor**: first name only, real photograph, no surname, no
employer named.
The credibility claim is deliberately **not** educational expertise. The About
page states plainly: *"I'm not an education expert."* Authority comes from two
things that are actually true:
- **Experience.** A parent going through primary admissions in south-west London
right now. Google's E-E-A-T leads with Experience, and lived experience of the
thing is exactly what the DfE's own service lacks.
- **Method.** Every number's provenance is stated, so a reader can check the
site rather than trust it.
This is more durable than borrowed expertise: it cannot be undermined by someone
noticing the author has no teaching qualification.
**Consequence for the design.** A `Person` entity with no surname is a weak
search signal and cannot be corroborated off-site. The credibility load
therefore shifts onto the methodology being visibly rigorous. That is a design
constraint, not a caveat — it is why the About page carries a substantial
"how this is built and where it can be wrong" section rather than a short bio.
### Voice rules
Applied to About and every post. Recorded here so the voice does not drift.
- First person singular. "I built", not "we provide".
- Concrete over general. "when we were looking at schools in Wandsworth" beats
any amount of stated warmth.
- State limits before someone else finds them. Every post that presents a
metric says what it does not show.
- No mission statements, no "passionate about", no invented team.
- No em dashes. One of the clearest tells of machine-written prose, which is
the exact problem this work exists to fix.
- Short sentences. The existing code comments in this repo are already written
this way; the prose should match.
## Scope
**In:**
- `/about` — a coded page (not CMS-managed).
- `/blog` and `/blog/[slug]` — Payload-backed, with an index and post pages.
- Payload CMS installed into the existing Next application.
- Footer and navigation links to both.
- `Person`, `Organization`, `BlogPosting`, `BreadcrumbList` JSON-LD.
- RSS feed and sitemap integration.
- One first post, so the blog does not launch empty.
**Out (deliberately):**
- Rewriting existing homepage/how-it-works copy into first person. Worth doing,
but it would double the review surface of this PR. Separate change.
- In-product signed notes on school pages (the "distributed humanity" idea).
Revisit once About and the blog exist.
- Comments, newsletter, author accounts beyond one.
- A team page. There is no team.
## Architecture
### Topology
Payload 3 installs **into the existing Next application** and serves `/admin`
from the same container. One image, one deploy, no new service. This is
Payload 3's native model and it makes on-demand revalidation trivial, because
the CMS hooks run in the same process as the Next cache.
Accepted costs: the public site's image now carries Payload, so a CMS security
patch redeploys the whole site; and the image grows substantially.
### Two collisions that must be handled
**1. `/api` is already taken.** `app/api/[...path]/route.ts` is a catch-all that
proxies `/api/*` to FastAPI at runtime. Payload's default API route is also
`/api`. Left alone, these fight, and the failure is not clean — the catch-all
would swallow Payload's admin API calls and forward them to FastAPI.
Payload's API route is therefore remapped:
```ts
routes: { api: '/cms-api', admin: '/admin' }
```
with its route group at `app/(payload)/cms-api/[...slug]/route.ts`. The
`/cms-api` prefix must also be added to the FastAPI proxy's excluded-paths list
as a defensive second line.
**2. `next.config.js` is CommonJS.** Payload's `withPayload()` wrapper is ESM
only. The config must become `next.config.mjs`, converting `module.exports` to
`export default` and wrapping the export. All existing content — the standalone
output, `outputFileTracingIncludes`, the staging `X-Robots-Tag` header block,
the CSP — carries over unchanged. This is mechanical but it touches the file
that controls staging's noindex, so it needs care and an explicit test.
### Database
Payload uses the existing `sc_database` Postgres instance, in its **own
`payload` schema**:
```ts
db: postgresAdapter({
pool: { connectionString: process.env.DATABASE_URL },
schemaName: 'payload',
})
```
The frontend container is already on the `backend` Docker network, so it can
reach `sc_database:5432` with no networking change. It needs a new
`DATABASE_URL` environment variable.
Schema isolation is not cosmetic. `public` currently holds the application
tables and Airflow's metadata, and `scripts/migrate_csv_to_db.py --drop` exists
to drop and reimport. Blog content living in its own schema means no data
pipeline operation can destroy it.
**Verified 2026-09-02** (this was an open question when the spec was written).
`--drop` calls `run_full_migration()` in `backend/migration.py`, which drops
exactly two tables by name:
```python
ks2_tables = ["school_results", "schools"]
for tname in ks2_tables:
if tname in existing:
Base.metadata.tables[tname].drop(bind=engine)
```
There is no `Base.metadata.drop_all()` anywhere in `backend/`, and no
`DROP SCHEMA`. The only other drop is `_apply_schema_drops()`, a single
schema-qualified `DROP TABLE IF EXISTS marts.fact_parent_view CASCADE`.
Nothing sets `search_path`, so the SQLAlchemy metadata resolves to `public`,
and `inspector.get_table_names()` does not even enumerate other schemas.
So the guarantee is stronger than schema isolation alone: `--drop` targets two
named tables that Payload does not have, and would not reach `posts`, `media`
or `users` even if they shared a schema. The `payload` schema remains the right
choice — it protects against a *future* broadening of that script rather than
today's behaviour — but the safety claim rests on verified code, not on
assumption.
Putting CMS tables in this instance is consistent with existing practice —
Airflow already stores its metadata there.
### Migrations
Payload's Postgres adapter auto-pushes schema in development and requires
explicit migrations in production. Use `prodMigrations`, which runs pending
migrations during server initialisation:
```ts
db: postgresAdapter({ /* ... */, prodMigrations: migrations })
```
This is preferred over a one-shot init container (the `airflow-init` pattern)
because the app is a single long-running process and there is no ordering
problem to solve. Migration files are generated with `payload migrate:create`
and committed, so schema changes travel through the same PR and staging gate as
code.
### Media
Uploads go to a Docker named volume, consistent with `postgres_data`,
`typesense_data` and `airflow_logs`.
- `staticDir` must be an **absolute** path in Payload 3: `/app/media`.
- The container runs as `nextjs` (uid 1001). The Dockerfile must
`mkdir -p /app/media && chown nextjs:nodejs /app/media` **before** the volume
is mounted, or Docker will create the mountpoint root-owned and every upload
will fail with EACCES.
- `sharp` moves from `devDependencies` to `dependencies` — Payload needs it at
runtime to generate `imageSizes`.
- The volume must be added to the backup routine alongside Postgres. A blog
post's images are not reproducible from the pipeline.
### Rendering
**Constraint:** CI builds the image with no database reachable. Blog pages
therefore cannot use build-time `generateStaticParams` — that would either fail
the build or bake in an empty post list.
Instead: ISR. Post and index pages declare a `revalidate` window and render on
first request, with Payload `afterChange` / `afterDelete` hooks calling
`revalidatePath('/blog')` and `revalidatePath('/blog/' + slug)` for immediate
publication. Because Payload runs in the same process, the hook calls
`revalidatePath` from `next/cache` directly — no webhook, no shared secret.
The ISR cache lives on container disk and is cleared by a redeploy. For a
single container serving a handful of posts this is fine.
### Collections
- **`posts`** — `title`, `slug`, `publishedAt`, `excerpt`, `heroImage`
(relation to `media`), `content` (Lexical rich text), `seo` group
(`metaTitle`, `metaDescription`), `_status` (drafts enabled).
- **`media`** — upload collection, `alt` required, `imageSizes` for thumbnail
and hero widths, public read access.
- **`users`** — Payload's auth collection. One account. Public creation
disabled.
Drafts are enabled so posts can be written over several sittings and previewed
before publication.
**Payload Blocks** are how posts embed live product components — a real trend
chart or comparison table inside a post, rendered from live data rather than
screenshotted. This is the main thing the CMS has to earn back against
file-based MDX, and it directly serves the goal: showing judgement in context.
Ship with one block (a callout/aside for "what this number doesn't tell you");
add a live-chart block once a post needs it.
### Security
`/admin` is the first authenticated surface on this site. Public, hardened:
- `PAYLOAD_SECRET` — long, random, set in the Portainer stack environment, never
committed. The same variable must exist in staging with a *different* value.
- Strong unique password on the single admin account.
- Login rate limiting via Payload's `maxLoginAttempts` / `lockTime`.
- `X-Robots-Tag: noindex, nofollow` on `/admin/*` and `/cms-api/*`, and a
`robots.ts` disallow. The admin panel must never be indexed.
- Public user creation disabled; no open registration.
- Verify the existing CSP `frame-ancestors` directive does not break the admin
panel.
Residual risk, accepted: a future Payload authentication CVE is live against the
public internet. Mitigation is prompt patching, which the staging→prod pipeline
already supports. If this becomes uncomfortable, restricting `/admin` at the
proxy to LAN/VPN is a one-line change later.
Staging note: staging runs the same image on `stx.`, so it gets its own admin
panel and its own database. It must have its own `PAYLOAD_SECRET` and its own
credentials — never production's.
## Deployment changes
- `nextjs-app/Dockerfile` — create and chown `/app/media`; ensure Payload's
admin bundle and `sharp` survive standalone output file tracing.
- `docker-compose.portainer.yml` and the staging equivalent — add
`DATABASE_URL` and `PAYLOAD_SECRET` to the `frontend` service, add a
`payload_media` volume mounted at `/app/media`, and add
`depends_on: sc_database`.
- Document both new environment variables in the compose header comment block,
which is where this stack records its configuration.
## SEO
- `Person` (Tudor, with photo) and `Organization` JSON-LD on `/about`.
- `BlogPosting` + `BreadcrumbList` on post pages, with `author` referencing the
same `Person`.
- Canonical URLs on `/blog` and every post.
- Posts and `/about` added to the existing sitemap (`app/sitemap.xml/route.ts`
and `app/sitemaps/[...parts]`). Post URLs come from Payload at request time.
- RSS feed at `/blog/rss.xml`.
- Footer links to both pages, under a new "About" column.
**Navigation is deliberately left alone.** `Navigation.tsx` renders a bottom tab
bar on mobile that already carries four items (Search, Compare, Rankings,
Admissions). A fifth tab makes each one cramped at 320px, and About and Blog are
both lower-intent than any of the four. Both live in the footer; About
additionally gets a byline link from every post, which is where a reader who
cares actually asks the question. Revisit only if analytics show people hunting
for it.
## Testing
Unit (Jest):
- Post rendering, including a post with no hero image and one with no excerpt.
- Slug generation and collision handling.
- JSON-LD shape for `BlogPosting` and `Person`.
- The `next.config.mjs` conversion preserves the staging `X-Robots-Tag` rule —
this guards the riskiest mechanical change in the plan.
E2E (Playwright, `e2e/`, required by CLAUDE.md for user-facing change):
- `/about` renders, shows the author name and photo, and is reachable from the
footer and nav.
- `/blog` lists at least one post; clicking through reaches the post.
- A post page renders title, date, body and byline.
- `/admin` responds with `noindex` and does not leak a stack trace when
unauthenticated.
Note the known constraint: new journeys cannot be proven in PR checks, because
the staging E2E gate runs post-merge.
## Risks
| Risk | Mitigation |
|---|---|
| `next.config.mjs` conversion silently drops the staging noindex header, making staging a crawlable duplicate | Unit test asserting the header rule; verify on staging before promotion |
| Payload API route collides with the FastAPI `/api` proxy | Remap to `/cms-api`; add to the proxy's exclusion list |
| Media volume mounts root-owned; all uploads fail with EACCES | `mkdir`+`chown` in the Dockerfile before the mount; test an upload on staging |
| Build fails or bakes empty content because CI has no DB | No build-time DB access; ISR only |
| A pipeline `--drop` destroys blog content | Separate `payload` schema; verify `--drop` blast radius before building |
| Media volume not backed up; images unrecoverable | Add `payload_media` to the backup routine |
| Payload auth CVE exposed publicly | Prompt patching; proxy restriction available as a fallback |
| Blog launches empty or goes stale | Ship with one post; cadence is explicitly "a few times a year", so no cadence is promised anywhere on the page — no dates implying a schedule |
## Sequence
Each step is independently reviewable and mergeable.
1. **Payload foundation** — install, `next.config.mjs` conversion, `payload`
schema, `/cms-api` remap, `users` collection, `/admin` hardening, compose and
Dockerfile changes. No public-facing change yet. Verify on staging that the
site is unchanged and `/admin` works.
2. **`/about`** — coded page, photo, `Person`/`Organization` JSON-LD, footer and
nav links, e2e journey. Independently valuable and does not depend on the
blog.
3. **Blog** — `posts` and `media` collections, `/blog` index and post pages, ISR
plus revalidation hooks, RSS, sitemap, structured data, e2e journeys.
4. **First post** — written in the admin panel, published through the normal
flow, proving the whole path end to end.
Step 1 carries all the infrastructure risk and none of the visible benefit, so
it should be verified on staging carefully before step 2 starts.
## Dependencies on Tudor
- **A photograph.** Blocks step 2. Nothing else in the plan is blocked by it.
- **The first post's subject.** Blocks step 4 only. Suggested: what school
performance data cannot tell you — it demonstrates judgement, is genuinely
useful, and is the kind of thing an anonymous or machine-written site will not
publish.
- ~~Confirmation that `scripts/migrate_csv_to_db.py --drop` is schema-scoped.~~
**Resolved 2026-09-02** — verified in `backend/migration.py`; see the
Database section. No action needed.
@@ -0,0 +1,432 @@
# Other Schools Nearby — Design
**Date:** 2026-09-21, revised 2026-09-22
**Status:** revised after staging review
**Scope:** school detail pages, both phase templates
> **Revision, 2026-09-22.** The first build ranked by intake similarity and used
> distance as a tiebreak. On staging a Catholic primary showed six Catholic
> primaries, none of them close enough to be a real option, and omitted the
> community school down the road. Distance now decides the order and nothing
> else does; the tier system is gone. The reasoning is kept below rather than
> quietly overwritten, because the mistake is the instructive part.
## Goal
Give a school detail page an answer to the question every reader arrives with
after the results tables: *and what else is around here?*
Today a school page links outward to its place pages through
`components/school/NearbyPlaces.tsx` and nowhere else. It never links to another
school. This section adds that edge — up to six nearby schools of the same phase
and a comparable intake, three at a time in a carousel, each a crawlable link and
each addable to the comparison basket in one click.
Mockup, in both themes (drawn against the original tiered design, so its ledes
and chip fallbacks are one revision behind the copy specified below):
<https://claude.ai/artifact/168KdUMcfkUeGWW2FGjuec>
Source of the same page in the repo: `mockups/similar-schools-nearby.html`.
## The constraint that shapes everything
**A nearby school is not automatically a comparable school.**
The section's whole value is that a reader treats what it shows as a shortlist.
That makes every card an implicit claim that the school is a realistic
alternative, and there are three ways that claim goes wrong:
1. A **selective** school beside a non-selective one. Their intakes are
different by construction, so putting their Attainment 8 figures side by side
invites a conclusion the data cannot support.
2. A **special school, PRU or AP** beside a mainstream school. This is the same
error PR #70 fixed for the England benchmark, where Greenmead (URN 101099)
rendered "0% — 62 below England".
3. A **single-sex** school of the opposite sex. Not a weak match — not an option
at all.
So the design separates two kinds of fact, and never confuses them:
- **Hard filters** encode the claims above. They decide eligibility, and are
never relaxed at any distance, even if that means the section does not render.
- **Shared characteristics** — gender, religious character, selectivity —
describe how closely an intake resembles this school's. They are *reported on
the card and never ranked on*, so the reader weighs them rather than having
them weighed for them.
Everything below follows from that split. The revision at the top of this
document is what happens when the second kind is treated as the first.
## Selection algorithm
A backend helper, `_nearby_schools_payload(urn)` in `backend/app.py`, modelled
on the existing `_places_payload(urn)` and delegating to
`backend/nearby_schools.select_nearby(frame, urn)`, which operates on the cached
`load_latest_school_data()` frame — one row per URN, already carrying
`latitude`, `longitude`, `phase`, `gender`, `religious_denomination`,
`admissions_policy`, `school_type` and `status`.
### Hard filters
| Filter | Rule |
|---|---|
| Self | `urn` is excluded |
| Status | GIAS status must be open |
| Coordinates | both `latitude` and `longitude` present on both schools |
| Phase | same phase group via the existing `PHASE_GROUPS` map |
| Provision | special/PRU/AP match only each other |
| Selectivity | selective matches selective; non-selective matches non-selective |
| Gender | Boys never matches Girls; Mixed is compatible with both |
`PHASE_GROUPS` is reused rather than re-derived so an all-through school is
offered correctly on both the primary and secondary sides, exactly as it already
behaves in search.
The provision filter needs a backend counterpart to the frontend's
`isSpecialSchool()` in `nextjs-app/lib/utils.ts:897`, reading the same GIAS
establishment types through `backend/gias_codes.py`. The two must agree: a
school the frontend treats as special for benchmarking but the backend treats as
mainstream for matching would be dropped from its own England comparison and
then offered as a peer to a mainstream school on the next page along.
**Up to six cards, three visible.** Three fit the row; the rest are reached with
the carousel arrows. Two is the minimum that renders at all.
### Order: distance, and nothing else
The nearest eligible schools, closest first. Similarity does not enter the
ranking at any point.
**Why not, having built it the other way first.** The original design ranked by
tiers — same gender and faith within 3 miles, then same gender within 5, then
anything within 10 — and used distance only to order the result. Two things
followed, and both showed up on the first Catholic primary anyone looked at:
- A faith match at 2.9 miles outranked a community school at 0.3 miles. For a
primary, whose catchment is routinely under a mile, the far school is not a
weaker option; it is not an option.
- Because the row filled from the best tier before widening, three Catholic
schools within 3 miles were enough to fill all six slots with Catholic
schools. The stopping rule that produced this had been added to prevent the
*opposite* failure — padding a row with weak distant matches — and made this
one certain.
The premise was backwards. **Distance is a constraint and intake is a
preference.** A parent cannot act on a school outside their reach however well
it matches, and they are perfectly capable of noticing a shared denomination
for themselves if we show it to them. So similarity moved from the ranking to
the card: `shared` reports what a school genuinely has in common, and the reader
applies their own weighting.
The hard filters above were always where the defensibility lived. They are
untouched.
### Reach: a sanity bound, not a target
| Phase | Reach |
|---|---|
| Primary, middle deemed primary, all-through | 2 miles |
| Secondary, middle deemed secondary | 6 miles |
| 16 plus | 10 miles |
Ordering by distance already handles density — a school in inner London fills
all six slots inside a mile and never approaches the cap. The cap decides one
thing: what happens where the area is sparse. It differs by phase because
catchments do, and because people travel furthest for post-16.
**A primary with nothing inside two miles renders no section**, and that is the
intended answer rather than a gap. The alternative is a section headed "nearby"
listing a school four miles from a five-year-old.
**Past the sixth school, the rest are dropped without a count.** The section
does not try to be the list: `NearbyPlaces` sits directly beneath and already
leads to the place pages, which are built for browsing a full set and which the
school page exists to feed.
**Fewer than two results renders nothing.** Not an empty state, not a single
lonely card. The section is absent, the nav item is absent, and the page is
unchanged from before it existed.
### Distance
Straight-line, from the vectorised haversine already used for postcode search at
`backend/app.py:831`, computed over the ~27k-row frame in numpy. Reported to one
decimal place in miles, consistent with the rest of the site.
Straight-line distance is not road distance and is not measured from the
reader's home. The section says so in its disclosure rather than leaving the
reader to assume otherwise.
## API
`/api/schools/{urn}` gains a `similar_schools` array. Each row:
| Field | Notes |
|---|---|
| `urn` | for the link and the compare basket |
| `school_name` | link text |
| `distance_miles` | one decimal place |
| `school_type` | GIAS type, translated, for the card's meta line |
| `age_range` | for the meta line |
| `shared` | what this school genuinely shares with the subject; may be empty |
| `metric_value` | the phase-appropriate headline figure, or null |
| `metric_key` | `rwm_expected_pct` or `attainment_8_score` — see below |
| `metric_year` | the year the figure is from |
The metric follows the subject school's phase side, not the neighbour's own
phase, so a row of cards never mixes two scales. The secondary
side uses `attainment_8_score`; the primary side uses `rwm_expected_pct`. Where
the neighbour has no value for that key, the card reads "Not published" rather
than falling back to the other key.
Which side a school takes is decided once, in
`similar_schools.is_secondary_phase`, by membership of `PHASE_GROUPS["secondary"]`
minus all-through — never by testing for the substring "secondary", which misses
`16 plus` (GIAS phase 6) and hands a sixth-form college the primary bucket.
All-through is the exception in the other direction: `PHASE_GROUPS` lists it on
both sides, but it takes the primary metric.
This is usually the same thing as "the template the page renders", but not
always. `computeSchoolFlags` decides the template with that same substring test,
so a `16 plus` school renders `PrimarySchoolSections` while being matched —
correctly — against secondaries. The section therefore takes its lede noun from
the school's own phase rather than from its template, or it would print "Other
primary schools near <sixth form college>" above a row of secondaries.
Up to six rows of roughly 130 bytes each. It rides in the existing detail payload
rather than a new endpoint because the page already makes exactly one server
fetch for its data, and `/school/[slug]` regenerates at most weekly
(`revalidate = 604800`), so the per-request cost is paid once per school per
week.
**The key is absent, not null, on a backend that does not have this code.** The
frontend treats absent and empty identically, which is what allowed
`NearbyPlaces` to ship without a lockstep deploy of the two images.
`shared` is computed on the backend, beside the data it is derived from, not
re-derived on the frontend. Deriving it twice is how a card comes to claim
something the selection never established. An empty list is a real answer and
renders no chips: a bare card costs a school nothing but the likeness it does
not have, since the order was already settled by distance.
## Frontend
### Components
`components/school/SimilarSchoolsSection.tsx` — a server component wrapped in
the shared `Section` shell from `sectionShared.tsx`. It renders the heading,
the lede, the card grid, the footer CTA and one caption line. Every
card's title is an `<a>` to the school's canonical slug URL via `schoolUrl()`.
`components/school/AddToCompareButton.tsx` — calls `addSchool` from
`ComparisonProvider` and reports the selection with a `from: 'similar_schools'`
attribution, mirroring `addSchoolFromSearch` in `HomeView.tsx:442`.
`components/school/SimilarSchoolsCarousel.tsx` — the scroller and its arrows. It
takes the server-rendered cards as `children` and the server-rendered heading and
lede as a `header` prop, so those stay server components while the client
component owns only the ref, the scroll handler and the arrows' disabled state.
The split matters: the links — the part with SEO value and the part that must
work without JavaScript — are server-rendered into the initial HTML, and only
the basket interaction and the arrows are hydrated.
### The carousel
**Every card is in the initial HTML.** The arrows scroll a list; they never swap
a view. Six `<a>` elements are in the markup whether or not anything is
hydrated, which is the whole reason the section exists — a paginated widget that
mounts cards on click would put four of the six links beyond a crawler and
beyond a reader with no JavaScript.
So the scroller is a plain overflowing `<ul>` with `scroll-snap-type: x
mandatory`, and the arrows call `scrollBy` on it. With no JavaScript it
degrades to a horizontally scrollable row that still works by touch and by
trackpad. Three cards are visible at desktop width and two below 820px.
**Arrows appear only when there is somewhere to go** — that is, only when more
than three schools were found. Each disables itself at its own end of the
travel.
#### Below 640px the arrows go away
This follows [MOBILE.md](../../../MOBILE.md), which makes 360px the design
floor and mobile the primary target at ≥55% of traffic.
Kept in the heading's flex row at 360px, the two arrow buttons take 96px from a
328px card and crush the lede into a four-line column — measured, not guessed.
And swiping already does what they do. So below 640px the header becomes a
single column, the arrows are not rendered, one card shows at 86% width so the
next one peeks, and the affordance is carried by the right-edge scroll-fade that
MOBILE.md documents for exactly this case:
```css
mask-image: linear-gradient(to right, #000 calc(100% - 28px), transparent);
```
The fade lifts at the end of the travel, where there is nothing left to hint
at. That means the at-end state must be computed whether or not an arrow exists
to consume it — on mobile it drives the mask alone.
**Every interactive element clears 44×44px**, per MOBILE.md's iOS HIG check: the
arrow buttons and the add-to-compare button are both 44px, up from the 40px they
were first drawn at. A card title's own box is shorter than that, but its hit
area is the whole card through the `::after` overlay, so it passes on the target
that actually receives the tap.
**The edge test needs a tolerance, and this is not fussiness.** The scroller
carries 2px of padding so focus rings are not clipped, and scroll-snap treats
that padding as the first card's snap position: a scroller sitting at its start
reports `scrollLeft` of 2, not 0. Sub-pixel rounding moves it again at other
zoom levels. Testing `scrollLeft === 0` therefore leaves the back arrow live and
pointing nowhere on first paint — confirmed in the mockup before it was fixed.
Both ends compare against an 8px tolerance.
**Selecting a school must not move the row.** Adding to the basket re-renders
the footer; the scroll offset lives in the DOM rather than in React state, so
the carousel must not remount or reset on that render. A reader who ticks the
fifth school and is thrown back to the first has been punished for using the
feature.
### Placement and navigation
Rendered as the last section **inside** `SchoolDetailShell`, from both
`PrimarySchoolSections` and `SecondarySchoolSections`. Inside, not after, because
the sticky nav's scroll-spy locates sections with `document.getElementById` and
can only reach a section that lives in the shell.
`NearbyPlaces` stays where it is, outside the shell, immediately below. The
resulting order — this school, then similar schools, then the places containing
them — narrows before it widens, which is the order a reader leaves a page in.
`buildNavItems` and `buildSecondaryNavItems` both gain
`{ id: 'similar', label: 'Similar schools' }`, gated on the section rendering.
The id must match the `Section` id or the scroll-spy silently breaks.
### The comparison CTA
A plain `<a href="/compare?urns=…">`, built from this school's URN plus the
selected ones. `/compare` already parses `urns` from the query string
(`app/(frontend)/compare/page.tsx:55`), so this needs no new compare plumbing.
With nothing selected the CTA is disabled; the button also adds to the shared
basket so the site-wide comparison state stays consistent with what the page
shows.
## Copy, and what the section is allowed to claim
**The lede never claims an intake.** It reads "Other primary schools near X." —
one sentence, no variants. The earlier version varied the wording by tier, which
only existed to soften a claim the section should not have been making.
**The heading is "Other schools nearby", not "Similar schools nearby".** The
hard filters do guarantee a comparable set — same phase, same selectivity,
mainstream never beside special — but nothing ranks on likeness, so the heading
does not say it does. The nav item reads "Nearby schools" and the section id is
`nearby`.
**Chips state only what is shared, and may be absent entirely.** A card with
nothing in common renders no chip row rather than falling back to a filler.
Since chips no longer affect the order, an empty one costs that school nothing
except a claim it cannot support — and a Catholic parent scanning the row still
spots "Roman Catholic" on the card that carries it, and weighs it themselves.
**The neighbour's metric carries no valence colour.** Green and terracotta are
reserved site-wide for comparison against the England average. Colouring a
neighbour's figure against this school's would read as ranking the neighbours
against each other, which is precisely the endorsement this section must not
make. The figure sits in neutral ink above a plain "72% at this school"
reference line, and the reader draws their own conclusion.
**A missing figure reads "Not published".** Never 0, never blank, never an
em dash. This follows the same rule the rest of the detail page uses: a school
with no published result has not scored zero.
**There is no "how these are chosen" disclosure.** The method is visible in what
the section already shows — the phase in the lede, the shared characteristics on
each card, the distance above each name — and a collapsed panel restating it
earns less than the space it costs.
**One caption line survives, and only one:** that distances are straight-line
from the school and not road distance. This is not a method note. A reader who
sees "0.6 miles away" and takes it for the walk has been misled by us, and no
other element on the card corrects that. The remaining notes — that listing is
not a recommendation, that special schools only meet special schools — are
statements the selection rules already keep true without being narrated.
## Degradation
| Condition | Behaviour |
|---|---|
| `similar_schools` absent (older backend image) | no section, no nav item |
| fewer than 2 qualifying schools | no section, no nav item |
| this school has no coordinates | no section |
| the helper raises | returns `[]`; the page renders without the section |
The helper is wrapped so a failure inside it never 500s a page that is otherwise
complete — the posture `get_supplementary_data` already takes for its own
queries.
## Testing
**Backend**, in a new `backend/tests/test_similar_schools.py`, against a
synthetic frame rather than live marts:
- a selective school never returns a non-selective one, and vice versa
- a special school returns only special schools; a mainstream school returns none
- a Boys school never returns a Girls school; Mixed matches both
- closed schools and schools without coordinates are never returned
- results are ordered by distance ascending, always
- a faith match never outranks a closer school (the staging defect, pinned)
- the nearest eligible school is always present
- more than six qualifying schools returns the six nearest
- reach is capped per phase, and a primary beyond two miles returns `[]`
- an all-through school is offered on both phase sides
- a `16 plus` school is matched against secondaries and colleges, never primaries
- `is_secondary_phase` and `PHASE_GROUPS` agree on every GIAS phase value
- fewer than two qualifying schools returns `[]`
- distances match a hand-computed haversine for a known pair
**Frontend**, in `nextjs-app/__tests__`:
- the section renders nothing for absent, empty and single-row inputs
- the lede never claims a similar intake
- an empty `shared` renders no chips rather than a filler
- a null metric renders "Not published"
- the nav item appears only alongside the section
- every card is in the DOM, including the ones scrolled out of view
- arrows render only when more than three schools were found
jsdom has no layout, so `scrollWidth` and `clientWidth` are both 0 there and the
arrows' disabled state cannot be meaningfully asserted in Jest. That behaviour is
covered in the journey instead, against a real engine, rather than by a unit test
that would pass on a measurement that does not exist.
**E2E**, added to the existing journeys in `e2e/tests` in the same PR, per the
repository's rule on user-facing behaviour:
- the section renders on a known staging URN, with resolving links
- where arrows are present, the back arrow starts disabled and the forward arrow
moves the row
- selecting a school does not reset the scroll position
- add-to-compare reaches `/compare` with the expected `urns`
- at 360, 390 and 430px: no horizontal overflow, every interactive element in the
section clears 44×44px, and no arrows are rendered
The E2E gate runs after merge on this project, so these journeys are not
provable in the PR checks; the PR is verified on the unit tests, and the
journeys are confirmed on the post-merge staging run.
## Out of scope
- A map of the nearby schools. The section is a list; the page already has a map.
- Autoplay, dots, or an infinite loop on the carousel. It is a short list a
reader scans deliberately, not a banner competing for attention, and a row
that moves on its own is a row that moves while someone is reading it.
- Statistical neighbours on deprivation, size or cohort profile. If plain
distance proves too blunt, that is the trigger to move this computation into a
dbt mart — `select_nearby` is a deliberate seam for exactly that swap.
- Precomputing neighbours in `marts.*`. Rejected for now: a new mart is inert
until Airflow runs, so the feature would ship dark, and every tuning change to
the rules would become a pipeline round-trip instead of a deploy. The revision
at the top of this document is the argument for keeping that loop short.
- Any change to `/api/compare`, the compare page, or the comparison basket.
+708 -6
View File
@@ -1,4 +1,4 @@
import { test, expect, Page } from '@playwright/test';
import { test, expect, Locator, Page } from '@playwright/test';
/**
* Journey tests for SchoolCompare, run against the staging environment as the
@@ -19,6 +19,31 @@ function schoolLinks(page: Page) {
return page.locator('a[href^="/school/"]');
}
/**
* A scroll offset that has stopped moving.
*
* The carousel arrows scroll with `behavior: 'smooth'`, so a reading taken
* straight after a click lands mid-animation. Measured against staging: the
* animation runs ~700ms, and a poll for "has it moved at all" is satisfied
* 50ms in, at 13px of a 1300px journey. A test that then records an offset,
* does something, and records again is measuring the tail of the arrow's
* animation rather than the effect of whatever it did in between.
*
* Two identical readings in a row is the cheapest sound definition of settled.
*/
async function settledScrollLeft(scroller: Locator): Promise<number> {
let previous = -1;
await expect
.poll(async () => {
const current = await scroller.evaluate((node: HTMLElement) => Math.round(node.scrollLeft));
const settled = current === previous;
previous = current;
return settled;
}, { timeout: 10_000 })
.toBe(true);
return previous;
}
/**
* Two URNs guaranteed to be pure-primary (same phase). The compare page's
* phase tabs split all-through schools (which carry KS4 data) onto the
@@ -1304,6 +1329,52 @@ test('with the distance feature off, the section is absent rather than empty', a
.toHaveCount(0);
});
/**
* A secondary school carrying an EES admissions row, which is what makes its
* Admissions section render while the distance feature is dark.
*/
async function secondarySchoolWithAdmissions(page: Page) {
const list = await page.request.get('/api/schools?phase=secondary&page_size=40');
if (!list.ok()) return null;
const body = await list.json();
for (const s of (body?.schools ?? []).slice(0, 25)) {
const res = await page.request.get(`/api/schools/${s.urn}`);
if (!res.ok()) continue;
const detail = await res.json();
if (detail?.admissions == null) continue;
return { urn: s.urn as number };
}
return null;
}
test('with the distance feature off, a secondary page makes no claim about publication', async ({ page }) => {
/*
* Shipping dark must not put words in the council's mouth. The secondary
* template is the only one that words the absence, and "X has not published
* a cut-off distance for this school" is false wherever X does publish and
* we are simply withholding it.
*
* This is why the API omits the key rather than sending null: absent means
* "cut-offs are not published at all", null means "this school has none".
* Only the second is a fact about the school, and only the second is sayable.
*/
test.skip(await distanceFeatureIsOn(page),
'the admission_distance flag is on in this environment');
const found = await secondarySchoolWithAdmissions(page);
test.skip(found === null, 'no secondary school in the sample has an admissions row');
await page.goto(`/school/${found!.urn}`);
await expect(page.locator('h1').first()).toBeVisible({ timeout: 15_000 });
// The Admissions section is still there — this is not a test that the whole
// section vanished, which would pass for the wrong reason.
await expect(page.locator('#admissions')).toHaveCount(1);
await expect(page.getByText(/has not published a cut-off distance/)).toHaveCount(0);
await expect(page.getByText(/Contact the admissions authority/)).toHaveCount(0);
});
test('/api/flags is not reachable from the public internet', async ({ page }) => {
// It names every unreleased feature and whether it is on. Next reads it
// server-side over the Docker network; the public proxy must deny it.
@@ -1889,6 +1960,63 @@ async function firstPlaceOfKind(page: Page, kind: string) {
return hit as { kind: string; slug: string; name: string; count: number };
}
/**
* The round trip. Place pages always linked down to school pages; school
* pages linked nowhere on the site, so the ~27k of them that carry most of
* the inbound authority stranded it — their only anchor pointed at the
* school's own website.
*
* Asserting both directions is the point. A one-way link is what already
* existed and is not what this journey is for.
*/
test('a school page links back into the location layer, and the place page links down', async ({ page }) => {
const town = await firstPlaceOfKind(page, 'town');
// Start from the place page and take its first school, so the pair is
// guaranteed to be genuinely related rather than a hardcoded guess.
await page.goto(`/schools/${town.slug}`);
const schoolHref = await page.locator('a[href^="/school/"]').first()
.getAttribute('href');
expect(schoolHref, 'the town page listed no school to follow').toBeTruthy();
await page.goto(schoolHref!);
// Down: the school page must offer a link back to the town it sits in.
const backToTown = page.locator(`a[href="/schools/${town.slug}"]`);
await expect(backToTown).toHaveCount(1);
await expect(backToTown).toBeVisible();
// The anchor says what it leads to, which is worth more than "see more".
await expect(backToTown).toContainText(town.name, { ignoreCase: true });
await expect(backToTown).toContainText(/\d+ schools?/);
// And the breadcrumb resolves the school into a real hierarchy.
const blocks = await page.locator('script[type="application/ld+json"]')
.allTextContents();
const graph = blocks.join(' ');
expect(graph).toContain('"BreadcrumbList"');
// The narrower type, not the EducationalOrganization parent it used to be.
expect(graph).toContain('"School"');
/*
* The phase variants are the pages this most needs to reach: ~950 of them
* were once reachable by nothing at all, absent from every sitemap and
* unlinked from the place page. Conditional because not every school sits
* in a town that publishes one.
*/
const phaseLink = page.locator(`a[href^="/schools/${town.slug}/"]`).first();
if (await phaseLink.count()) {
const phaseHref = await phaseLink.getAttribute('href');
expect((await page.request.get(phaseHref!)).status()).toBe(200);
await expect(phaseLink).toContainText(/primary|secondary/);
}
// Following it lands on a real page, not a 404.
await backToTown.click();
await page.waitForURL(new RegExp(`/schools/${town.slug}$`));
await expect(page.locator('h1')).toContainText(town.name, { ignoreCase: true });
});
for (const [kind, prefix, article] of [
['town', '/schools/', 'a'],
['authority', '/schools/authority/', 'an'],
@@ -1978,6 +2106,59 @@ test('a place page links its phase variants, and they resolve', async ({ page })
await expect(page.locator('h1')).toContainText(new RegExp(`${phase} schools in`, 'i'));
});
/*
* The table shipped with one column of scores. A parent shortlisting from a
* town page needs to know whether a school takes their child's age, whether
* it is a faith school, and — for a primary — whether it has a nursery,
* before a percentage means anything.
*
* These assert the column headings rather than the values: nursery_provision
* and parliamentary_constituency are optional mart columns, and on an
* environment whose pipeline has not rebuilt them the API degrades them to
* absent. A value assertion would then fail for a data reason, not a code one.
*/
async function phasedPlace(page: Page, phase: 'primary' | 'secondary') {
const place = await firstPlaceOfKind(page, 'town');
const detail = await (await page.request.get(`/api/places/town/${place.slug}`)).json();
test.skip(!(detail.place.phases ?? []).includes(phase),
`no ${phase} page clears the threshold here`);
return place;
}
test('a primary place page names each school as well as scoring it', async ({ page }) => {
const place = await phasedPlace(page, 'primary');
await page.goto(`/schools/${place.slug}/primary`);
for (const heading of ['Ages', 'Religious character', 'Nursery', 'Constituency']) {
await expect(page.getByRole('columnheader', { name: heading, exact: true }))
.toBeVisible();
}
// age_range rides in on SCHOOL_COLUMNS and predates the optional columns,
// so it is the one attribute safe to assert a value for anywhere.
await expect(page.locator('table tbody td').filter({ hasText: /^\d+–\d+$/ }).first())
.toBeVisible();
});
test('a secondary place page does not ask about nurseries', async ({ page }) => {
const place = await phasedPlace(page, 'secondary');
await page.goto(`/schools/${place.slug}/secondary`);
await expect(page.getByRole('columnheader', { name: 'Ages', exact: true }))
.toBeVisible();
await expect(page.getByRole('columnheader', { name: 'Nursery', exact: true }))
.toHaveCount(0);
});
test('the measure stays beside the school name, not behind a swipe', async ({ page }) => {
// Six columns overflow a phone; .tableWrap turns that into a horizontal
// scroll. With the measure last, the number the page exists for is the one
// off the screen.
const place = await phasedPlace(page, 'primary');
await page.setViewportSize({ width: 390, height: 844 });
await page.goto(`/schools/${place.slug}/primary`);
const second = page.locator('table thead th').nth(1);
await expect(second).toContainText(/reading, writing/i);
await expect(second).toBeInViewport();
});
test('phase variants are submitted in the places sitemap', async ({ page }) => {
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toMatch(/\/schools\/[a-z0-9-]+\/primary</);
@@ -2094,12 +2275,32 @@ test('a place page lists its schools alphabetically', async ({ page }) => {
expect(town).toBeTruthy();
await page.goto(`/schools/${town.slug}`);
const names = await page.locator('a[href^="/school/"]').allTextContents();
expect(names.length).toBeGreaterThan(1);
const sorted = [...names].sort((a, b) =>
a.toLowerCase().localeCompare(b.toLowerCase()));
expect(names).toEqual(sorted);
/*
* Per table, not per page.
*
* An unphased place page renders one table per phase, and an all-through
* school legitimately appears in both — so the page's school links are not
* one alphabetical run and never were. This assertion used to collect them
* all together and only passed because no town it picked happened to hold an
* all-through school; when the data gave Abbots Langley one, Breakspeare
* School showed up in the primary table and again in the secondary, and the
* test failed on correct behaviour.
*/
const tables = page.locator('table');
const tableCount = await tables.count();
expect(tableCount).toBeGreaterThan(0);
let checked = 0;
for (let i = 0; i < tableCount; i++) {
const names = await tables.nth(i).locator('a[href^="/school/"]').allTextContents();
if (names.length < 2) continue; // a one-row table says nothing about order
const sorted = [...names].sort((a, b) =>
a.toLowerCase().localeCompare(b.toLowerCase()));
expect(names, `table ${i + 1} is not alphabetical`).toEqual(sorted);
checked++;
}
expect(checked, 'no table had enough rows to check the ordering').toBeGreaterThan(0);
});
test('the rankings page still orders by score, not name', async ({ page }) => {
@@ -2112,6 +2313,66 @@ test('the rankings page still orders by score, not name', async ({ page }) => {
expect(scores).toEqual([...scores].sort((a: number, b: number) => b - a));
});
/*
* Analytics on the location layer.
*
* Umami counts a pageview for every one of these URLs already. What it cannot
* say is which *kind* of location page earns engagement, because all four
* families share the /schools/ prefix — and that is the question that decides
* whether to keep investing in them.
*/
/** Capture Umami events, with the real script blocked so it cannot clobber
* the stub. Must be called before the first navigation. */
async function captureEvents(page: Page) {
const events: Array<{ name: string; data: Record<string, unknown> }> = [];
await page.route('**/analytics.schoolcompare.co.uk/**', (route) => route.abort());
await page.exposeFunction('__capture',
(name: string, data: Record<string, unknown>) => { events.push({ name, data }); });
await page.addInitScript(() => {
(window as unknown as { umami: unknown }).umami = {
track: (name: string, data: unknown) =>
(window as unknown as { __capture: (n: string, d: unknown) => void })
.__capture(name, data),
};
});
return events;
}
test('a location page reports which kind of place it is', async ({ page }) => {
const events = await captureEvents(page);
const place = await firstPlaceOfKind(page, 'authority');
await page.goto(`/schools/authority/${place.slug}`);
await expect.poll(() => events.find((e) => e.name === 'place_viewed'),
{ timeout: 10_000 }).toBeTruthy();
const event = events.find((e) => e.name === 'place_viewed')!;
expect(event.data.kind).toBe('authority');
expect(event.data.slug).toBe(place.slug);
expect(event.data.phase).toBe('all');
});
test('a school reached from a location page is attributed to it, not to direct', async ({ page }) => {
/*
* The defect this was written for. getNavigationSource had no case for
* /schools/, so every school view that came through the location layer was
* filed as 'direct' — the bucket you read as "typed the URL". The one
* measurement that says whether ~3,900 SEO pages work was reporting the
* wrong answer, confidently.
*/
const events = await captureEvents(page);
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
await page.locator('a[href^="/school/"]').first().click();
await page.waitForURL(/\/school\//);
await expect.poll(() => events.find((e) => e.name === 'school_viewed'),
{ timeout: 10_000 }).toBeTruthy();
expect(events.find((e) => e.name === 'school_viewed')!.data.from).toBe('place');
});
/*
* School autosuggest (spec 2026-08-26).
*/
@@ -2163,6 +2424,65 @@ test('typing a school name suggests it, and choosing it opens that school', asyn
await expect(page).toHaveURL(/\/school\/\d+/);
});
test('the whole dropdown is reachable, not clipped by the hero', async ({ page }) => {
/*
* The hero panel had overflow: hidden to clip its artwork to the rounded
* corners, and it clipped the dropdown too — 320px of list against 145px of
* panel below the input, so roughly half was cut off with nothing to say so.
*
* toBeVisible() does not catch this: it checks the box is non-empty and not
* visibility:hidden, and an ancestor's overflow clips neither. The invariant
* that does catch it is that the LAST option is the thing actually painted
* at its own coordinates — which fails for clipping and for occlusion alike.
*/
test.skip(!(await autosuggestIsOn(page)),
'the school_autosuggest flag is off in this environment');
const { schools } = await (await page.request.get('/api/schools?page_size=1')).json();
test.skip(!schools?.length, 'no schools in this environment');
await page.goto('/');
await page.getByRole('combobox').first().fill(
(schools[0].school_name as string).slice(0, 6));
const options = page.getByRole('option');
await expect(options.first()).toBeVisible();
const count = await options.count();
const painted = await options.nth(count - 1).evaluate((el) => {
const r = el.getBoundingClientRect();
const hit = document.elementFromPoint(r.left + r.width / 2, r.top + r.height / 2);
return { inside: el.contains(hit) || el === hit, bottom: Math.round(r.bottom) };
});
expect(painted.inside,
`the last option is not painted at its own coordinates (bottom ${painted.bottom}) `
+ '— an ancestor is clipping or covering the dropdown').toBeTruthy();
});
test('the dropdown does not survive into the results it produced', async ({ page }) => {
/*
* The bug that took the staging gate down, and it was not a test problem:
* after a search the results-page bar still holds the term, so the dropdown
* reopened on top of the results and swallowed the click on the first one.
* Playwright reported it as "<li role=option> intercepts pointer events"; a
* reader would simply have found their first result unclickable.
*/
test.skip(!(await autosuggestIsOn(page)),
'the school_autosuggest flag is off in this environment');
await page.goto('/');
await page.getByRole('combobox').first().fill('school');
await expect(page.getByRole('option').first()).toBeVisible();
await page.getByRole('button', { name: /Search/i }).first().click();
await page.waitForURL(/search=school/);
await expect(page.getByRole('listbox')).toHaveCount(0);
// And the results underneath are actually reachable, which is the point.
await page.locator('a[href^="/school/"]').first().click({ timeout: 15_000 });
await expect(page).toHaveURL(/\/school\//);
});
test('with autosuggest off, the search box is a plain input', async ({ page }) => {
test.skip(await autosuggestIsOn(page),
'the school_autosuggest flag is on in this environment');
@@ -2174,3 +2494,385 @@ test('with autosuggest off, the search box is a plain input', async ({ page }) =
await page.getByRole('button', { name: /Search/i }).first().click();
await expect(page).toHaveURL(/search=abbey/);
});
// ── Destination measures ───────────────────────────────────────────────────
//
// Two failure modes have to be told apart here, and conflating them is how
// this suite would either hide a regression or block the promotion pipeline:
//
// * the backend does not serve the `destinations` field at all — a code
// regression, or a deploy that did not land. FAILS.
// * the field is served but every school is empty — the annual EES DAG has
// not run on this environment yet. SKIPS, loudly.
//
// The second is a data-load precondition, not a defect, and it is true for
// every commit between this merging and the DAG being triggered. Failing on it
// would redden the staging gate for unrelated work. This is not the quiet skip
// 4f01fbd removed from the distance journeys: that one hid a broken feature
// behind a flag check, whereas the assertion that the code is deployed and
// correctly shaped still runs here on every commit.
async function secondaryWithDestinations(page: Page): Promise<{
urn: string; destinations: any;
}> {
const res = await page.request.get('/api/schools?search=school&per_page=100');
expect(res.ok()).toBeTruthy();
const body = await res.json();
const urns: string[] = (body.schools ?? [])
.filter((s: { phase?: string; attainment_8_score?: number | null }) =>
s.phase === 'Secondary' && s.attainment_8_score != null)
.map((s: { urn: number }) => String(s.urn));
expect(urns.length).toBeGreaterThan(0);
let served = false;
for (const urn of urns.slice(0, 25)) {
const detail = await page.request.get(`/api/schools/${urn}`);
if (!detail.ok()) continue;
const data = await detail.json();
// The key must exist, even as null. Its absence means the backend in front
// of us does not know about destinations at all.
if ('destinations' in data) served = true;
if (data.destinations?.ks4) return { urn, destinations: data.destinations };
}
expect(served,
'GET /api/schools/{urn} served no `destinations` key at all — the backend '
+ 'is missing this feature, not merely missing its data').toBeTruthy();
test.skip(true,
'No school has destination data yet: the annual EES DAG has not run on '
+ 'this environment. The API shape is correct, so this is a data-load '
+ 'precondition rather than a regression.');
throw new Error('unreachable');
}
test('a secondary school page says where its Year 11 leavers went', async ({ page }) => {
const { urn } = await secondaryWithDestinations(page);
await page.goto(`/school/${urn}`);
const section = page.locator('#destinations');
await expect(section).toBeVisible({ timeout: 15_000 });
await expect(section.getByRole('heading', { name: 'After Year 11' })).toBeVisible();
// The section must date its own cohort: destinations run about two GCSE
// years behind the results above them, and an undated figure reads as stale.
await expect(section).toContainText(/20\d{2}\/\d{2}/);
});
test('the destinations bar is absent entirely whenever a figure is withheld', async ({ page }) => {
const { urn, destinations } = await secondaryWithDestinations(page);
await page.goto(`/school/${urn}`);
const section = page.locator('#destinations');
await expect(section).toBeVisible({ timeout: 15_000 });
const allGroup = destinations.ks4.groups.all;
const suppressed = (allGroup?.categories ?? [])
.filter((c: { status: string }) => c.status === 'suppressed');
if (suppressed.length > 0) {
// R1: a bar drawn from the published segments leaves a gap whose width is
// the withheld figure, readable straight off the axis.
await expect(section.locator('[data-destination-segment]')).toHaveCount(0);
await expect(section.getByText(/withheld/i).first()).toBeVisible();
} else {
const published = (allGroup?.categories ?? [])
.filter((c: { status: string }) => c.status === 'published');
await expect(section.locator('[data-destination-segment]'))
.toHaveCount(published.length);
}
});
test('switching to disadvantaged pupils never reveals a withheld figure', async ({ page }) => {
const { urn, destinations } = await secondaryWithDestinations(page);
const disadvantaged = destinations.ks4.groups.disadvantaged;
test.skip(!disadvantaged, 'this school publishes no disadvantaged breakdown');
await page.goto(`/school/${urn}`);
const section = page.locator('#destinations');
await expect(section).toBeVisible({ timeout: 15_000 });
const radio = section.getByRole('radio', { name: /disadvantaged/i });
await expect(radio).toBeVisible();
await radio.click();
const suppressed = (disadvantaged.categories ?? [])
.filter((c: { status: string }) => c.status === 'suppressed');
if (suppressed.length > 0) {
await expect(section.locator('[data-destination-segment]')).toHaveCount(0);
// The residual must appear nowhere on the page — it is the withheld figure.
const cohort: number = disadvantaged.cohort;
const publishedTotal = (disadvantaged.categories ?? [])
.filter((c: { status: string }) => c.status === 'published')
.reduce((sum: number, c: { pupils: number }) => sum + c.pupils, 0);
const residual = cohort - publishedTotal;
const text = (await section.textContent()) ?? '';
expect(text).not.toMatch(new RegExp(`\\b${residual}\\b`));
}
});
test('a school with no sixth form has no post-16 destinations section', async ({ page }) => {
const res = await page.request.get('/api/schools?search=school&per_page=100');
const body = await res.json();
const noSixthForm = (body.schools ?? [])
.filter((s: { phase?: string; has_sixth_form?: boolean }) =>
s.phase === 'Secondary' && s.has_sixth_form === false)
.map((s: { urn: number }) => String(s.urn));
test.skip(noSixthForm.length === 0, 'no sixth-form-less secondary in this dataset');
await page.goto(`/school/${noSixthForm[0]}`);
await expect(page.locator('h1').first()).toBeVisible({ timeout: 15_000 });
// Absence is the correct statement, so there must be no placeholder either.
await expect(page.locator('#post16-destinations')).toHaveCount(0);
await expect(page.getByText(/destination data coming soon/i)).toHaveCount(0);
});
test('the destinations section never claims a pupil stayed at this school', async ({ page }) => {
const { urn } = await secondaryWithDestinations(page);
await page.goto(`/school/${urn}`);
const section = page.locator('#destinations');
await expect(section).toBeVisible({ timeout: 15_000 });
// The published file records the TYPE of place a leaver went to, never which
// one, so the page can never say a pupil stayed on here.
const text = (await section.textContent()) ?? '';
expect(text).not.toMatch(/stayed on (here|at this school)/i);
});
/**
* The About page and the blog exist to give the site a named human author.
* These journeys assert the load-bearing parts of that — a name, a face, the
* honesty claim, and a resolvable Person entity — rather than exact copy,
* which will be edited.
*
* Both are behind flags (about_page, blog), so each has a lit journey and a
* dark one. Flag state is read from the observable effect rather than from
* /api/flags, which the public proxy denies on purpose — the same approach
* distanceFeatureIsOn() takes above.
*/
async function aboutPageIsOn(page: Page): Promise<boolean> {
return (await page.request.get('/about')).ok();
}
async function blogIsOn(page: Page): Promise<boolean> {
return (await page.request.get('/blog')).ok();
}
test('with the about page off, it is absent rather than empty', async ({ page }) => {
test.skip(await aboutPageIsOn(page), 'the about_page flag is on in this environment');
// Dark means the URL does not exist, not that it renders empty: a 404 is
// what stops a crawler keeping the page in its index.
expect((await page.request.get('/about')).status()).toBe(404);
// A footer link into a 404 is the failure this flag has to avoid.
await page.goto('/');
await expect(page.locator('footer a[href="/about"]')).toHaveCount(0);
// And a sitemap must never advertise a URL that 404s.
const sitemap = await page.request.get('/content-sitemap.xml');
expect(await sitemap.text()).not.toContain('/about');
});
test('with the blog off, it is absent rather than empty', async ({ page }) => {
test.skip(await blogIsOn(page), 'the blog flag is on in this environment');
expect((await page.request.get('/blog')).status()).toBe(404);
expect((await page.request.get('/blog/rss.xml')).status()).toBe(404);
await page.goto('/');
await expect(page.locator('footer a[href="/blog"]')).toHaveCount(0);
const sitemap = await page.request.get('/content-sitemap.xml');
expect(await sitemap.text()).not.toContain('/blog');
// The admin panel is deliberately NOT flagged: posts have to be writable
// before the blog is readable, or there is nothing to turn on.
expect((await page.request.get('/admin')).status()).not.toBe(404);
});
test('the about page names a human author and is reachable from the footer', async ({ page }) => {
test.skip(!(await aboutPageIsOn(page)), 'the about_page flag is off in this environment');
await page.goto('/');
const aboutLink = page.locator('footer a[href="/about"]');
await expect(aboutLink).toBeVisible();
await aboutLink.click();
await page.waitForURL(/\/about$/);
await expect(page.getByRole('heading', { level: 1 })).toContainText('Tudor');
await expect(page.locator('img[alt*="Tudor"]')).toBeVisible();
// The credibility claim is lived experience plus stated provenance, not
// expertise. If this sentence ever disappears the positioning has drifted.
await expect(page.getByText(/not an education expert/i)).toBeVisible();
const jsonLd = await page
.locator('script[type="application/ld+json"]')
.first()
.textContent();
expect(jsonLd).toContain('"Person"');
// First name only — a surname here would be the one place it leaks.
expect(jsonLd).not.toMatch(/familyName/);
});
test('the blog lists posts and each one renders with a byline', async ({ page }) => {
test.skip(!(await blogIsOn(page)), 'the blog flag is off in this environment');
await page.goto('/blog');
await expect(page.getByRole('heading', { level: 1 })).toBeVisible();
const postLinks = page.locator('a[href^="/blog/"]');
// Data invariant: staging must carry at least one published post. If this
// fails, the environment has no content rather than the code being broken.
expect(await postLinks.count()).toBeGreaterThan(0);
await postLinks.first().click();
await page.waitForURL(/\/blog\/.+/);
await expect(page.getByRole('heading', { level: 1 })).toBeVisible();
await expect(page.getByText(/^By Tudor/)).toBeVisible();
const jsonLd = await page
.locator('script[type="application/ld+json"]')
.first()
.textContent();
expect(jsonLd).toContain('"BlogPosting"');
});
test('the admin panel is not indexable', async ({ page }) => {
const response = await page.request.get('/admin');
expect(response.headers()['x-robots-tag']).toContain('noindex');
});
test('the content sitemap lists the about page and is advertised in robots', async ({ page }) => {
const sitemap = await page.request.get('/content-sitemap.xml');
// Served whatever the flags say: robots.txt names it unconditionally, and
// with both dark it is a valid empty urlset rather than a 404.
expect(sitemap.ok()).toBeTruthy();
if (await aboutPageIsOn(page)) {
expect(await sitemap.text()).toContain('/about');
}
// The school corpus sitemap is proxied from FastAPI; this one is Next's.
// robots.txt must advertise both or the blog never gets discovered.
const robots = await page.request.get('/robots.txt');
const body = await robots.text();
expect(body).toContain('/sitemap.xml');
expect(body).toContain('/content-sitemap.xml');
});
/**
* Other schools nearby.
*
* The section is absent by design where fewer than two schools qualify, and the
* arrows are absent where three cards fit, so this asserts each part of the
* contract only where it applies.
*
* Two things here cannot be tested anywhere else: the arrows' disabled state,
* which jsdom cannot measure because it has no layout, and the scroll position
* surviving a selection, which is DOM state rather than React state.
*/
test('nearby schools link on to other schools and into compare', async ({ page }) => {
await searchByName(page, 'Primary');
await schoolLinks(page).first().click();
await page.waitForURL(/\/school\//);
const section = page.locator('#nearby');
if ((await section.count()) === 0) {
test.skip(true, 'No qualifying similar schools for this school');
}
// Every card is a real link to another school page — including the ones
// behind the arrows, which is the whole reason this is a scroller and not a
// paginated widget.
const links = section.locator('a[href^="/school/"]');
const linkCount = await links.count();
expect(linkCount).toBeGreaterThanOrEqual(2);
expect(linkCount).toBeLessThanOrEqual(6);
expect(await links.first().getAttribute('href')).toMatch(/^\/school\/\d{6}-/);
await expect(section.getByText(/miles away/).first()).toBeVisible();
const scroller = section.locator('ul').first();
// The carousel, where this school had more than three matches.
const forward = section.getByRole('button', { name: 'More schools' });
if (await forward.count()) {
const back = section.getByRole('button', { name: 'Previous schools' });
await expect(back).toBeDisabled();
await forward.click();
expect(await settledScrollLeft(scroller)).toBeGreaterThan(8);
await expect(back).toBeEnabled();
}
// The compare hand-off, and the row must not jump back to the start when the
// footer re-renders underneath it.
//
// Click the LAST card's button, not the first. Playwright scrolls a target
// into view before clicking it, so clicking card one while the row is paged
// to the end scrolls the container back to the start — and the assertion
// below then measures Playwright's own scrolling rather than the app's.
// That is what this test did on its first staging run: 537 → 2, reproduced
// afterwards on a static page with no React on it at all.
//
// A few pixels of snap or sub-pixel adjustment are fine; a reset to the
// start is not, which is the whole point of the check.
const offsetBefore = await settledScrollLeft(scroller);
await section.getByRole('button', { name: /Add to compare/ }).last().click();
await expect(
section.getByRole('button', { name: /Added to compare/ }).first(),
).toBeVisible();
const offsetAfter = await settledScrollLeft(scroller);
expect(Math.abs(offsetAfter - offsetBefore)).toBeLessThanOrEqual(8);
});
/**
* The nearby-schools section at MOBILE.md's three reference widths.
*
* MOBILE.md asks for exactly this check and records that it was not written
* because "Playwright isn't currently in the project dependency set". That is
* no longer true — this suite is Playwright — so the check exists now, scoped
* to the page this feature touches.
*/
for (const width of [360, 390, 430]) {
test(`nearby schools survives a ${width}px viewport`, async ({ page }) => {
await page.setViewportSize({ width, height: 800 });
await searchByName(page, 'Primary');
await schoolLinks(page).first().click();
await page.waitForURL(/\/school\//);
const section = page.locator('#nearby');
if ((await section.count()) === 0) {
test.skip(true, 'No qualifying similar schools for this school');
}
// 1. Nothing bleeds past the right edge.
expect(
await page.evaluate(() => document.documentElement.scrollWidth - window.innerWidth),
).toBe(0);
// 2. No arrows on touch widths — swiping does the job, and they would take
// 96px from a 328px card.
await expect(section.getByRole('button', { name: 'More schools' })).toHaveCount(0);
// 3. Every tap target in the section clears 44px. A card title's own box is
// shorter, but its hit area is the whole card via ::after.
const failing = await section.evaluate((root: HTMLElement) =>
Array.from(root.querySelectorAll('a, button'))
.filter((el) => (el as HTMLElement).offsetParent)
.map((el) => {
const card = el.closest('li');
const box = el.matches('h3 a') && card
? card.getBoundingClientRect()
: el.getBoundingClientRect();
return {
text: (el as HTMLElement).innerText.trim().slice(0, 24),
w: box.width,
h: box.height,
};
})
.filter((o) => o.w < 44 || o.h < 44),
);
expect(failing).toEqual([]);
});
}
+41
View File
@@ -0,0 +1,41 @@
import { test, expect, Route } from '@playwright/test';
test('the deployed frontend and backend report the tested build', async ({ request }) => {
const response = await request.get('/release.json');
expect(response.ok()).toBeTruthy();
expect(response.headers()['cache-control']).toContain('no-store');
const identity = await response.json();
expect(identity.frontend).toEqual(identity.backend);
expect(identity.frontend.sha).toMatch(/^[a-f0-9]{40}$/);
expect(identity.frontend.build_id).toMatch(/^[a-f0-9]{32}$/);
if (process.env.EXPECTED_SHA) expect(identity.frontend.sha).toBe(process.env.EXPECTED_SHA);
if (process.env.EXPECTED_BUILD_ID) expect(identity.frontend.build_id).toBe(process.env.EXPECTED_BUILD_ID);
});
test('changing search while loading another page does not append old results', async ({ page }) => {
await page.goto('/?phase=primary');
await expect(page.getByRole('button', { name: 'Load more schools' })).toBeVisible();
let received!: (route: Route) => void;
const pending = new Promise<Route>(resolve => { received = resolve; });
await page.route('**/api/schools?**', async route => {
if (new URL(route.request().url()).searchParams.get('page') === '2') {
received(route);
return;
}
await route.continue();
});
await page.getByRole('button', { name: 'Load more schools' }).click();
const oldRequest = await pending;
const search = page.getByPlaceholder('School name or postcode').first();
await search.fill('secondary');
await search.press('Enter');
await page.waitForURL(/search=secondary/);
// A cancelled fetch may prevent route fulfilment altogether; either way,
// this deliberately late response must not become part of the new results.
await oldRequest.fulfill({ json: {
schools: [{ urn: 999998, school_name: 'P1 stale result sentinel', phase: 'Primary' }],
total: 2, page: 2, page_size: 1, total_pages: 2,
} }).catch(() => {});
await expect(page.getByText('P1 stale result sentinel')).toHaveCount(0);
await expect(page.getByRole('button', { name: 'Loading...' })).toHaveCount(0);
});
+373
View File
@@ -0,0 +1,373 @@
<!DOCTYPE html>
<html lang="en">
<head>
<meta charset="utf-8">
<meta name="viewport" content="width=device-width, initial-scale=1">
<title>Similar Schools Nearby</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link href="https://fonts.googleapis.com/css2?family=Inter:wght@400;500;600&family=Manrope:wght@500;600;700&display=swap" rel="stylesheet">
<style>
/* Tokens copied verbatim from nextjs-app/app/(frontend)/globals.css so this
mockup cannot drift from the shipped palette. Light values first, dark
under prefers-color-scheme, both overridable by the theme switch. */
:root {
color-scheme: light;
--bg-primary:#FAFAF8; --bg-secondary:#F5EFE6; --bg-card:#FFFFFF;
--text-primary:#1C2731; --text-secondary:#4A5560; --text-muted:#5F6A75;
--border:#E5E7EB; --border-strong:#D3D7DD;
--brand:#0F766E; --brand-strong:#0C5F58; --brand-bg:rgba(15,118,110,.10); --brand-on:#FFFFFF;
--action:#BE3C27; --action-strong:#A33320; --action-on:#FFFFFF;
--sand:#F5EFE6;
--font-display:Manrope,-apple-system,BlinkMacSystemFont,sans-serif;
--font-ui:Inter,-apple-system,BlinkMacSystemFont,sans-serif;
--radius-md:8px; --radius-lg:16px;
--shadow:0 1px 2px rgba(28,39,49,.06),0 1px 3px rgba(28,39,49,.05);
}
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) {
color-scheme: dark;
--bg-primary:#111A20; --bg-secondary:#16222A; --bg-card:#18242C;
--text-primary:#E9EEF0; --text-secondary:#B4C2C7; --text-muted:#8B9AA1;
--border:#26343D; --border-strong:#35454F;
--brand:#5FC7BB; --brand-strong:#7BD6CC; --brand-bg:rgba(95,199,187,.14); --brand-on:#0A1418;
--action:#F08A72; --action-strong:#F5A492; --action-on:#241009;
--sand:#1B2730;
--shadow:0 1px 2px rgba(0,0,0,.3),0 1px 3px rgba(0,0,0,.25);
}
}
:root[data-theme="dark"] {
color-scheme: dark;
--bg-primary:#111A20; --bg-secondary:#16222A; --bg-card:#18242C;
--text-primary:#E9EEF0; --text-secondary:#B4C2C7; --text-muted:#8B9AA1;
--border:#26343D; --border-strong:#35454F;
--brand:#5FC7BB; --brand-strong:#7BD6CC; --brand-bg:rgba(95,199,187,.14); --brand-on:#0A1418;
--action:#F08A72; --action-strong:#F5A492; --action-on:#241009;
--sand:#1B2730;
--shadow:0 1px 2px rgba(0,0,0,.3),0 1px 3px rgba(0,0,0,.25);
}
* { box-sizing:border-box; }
body {
margin:0; padding:32px 16px 80px; background:var(--bg-primary);
color:var(--text-primary); font:15px/1.55 var(--font-ui);
-webkit-font-smoothing:antialiased;
}
.page { max-width:960px; margin:0 auto; }
.page > header { margin-bottom:28px; display:flex; flex-wrap:wrap; gap:16px; align-items:flex-start; justify-content:space-between; }
.page > header h1 { font:700 25px/1.25 var(--font-display); letter-spacing:-.6px; margin:0 0 6px; }
.page > header p { margin:0; color:var(--text-muted); font-size:14px; max-width:60ch; }
.theme-switch { border:1px solid var(--border-strong); background:var(--bg-card); color:var(--text-secondary); border-radius:999px; padding:8px 14px; font:500 13px var(--font-ui); cursor:pointer; min-height:44px; }
.theme-switch:hover { border-color:var(--brand); color:var(--brand); }
/* ── The page context each variant is shown inside ─────────────────── */
.variant { margin-bottom:40px; }
.variant > .context { padding:0 4px 12px; }
.variant .eyebrow { margin:0 0 4px; font-size:12px; letter-spacing:.04em; text-transform:uppercase; color:var(--text-muted); }
.variant .context h2 { font:600 18px/1.35 var(--font-display); margin:0; color:var(--text-secondary); }
.variant .note { margin:10px 4px 0; font-size:12.5px; color:var(--text-muted); }
.variant .note b { color:var(--text-secondary); font-weight:600; }
/* ── The section itself — mirrors components/school/Section ────────── */
.card {
background:var(--bg-card); border:1px solid var(--border);
border-radius:var(--radius-lg); padding:28px; box-shadow:var(--shadow);
}
.top { display:flex; align-items:flex-start; justify-content:space-between; gap:16px; }
.top h2 { font:700 22px/1.25 var(--font-display); letter-spacing:-.4px; margin:0; }
.lede { margin:8px 0 20px; color:var(--text-secondary); font-size:14.5px; max-width:64ch; }
/* ── Carousel ───────────────────────────────────────────────────────
Every card is in the DOM and in the initial HTML — the arrows scroll a
list, they do not swap a view. That keeps all six links crawlable and
keeps the section usable with no JavaScript, where it degrades to a
plain horizontally scrollable row. */
.arrows { display:flex; gap:8px; flex:none; }
.arrow {
width:44px; height:44px; display:grid; place-items:center; cursor:pointer;
border:1px solid var(--border-strong); border-radius:999px;
background:var(--bg-card); color:var(--brand);
}
.arrow:hover:not(:disabled) { border-color:var(--brand); background:var(--brand-bg); }
.arrow:disabled { opacity:.35; cursor:default; }
.arrow:focus-visible { outline:2px solid var(--brand); outline-offset:2px; }
.arrow svg { width:17px; height:17px; }
.scroller {
display:grid; grid-auto-flow:column;
grid-auto-columns:calc((100% - 28px) / 3);
gap:14px; overflow-x:auto; scroll-snap-type:x mandatory;
padding:2px; margin:-2px; /* room for focus rings */
scrollbar-width:none; -ms-overflow-style:none;
list-style:none;
}
.scroller::-webkit-scrollbar { display:none; }
.scroller:focus-visible { outline:2px solid var(--brand); outline-offset:4px; border-radius:var(--radius-md); }
@media (max-width:820px) { .scroller { grid-auto-columns:calc((100% - 14px) / 2); } }
/* Touch widths: the arrows would squeeze the lede into a four-line column for a
control that swiping already provides, so they go and the documented
right-edge fade carries the affordance instead (MOBILE.md). The fade lifts at
the end of the travel, where there is nothing more to hint at. */
@media (max-width:640px) {
.top { display:block; }
.arrows { display:none; }
.scroller { grid-auto-columns:86%; mask-image:linear-gradient(to right, #000 calc(100% - 28px), transparent); }
.scroller[data-at-end=true] { mask-image:none; }
.card { padding:20px; }
}
.school {
position:relative; display:flex; flex-direction:column; scroll-snap-align:start;
border:1px solid var(--border); border-radius:var(--radius-md);
padding:16px; background:var(--bg-card);
}
.school:has(.add[aria-pressed=true]) { border-color:var(--brand); background:var(--brand-bg); }
.distance { display:flex; align-items:center; gap:5px; font-size:12px; color:var(--text-muted); margin:0 0 10px; }
.distance svg { width:13px; height:13px; flex:none; }
.school h3 { font:600 16px/1.35 var(--font-display); margin:0 0 6px; }
/* The whole card is the link target; the button sits above it on z-index so
it stays independently clickable. */
.school h3 a { color:var(--text-primary); text-decoration:none; }
.school h3 a::after { content:""; position:absolute; inset:0; border-radius:var(--radius-md); }
.school:hover { border-color:var(--border-strong); }
.school h3 a:hover { color:var(--brand); text-decoration:underline; }
.school h3 a:focus-visible { outline:none; }
.school:has(h3 a:focus-visible) { outline:2px solid var(--brand); outline-offset:2px; }
.meta { margin:0 0 12px; font-size:12.5px; color:var(--text-muted); }
.shared { display:flex; flex-wrap:wrap; gap:6px; margin:0 0 14px; padding:0; list-style:none; }
.shared li { font-size:11.5px; line-height:1.4; padding:4px 8px; border-radius:999px; background:var(--brand-bg); color:var(--brand); border:1px solid transparent; }
.shared li.loose { background:transparent; color:var(--text-muted); border-color:var(--border); }
.metric { margin-top:auto; padding-top:13px; border-top:1px solid var(--border); }
.value { font:700 26px/1.1 var(--font-display); letter-spacing:-.6px; margin:0; }
.value.absent { font-size:15px; font-weight:600; color:var(--text-muted); letter-spacing:0; }
.metric .label { margin:4px 0 0; font-size:12px; color:var(--text-secondary); }
.metric .ref { margin:2px 0 0; font-size:12px; color:var(--text-muted); }
.add {
position:relative; z-index:1; margin-top:14px; width:100%; min-height:44px;
font:500 13px var(--font-ui); cursor:pointer; border-radius:var(--radius-md);
border:1px solid var(--border-strong); background:var(--bg-card); color:var(--brand);
}
.add:hover { border-color:var(--brand); background:var(--brand-bg); }
.add[aria-pressed=true] { border-color:var(--brand); background:var(--brand-bg); font-weight:600; }
.add:focus-visible { outline:2px solid var(--brand); outline-offset:2px; }
.footer {
display:flex; flex-wrap:wrap; align-items:center; justify-content:space-between;
gap:14px; margin-top:20px; padding-top:18px; border-top:1px solid var(--border);
}
.footer p { margin:0; font-size:12.5px; color:var(--text-muted); }
.footer strong { display:block; font:600 14px var(--font-ui); color:var(--text-primary); }
/* Coral: the one decisive action in this section, and there is only one. */
.compare {
min-height:44px; padding:0 20px; border-radius:var(--radius-md); cursor:pointer;
font:600 14px var(--font-ui); background:var(--action); color:var(--action-on);
border:1px solid var(--action); text-decoration:none; display:inline-flex; align-items:center; gap:8px;
}
.compare:hover { background:var(--action-strong); border-color:var(--action-strong); }
.compare[aria-disabled=true] { opacity:.45; pointer-events:none; }
.caption { margin:16px 0 0; font-size:11.5px; color:var(--text-muted); }
.sr { position:absolute; width:1px; height:1px; padding:0; margin:-1px; overflow:hidden; clip:rect(0 0 0 0); white-space:nowrap; border:0; }
</style>
</head>
<body>
<div class="page">
<header>
<div>
<h1>Similar schools nearby</h1>
<p>A new section on the school detail page. Fictional schools and figures; shipped
colour, type and section shell taken from <code>globals.css</code>.</p>
</div>
<button class="theme-switch" type="button" id="theme">Dark theme</button>
</header>
<div id="variants"></div>
</div>
<script>
const PIN = '<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="M20 10c0 6-8 12-8 12s-8-6-8-12a8 8 0 0 1 16 0Z"/><circle cx="12" cy="10" r="3"/></svg>';
const CHEV = (dir) => `<svg viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" aria-hidden="true"><path d="${dir === 'prev' ? 'M15 18l-6-6 6-6' : 'M9 18l6-6-6-6'}"/></svg>`;
const variants = [
{
id: 'dense',
eyebrow: 'Variant 1 · Dense urban primary — six matches, carousel active',
context: 'Meadowbrook Primary School — Ages 4–11 · Mixed · No religious character · Community school',
lede: 'Other primary schools near Meadowbrook Primary School, with a similar intake.',
metric: 'Reading, writing & maths',
caption: 'Meeting the expected standard at key stage 2, 2025.',
thisValue: '72%',
note: 'Fourteen schools cleared <b>tier 1</b> within three miles, so the section takes the six nearest and stops there. The arrows scroll a list that is entirely in the HTML — all six links are crawlable, and with JavaScript off the row still scrolls.',
schools: [
{ name:'Willow Lane Primary School', distance:'0.4', meta:'Community school · Ages 4–11', shared:['Mixed','No religious character'], value:'74%' },
{ name:'Oakfield Primary School', distance:'0.6', meta:'Academy converter · Ages 3–11', shared:['Mixed','No religious character'], value:'69%' },
{ name:'Brookside Primary School', distance:'0.9', meta:'Community school · Ages 4–11', shared:['Mixed','No religious character'], value:'Not published' },
{ name:'Hollytree Primary School', distance:'1.3', meta:'Academy converter · Ages 4–11', shared:['Mixed','No religious character'], value:'81%' },
{ name:'Marsh Green Primary School', distance:'1.8', meta:'Community school · Ages 3–11', shared:['Mixed','No religious character'], value:'64%' },
{ name:'Kingsway Primary School', distance:'2.2', meta:'Foundation school · Ages 4–11', shared:['Mixed','No religious character'], value:'77%' },
],
},
{
id: 'secondary',
eyebrow: 'Variant 2 · Secondary — four matches, mixed tiers',
context: 'Meadowbrook High School — Ages 11–18 · Mixed · Non-selective · Academy',
lede: 'Other secondary schools near Meadowbrook High School, with a similar intake.',
metric: 'Attainment 8',
caption: 'Average GCSE attainment score across eight qualifications, 2025.',
thisValue: '51.2',
note: 'Only two schools cleared tier 1, so the search widened to <b>tier 2</b> and found two more. It stops there rather than widening again to reach six — tiers relax to reach a usable set, never to fill the last slots. Selectivity never relaxes, so no grammar school can appear here.',
schools: [
{ name:'Rivermead High School', distance:'0.9', meta:'Academy converter · Ages 11–18', shared:['Mixed','Non-selective','No religious character'], value:'52.8' },
{ name:'Oakfield Academy', distance:'1.7', meta:'Academy sponsor led · Ages 11–16', shared:['Mixed','Non-selective','No religious character'], value:'49.6' },
{ name:'St Aidan’s Catholic High School', distance:'2.4', meta:'Voluntary aided · Ages 11–18', shared:['Mixed','Non-selective'], value:'53.4' },
{ name:'Parkside Community School', distance:'3.8', meta:'Community school · Ages 11–16', shared:['Mixed','Non-selective'], value:'50.9' },
],
},
{
id: 'sparse',
eyebrow: 'Variant 3 · Rural — two matches, no arrows',
context: 'Little Ashby Church of England Primary School — Ages 4–11 · Mixed · Church of England · Voluntary controlled',
lede: 'Other primary schools near Little Ashby Church of England Primary School.',
metric: 'Reading, writing & maths',
caption: 'Meeting the expected standard at key stage 2, 2025.',
thisValue: '66%',
note: 'Nothing matched on religious character within range. At <b>tier 3</b> the lede drops the phrase “with a similar intake” and the chips fall back to the plain phase. Two cards fit the row, so the arrows are not rendered at all. One school fewer and the section would not render either.',
schools: [
{ name:'Great Marden Primary School', distance:'4.2', meta:'Community school · Ages 4–11', shared:['Primary school'], loose:true, value:'71%' },
{ name:'Ashby Vale Academy', distance:'7.8', meta:'Academy converter · Ages 4–11', shared:['Primary school'], loose:true, value:'58%' },
],
},
];
const selected = Object.fromEntries(variants.map((v) => [v.id, new Set()]));
function card(v, s, i) {
const on = selected[v.id].has(i);
const absent = s.value === 'Not published';
return `
<li class="school">
<p class="distance">${PIN}${s.distance} miles away</p>
<h3><a href="#">${s.name}</a></h3>
<p class="meta">${s.meta}</p>
<ul class="shared">${s.shared.map((c) => `<li class="${s.loose ? 'loose' : ''}">${c}</li>`).join('')}</ul>
<div class="metric">
<p class="value ${absent ? 'absent' : ''}">${s.value}</p>
<p class="label">${v.metric}</p>
<p class="ref">${v.thisValue} at this school</p>
</div>
<button class="add" type="button" data-variant="${v.id}" data-index="${i}" aria-pressed="${on}">
${on ? '✓ Added to compare' : '+ Add to compare'}<span class="sr"> — ${s.name}</span>
</button>
</li>`;
}
function render() {
document.getElementById('variants').innerHTML = variants.map((v) => {
const count = selected[v.id].size;
// Three fit the row, so anything more is what the arrows are for.
const scrollable = v.schools.length > 3;
return `
<section class="variant">
<div class="context">
<p class="eyebrow">${v.eyebrow}</p>
<h2>${v.context}</h2>
</div>
<div class="card">
<div class="top">
<div>
<h2 id="h-${v.id}">Similar schools nearby</h2>
<p class="lede">${v.lede}</p>
</div>
${scrollable ? `<div class="arrows">
<button class="arrow" type="button" data-scroll="prev" data-variant="${v.id}" aria-label="Previous schools" aria-controls="sc-${v.id}">${CHEV('prev')}</button>
<button class="arrow" type="button" data-scroll="next" data-variant="${v.id}" aria-label="More schools" aria-controls="sc-${v.id}">${CHEV('next')}</button>
</div>` : ''}
</div>
<ul class="scroller" id="sc-${v.id}" ${scrollable ? `tabindex="0" role="group" aria-labelledby="h-${v.id}"` : ''}>
${v.schools.map((s, i) => card(v, s, i)).join('')}
</ul>
<div class="footer">
<p aria-live="polite">
<strong>${count ? `${count} school${count === 1 ? '' : 's'} selected` : 'Compare side by side'}</strong>
${count ? 'This school is included automatically.' : 'Add a school to compare it with this one.'}
</p>
<a class="compare" href="#" aria-disabled="${count ? 'false' : 'true'}">
${count ? `Compare ${count + 1} schools` : 'Compare'} →
</a>
</div>
<p class="caption">Distances are straight-line from this school, not road distance.
${v.caption} Fictional schools and figures for this mockup.</p>
</div>
<p class="note">${v.note}</p>
</section>`;
}).join('');
variants.forEach((v) => {
const scroller = document.getElementById(`sc-${v.id}`);
if (scroller) syncArrows(v.id, scroller);
});
}
/** An arrow that scrolls nowhere is a dead control, so each end disables its own.
*
* EDGE is not paranoia. The scroller carries 2px of padding so focus rings are
* not clipped, and scroll-snap treats that padding as the first card's snap
* position — so a scroller sitting at its start reports scrollLeft 2, not 0.
* Sub-pixel rounding at other zoom levels moves it again. Testing against an
* exact 0 leaves the back arrow live at the start, pointing nowhere. */
const EDGE = 8;
function syncArrows(id, scroller) {
const max = scroller.scrollWidth - scroller.clientWidth;
const atStart = scroller.scrollLeft <= EDGE;
const atEnd = scroller.scrollLeft >= max - EDGE;
// Drives the mobile scroll-fade, so it is computed even where no arrow is
// rendered to consume it.
scroller.dataset.atEnd = String(atEnd);
const prev = document.querySelector(`.arrow[data-scroll="prev"][data-variant="${id}"]`);
const next = document.querySelector(`.arrow[data-scroll="next"][data-variant="${id}"]`);
if (!prev || !next) return;
prev.disabled = atStart;
next.disabled = atEnd;
}
document.getElementById('variants').addEventListener('click', (event) => {
const arrow = event.target.closest('.arrow');
if (arrow) {
const scroller = document.getElementById(`sc-${arrow.dataset.variant}`);
// A page is what the reader can see, so the viewport is the step.
scroller.scrollBy({ left: (arrow.dataset.scroll === 'next' ? 1 : -1) * scroller.clientWidth, behavior: 'smooth' });
return;
}
const button = event.target.closest('.add');
if (!button) return;
const { variant, index } = button.dataset;
const set = selected[variant];
const i = Number(index);
// Scroll position is DOM state, not React state; keep it across the re-render.
const offset = document.getElementById(`sc-${variant}`).scrollLeft;
set.has(i) ? set.delete(i) : set.add(i);
render();
const scroller = document.getElementById(`sc-${variant}`);
scroller.scrollLeft = offset;
syncArrows(variant, scroller);
document.querySelector(`.add[data-variant="${variant}"][data-index="${index}"]`).focus();
}, true);
document.getElementById('variants').addEventListener('scroll', (event) => {
const scroller = event.target.closest('.scroller');
if (scroller) syncArrows(scroller.id.replace('sc-', ''), scroller);
}, true);
const themeButton = document.getElementById('theme');
themeButton.addEventListener('click', () => {
const dark = document.documentElement.dataset.theme === 'dark';
document.documentElement.dataset.theme = dark ? 'light' : 'dark';
themeButton.textContent = dark ? 'Dark theme' : 'Light theme';
});
render();
</script>
</body>
</html>
+12 -4
View File
@@ -1,8 +1,16 @@
# API Configuration
NEXT_PUBLIC_API_URL=http://localhost:8000/api
# Browser requests use the same-origin Next.js proxy.
NEXT_PUBLIC_API_URL=/api
# Production API URL (for deployment)
# NEXT_PUBLIC_API_URL=https://api.schoolcompare.co.uk/api
# Absolute URL for server-side fetching and the proxy; include /api.
# In the managed container network this is http://backend:80/api (staging differs).
FASTAPI_URL=http://localhost:8000/api
# Payload CMS runtime configuration. Use the managed environment's database;
# Payload owns the payload schema, independently of the school marts.
DATABASE_URL=postgresql://schoolcompare:CHANGE_THIS_PASSWORD@localhost:5432/schoolcompare
# Generate a secret: python -c "import secrets; print(secrets.token_urlsafe(32))"
# Use distinct secrets for staging and production.
PAYLOAD_SECRET=CHANGE_THIS_TO_A_SECURE_RANDOM_SECRET
# Node Environment
NODE_ENV=development
+1
View File
@@ -39,3 +39,4 @@ yarn-error.log*
# typescript
*.tsbuildinfo
next-env.d.ts
+12 -288
View File
@@ -1,291 +1,15 @@
# Deployment Guide
# Frontend deployment
This guide covers deployment options for the SchoolCompare Next.js application.
Next.js and Payload run in the same frontend container. The maintained deployment
procedure is [docs/DEPLOY.md](../docs/DEPLOY.md), with the production and staging
Portainer compose files at the repository root.
## Deployment Options
The frontend Dockerfile builds a standalone Next.js image. Runtime configuration
supplies `FASTAPI_URL`, `DATABASE_URL` and `PAYLOAD_SECRET`; uploaded CMS media is
persisted in a volume. Promote the built image through the repository's Gitea
workflow after human staging approval.
### Option 1: Vercel (Recommended for Next.js)
Vercel is the easiest and most optimized platform for Next.js applications.
#### Steps:
1. **Install Vercel CLI**:
```bash
npm install -g vercel
```
2. **Login to Vercel**:
```bash
vercel login
```
3. **Deploy**:
```bash
vercel --prod
```
4. **Configure Environment Variables** in Vercel dashboard:
- `NEXT_PUBLIC_API_URL`: Your FastAPI endpoint (e.g., `https://api.schoolcompare.co.uk/api`)
- `FASTAPI_URL`: Same as above for server-side requests
#### Benefits:
- Automatic HTTPS
- Global CDN
- Zero-config deployment
- Automatic preview deployments
- Built-in analytics
---
### Option 2: Docker (Self-hosted)
Deploy using Docker containers for full control.
#### Prerequisites:
- Docker 20+
- Docker Compose 2+
#### Steps:
1. **Build Docker Image**:
```bash
docker build -t schoolcompare-nextjs:latest .
```
2. **Run with Docker Compose**:
```bash
# Create .env file with production variables
echo "NEXT_PUBLIC_API_URL=https://api.schoolcompare.co.uk/api" > .env
echo "FASTAPI_URL=http://backend:8000/api" >> .env
# Start services
docker-compose up -d
```
3. **Verify Deployment**:
```bash
curl http://localhost:3000
```
#### Environment Variables:
- `NEXT_PUBLIC_API_URL`: Public API endpoint (client-side)
- `FASTAPI_URL`: Internal API endpoint (server-side)
- `NODE_ENV`: `production`
---
### Option 3: PM2 (Node.js Process Manager)
Deploy directly on a Node.js server using PM2.
#### Prerequisites:
- Node.js 24+
- PM2 (`npm install -g pm2`)
#### Steps:
1. **Build Application**:
```bash
npm run build
```
2. **Create PM2 Ecosystem File** (`ecosystem.config.js`):
```javascript
module.exports = {
apps: [{
name: 'schoolcompare-nextjs',
script: 'npm',
args: 'start',
cwd: '/path/to/nextjs-app',
instances: 'max',
exec_mode: 'cluster',
env: {
NODE_ENV: 'production',
PORT: 3000,
NEXT_PUBLIC_API_URL: 'https://api.schoolcompare.co.uk/api',
FASTAPI_URL: 'http://localhost:8000/api',
},
}],
};
```
3. **Start with PM2**:
```bash
pm2 start ecosystem.config.js
pm2 save
pm2 startup
```
---
### Option 4: Nginx Reverse Proxy
Use Nginx as a reverse proxy in front of Next.js.
#### Nginx Configuration:
```nginx
server {
listen 80;
server_name schoolcompare.co.uk;
# Redirect to HTTPS
return 301 https://$server_name$request_uri;
}
server {
listen 443 ssl http2;
server_name schoolcompare.co.uk;
# SSL Configuration
ssl_certificate /etc/ssl/certs/schoolcompare.crt;
ssl_certificate_key /etc/ssl/private/schoolcompare.key;
# Security Headers
# frame-ancestors replaces X-Frame-Options so the analytics subdomain
# (Umami heatmap/recorder) can embed the site in an iframe.
add_header Content-Security-Policy "frame-ancestors 'self' https://analytics.schoolcompare.co.uk" always;
add_header X-Content-Type-Options "nosniff" always;
add_header X-XSS-Protection "1; mode=block" always;
# Proxy to Next.js
location / {
proxy_pass http://localhost:3000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection 'upgrade';
proxy_set_header Host $host;
proxy_cache_bypass $http_upgrade;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
# Proxy to FastAPI
location /api/ {
proxy_pass http://localhost:8000;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
# Cache static files
location /_next/static/ {
proxy_pass http://localhost:3000;
add_header Cache-Control "public, max-age=31536000, immutable";
}
}
```
---
## Pre-Deployment Checklist
- [ ] Run `npm run build` successfully
- [ ] Run `npm test` - all tests pass
- [ ] Environment variables configured
- [ ] FastAPI backend accessible
- [ ] Database migrations applied
- [ ] SSL certificates configured (production)
- [ ] Domain DNS configured
- [ ] Monitoring/logging set up
- [ ] Backup strategy in place
---
## Post-Deployment Verification
1. **Health Check**:
```bash
curl https://schoolcompare.co.uk
```
2. **Test Routes**:
- Home: `https://schoolcompare.co.uk/`
- School Page: `https://schoolcompare.co.uk/school/100001`
- Compare: `https://schoolcompare.co.uk/compare`
- Rankings: `https://schoolcompare.co.uk/rankings`
3. **Check SEO**:
- Sitemap: `https://schoolcompare.co.uk/sitemap.xml`
- Robots: `https://schoolcompare.co.uk/robots.txt`
4. **Performance Audit**:
- Run Lighthouse in Chrome DevTools
- Target scores: 90+ for Performance, Accessibility, Best Practices, SEO
---
## Monitoring
### Recommended Tools:
- **Vercel Analytics** (if using Vercel)
- **Sentry** for error tracking
- **Google Analytics** for user analytics
- **Uptime Robot** for uptime monitoring
### Health Check Endpoint:
The application automatically serves health data at the root route.
---
## Rollback Procedure
### Vercel:
```bash
vercel rollback
```
### Docker:
```bash
docker-compose down
docker-compose up -d --force-recreate
```
### PM2:
```bash
pm2 stop schoolcompare-nextjs
# Restore previous build
pm2 start schoolcompare-nextjs
```
---
## Troubleshooting
### Issue: API requests failing
- **Solution**: Check `NEXT_PUBLIC_API_URL` and `FASTAPI_URL` environment variables
- **Verify**: FastAPI backend is accessible from Next.js container/server
### Issue: Build fails
- **Solution**: Check Node.js version (requires 24+)
- **Clear cache**: `rm -rf .next node_modules && npm install && npm run build`
### Issue: Slow page loads
- **Solution**: Enable caching in API calls
- **Check**: Network latency to FastAPI backend
- **Verify**: CDN is serving static assets
---
## Security Considerations
- ✅ HTTPS enabled
- ✅ Security headers configured (X-Frame-Options, CSP, etc.)
- ✅ API keys in environment variables (never in code)
- ✅ CORS properly configured
- ✅ Rate limiting on API endpoints
- ✅ Regular security updates
- ✅ Dependency vulnerability scanning
---
## Support
For deployment issues, contact the DevOps team or refer to:
- [Next.js Deployment Docs](https://nextjs.org/docs/deployment)
- [Vercel Documentation](https://vercel.com/docs)
- [Docker Documentation](https://docs.docker.com/)
Earlier Vercel and standalone deployment recipes have been retired from this file
because they do not describe the current CMS, persistence and promotion setup.
See [development](../docs/DEVELOPMENT.md) for checks and
[publishing](docs/PUBLISHING.md) for CMS operations.
+17
View File
@@ -28,6 +28,10 @@ ENV NODE_ENV=production
ARG FASTAPI_URL=http://backend:80/api
ENV FASTAPI_URL=${FASTAPI_URL}
ARG BUILD_SHA=development
ARG BUILD_ID=development
RUN node -e 'require("fs").writeFileSync("build-info.json", JSON.stringify({sha:process.argv[1],build_id:process.argv[2]}))' "$BUILD_SHA" "$BUILD_ID"
# Build application
RUN npm run build
@@ -53,6 +57,13 @@ COPY --from=builder /app/.next/static ./.next/static
# a miss here is a silent 500 on /opengraph-image, not a build failure.
COPY --from=builder /app/assets ./assets
# Payload writes uploads here, and the compose file mounts a named volume over
# it. The directory must exist and be owned by the runtime user BEFORE the
# mount: Docker seeds a fresh named volume from the image path, so a missing or
# root-owned directory here makes every upload fail with EACCES at runtime,
# long after the build passed. The chown below covers it.
RUN mkdir -p /app/media
# Set correct permissions
RUN chown -R nextjs:nodejs /app
@@ -63,6 +74,12 @@ USER nextjs
EXPOSE 3000
# Set environment variables
ARG BUILD_SHA=development
ARG BUILD_ID=development
LABEL io.schoolcompare.build-id=$BUILD_ID
LABEL io.schoolcompare.commit=$BUILD_SHA
COPY --from=builder /app/build-info.json ./build-info.json
ENV PORT=3000
ENV HOSTNAME="0.0.0.0"
+43 -141
View File
@@ -1,156 +1,58 @@
# SchoolCompare Next.js Application
# SchoolCompare frontend and CMS
Modern Next.js application for comparing primary school KS2 performance across England.
Next.js App Router with React, TypeScript, CSS Modules, Chart.js, Leaflet and
Payload CMS. It serves school search, comparisons, rankings, school/place detail
pages and editorial content across England.
## Features
Start with the [repository overview](../README.md),
[architecture](../docs/ARCHITECTURE.md) and [development checks](../docs/DEVELOPMENT.md).
- **Server-Side Rendering (SSR)**: Fast initial page loads with pre-rendered content
- **Individual School Pages**: Dedicated pages for each school with full SEO optimization
- **Side-by-Side Comparison**: Compare up to 5 schools simultaneously
- **School Rankings**: Top-performing schools by various metrics
- **Interactive Maps**: Leaflet integration for geographic visualization
- **Performance Charts**: Chart.js visualizations for historical data
- **Responsive Design**: Mobile-first approach with full responsive support
- **SEO Optimized**: Dynamic sitemaps, meta tags, and structured data
## Source map
## Tech Stack
| Path | Purpose |
|---|---|
| `app/(frontend)/` | Public root layout, server pages and FastAPI proxy |
| `app/(payload)/` | Payload root layout, `/admin` and `/cms-api` |
| `app/robots.ts`, `app/opengraph-image.tsx`, root icons | Site-wide metadata endpoints |
| `components/` | Client views and reusable display components |
| `components/school/` | School detail sections |
| `lib/api.ts`, `lib/types.ts` | Fetch wrappers and manual school API types |
| `lib/schoolSections.ts`, `lib/compareLogic.ts` | Presentation decisions and data preparation |
| `context/`, `hooks/` | Comparison state, suggestion state and responsive behaviour |
| `collections/`, `blocks/`, `migrations/` | CMS schema and production migrations |
| `__tests__/` | Jest and React Testing Library tests |
- **Framework**: Next.js 16 (App Router)
- **Language**: TypeScript 5
- **Styling**: CSS Modules + CSS Variables
- **State Management**: React Context API + URL state
- **Data Fetching**: SWR (client-side) + Next.js fetch (server-side)
- **Charts**: Chart.js + react-chartjs-2
- **Maps**: Leaflet + react-leaflet
- **Testing**: Jest + React Testing Library
- **Validation**: Zod
Do not introduce a shared `app/layout.tsx`: public pages and Payload have separate
root layouts. Keep root metadata files outside the route groups.
## Getting Started
## Data and state
### Prerequisites
Server pages fetch initial data directly from `FASTAPI_URL`. Browser fetches use
`/api` by default, forwarded by `app/(frontend)/api/[...path]/route.ts`.
`FASTAPI_URL` must include `/api`. See `.env.example` for CMS and API settings.
- Node.js 24+ (using nvm recommended)
- FastAPI backend running on port 8000
State uses React hooks/context, URL search parameters and localStorage for the
comparison basket. SWR is not installed. Maps use dynamic Leaflet wrappers.
Revalidation intervals are configured in fetch wrappers and pages; they vary by
resource. Backend reloads do not automatically invalidate every Next.js cache.
### Installation
## Commands
```bash
# Install dependencies
npm install
# Copy environment variables
cp .env.example .env.local
# Update .env.local with your configuration
```
### Development
```bash
# Start development server
npm run dev
# Open http://localhost:3000
```
### Building
```bash
# Build for production
```sh
npm ci
npm run typecheck
npm test -- --runInBand
npm run build
# Start production server
npm start
```
### Testing
`test:watch` and `test:coverage` are also available. There is no `lint` script.
A running application needs the backend/data environment described in the
[development guide](../docs/DEVELOPMENT.md).
```bash
# Run tests
npm test
After CMS field or editor changes, run `npm run generate:importmap`. Keep
`payload-types.ts` generated from the CMS schema rather than editing it by hand.
The build must work without a database connection; avoid module-scope CMS queries
and DB-backed `generateStaticParams` functions.
# Run tests in watch mode
npm run test:watch
# Run tests with coverage
npm run test:coverage
```
### Linting
```bash
# Run ESLint
npm run lint
```
## Project Structure
```
nextjs-app/
├── app/ # App Router pages
│ ├── layout.tsx # Root layout
│ ├── page.tsx # Home page
│ ├── compare/ # Compare page
│ ├── rankings/ # Rankings page
│ ├── school/[urn]/ # Individual school pages
│ ├── sitemap.ts # Dynamic sitemap
│ └── robots.ts # Robots.txt
├── components/ # React components
│ ├── SchoolCard.tsx # School card component
│ ├── FilterBar.tsx # Search/filter controls
│ ├── ComparisonView.tsx # Comparison interface
│ ├── RankingsView.tsx # Rankings table
│ └── ...
├── lib/ # Utility libraries
│ ├── api.ts # API client
│ ├── types.ts # TypeScript types
│ └── utils.ts # Helper functions
├── hooks/ # Custom React hooks
├── context/ # React Context providers
├── styles/ # Global styles
├── public/ # Static assets
└── __tests__/ # Test files
```
## Environment Variables
| Variable | Description | Default |
|----------|-------------|---------|
| `NEXT_PUBLIC_API_URL` | Public API endpoint (client-side) | `http://localhost:8000/api` |
| `FASTAPI_URL` | Server-side API endpoint | `http://localhost:8000/api` |
| `NODE_ENV` | Environment mode | `development` |
## Performance Optimizations
- **Server-Side Rendering**: Initial HTML rendered on server
- **Static Generation**: Where possible, pages are pre-generated
- **Image Optimization**: Next.js Image component with AVIF/WebP support
- **Code Splitting**: Automatic route-based code splitting
- **Dynamic Imports**: Heavy components loaded on demand
- **API Caching**: Configurable revalidation for data fetching
- **Bundle Optimization**: Tree shaking and minification
- **Compression**: Gzip compression enabled
## SEO Features
- **Dynamic Meta Tags**: Generated per page with Next.js Metadata API
- **Open Graph**: Social media optimization
- **JSON-LD**: Structured data for search engines
- **Sitemap**: Auto-generated from database
- **Robots.txt**: Search engine crawling rules
- **Canonical URLs**: Duplicate content prevention
## Browser Support
- Chrome (latest)
- Firefox (latest)
- Safari (latest)
- Edge (latest)
## License
Proprietary - SchoolCompare
## Support
For issues and questions, please contact the development team.
See [publishing](docs/PUBLISHING.md) for CMS operations and
[deployment](../docs/DEPLOY.md) for staging and production promotion.
@@ -8,7 +8,7 @@
// environment provides — under jsdom this suite fails on import, not on an
// assertion.
import { NextRequest } from 'next/server';
import { GET } from '@/app/api/[...path]/route';
import { GET } from '@/app/(frontend)/api/[...path]/route';
function request(path: string) {
return new NextRequest(`http://localhost:3000/api/${path}`);
@@ -0,0 +1,34 @@
import { metadata } from '@/app/(frontend)/about/page';
import { personJsonLd, organizationJsonLd } from '@/lib/jsonld';
describe('/about metadata', () => {
it('canonicalises to the bare path', () => {
expect(metadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/about');
});
});
describe('author structured data', () => {
it('describes a Person with a first name and a photo', () => {
const person = personJsonLd();
expect(person['@type']).toBe('Person');
expect(person.name).toBe('Tudor');
expect(person.image).toBe('https://www.schoolcompare.co.uk/brand/tudor.jpg');
expect(person.url).toBe('https://www.schoolcompare.co.uk/about');
});
it('never publishes a surname or an employer', () => {
// Author identity constraint: first name only. A surname here would be
// the one place it leaks, since JSON-LD is machine-read and archived.
const serialised = JSON.stringify(personJsonLd());
expect(serialised).not.toMatch(/familyName|Sitaru/i);
expect(serialised).not.toMatch(/worksFor|affiliation/i);
});
it('describes the site as an Organization the Person authors for', () => {
const org = organizationJsonLd();
expect(org['@type']).toBe('Organization');
expect(org.name).toBe('schoolcompare');
expect(org.url).toBe('https://www.schoolcompare.co.uk');
});
});
@@ -0,0 +1,66 @@
/**
* The blog index imports getCachedPayload, which pulls in Payload — ESM-only,
* and next/jest will not transform node_modules. Mocking that one module keeps
* the page's metadata testable without loading the CMS; the mock is never
* called, because `metadata` is a static export evaluated at import time.
*/
jest.mock('@/lib/payload', () => ({ getCachedPayload: jest.fn() }));
import { metadata } from '@/app/(frontend)/blog/page';
import { blogPostingJsonLd, breadcrumbJsonLd } from '@/lib/jsonld';
const post = {
title: 'What the data cannot tell you',
slug: 'what-the-data-cannot-tell-you',
excerpt: 'Results describe one year group on a handful of days.',
publishedAt: '2026-09-15T00:00:00.000Z',
};
describe('/blog metadata', () => {
it('canonicalises to the bare path', () => {
expect(metadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/blog');
});
});
describe('BlogPosting structured data', () => {
it('names the same Person entity the about page declares', () => {
// By @id, not by repeating the person: search engines must resolve every
// post and the about page to one author entity, or the site has several.
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld['@type']).toBe('BlogPosting');
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk/about#tudor');
expect(ld.publisher['@id']).toBe('https://www.schoolcompare.co.uk#organization');
});
it('attributes to the organization when the about page is dark', () => {
/*
* The two flags are independent, so blog-on-about-off is a reachable
* state. The Person entity lives at /about#tudor and that URL 404s while
* the flag is dark, so claiming it would declare an author that resolves
* to nothing — worse for the blog's credibility than having no named
* author at all. Attribute to the publisher instead.
*/
const ld = blogPostingJsonLd(post, { namedAuthor: false });
expect(ld.author['@id']).toBe('https://www.schoolcompare.co.uk#organization');
expect(JSON.stringify(ld)).not.toContain('/about');
});
it('carries a self-referencing canonical url and the publish date', () => {
const ld = blogPostingJsonLd(post, { namedAuthor: true });
expect(ld.url).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
expect(ld.datePublished).toBe('2026-09-15T00:00:00.000Z');
});
});
describe('breadcrumbs', () => {
it('places the post under the blog index', () => {
const ld = breadcrumbJsonLd(post);
expect(ld.itemListElement[0].item).toBe('https://www.schoolcompare.co.uk/blog');
expect(ld.itemListElement[1].item).toBe(
'https://www.schoolcompare.co.uk/blog/what-the-data-cannot-tell-you',
);
});
});
@@ -0,0 +1,51 @@
import { fireEvent, render, screen } from '@testing-library/react';
import HomePage from '@/app/(frontend)/page';
import SchoolPage from '@/app/(frontend)/school/[slug]/page';
import ErrorPage from '@/app/(frontend)/error';
import { APIFetchError, fetchSchools, fetchFilters, fetchSchoolDetails } from '@/lib/api';
import { fetchPlace, fetchPlaces } from '@/lib/places';
jest.mock('@/lib/api', () => ({
...jest.requireActual('@/lib/api'),
fetchSchools: jest.fn(),
fetchSchoolDetails: jest.fn(),
fetchFilters: jest.fn(async () => ({})),
fetchDataInfo: jest.fn(async () => null),
fetchNationalAverages: jest.fn(async () => null),
}));
jest.mock('@/lib/flags', () => ({ getFlags: jest.fn(async () => ({})) }));
jest.mock('next/navigation', () => ({
notFound: () => { throw new Error('NEXT_NOT_FOUND'); },
redirect: jest.fn(),
}));
const realFetch = global.fetch;
afterEach(() => { global.fetch = realFetch; jest.clearAllMocks(); });
test('school outages propagate; only a real 404 becomes not found', async () => {
const request = { params: Promise.resolve({ slug: '100001-school' }) };
const outage = new APIFetchError('unavailable', 503);
jest.mocked(fetchSchoolDetails).mockRejectedValueOnce(outage);
await expect(SchoolPage(request)).rejects.toBe(outage);
jest.mocked(fetchSchoolDetails).mockRejectedValueOnce(new APIFetchError('missing', 404));
await expect(SchoolPage(request)).rejects.toThrow('NEXT_NOT_FOUND');
});
test('homepage search failure is not returned as an empty successful page', async () => {
jest.mocked(fetchSchools).mockRejectedValueOnce(new APIFetchError('unavailable', 503));
await expect(HomePage({ searchParams: Promise.resolve({ search: 'school' }) })).rejects.toThrow('unavailable');
});
test('a place is absent only on 404; other failures propagate', async () => {
global.fetch = jest.fn().mockResolvedValue({ ok: false, status: 404 });
await expect(fetchPlace('town', 'example')).resolves.toBeNull();
jest.mocked(global.fetch).mockResolvedValue({ ok: false, status: 503 } as Response);
await expect(fetchPlace('town', 'example')).rejects.toMatchObject({ status: 503 });
await expect(fetchPlaces()).rejects.toMatchObject({ status: 503 });
});
test('the error boundary offers a retry without showing an empty search', () => {
const reset = jest.fn();
render(<ErrorPage error={new Error('offline')} reset={reset} />);
fireEvent.click(screen.getByRole('button', { name: 'Try again' }));
expect(reset).toHaveBeenCalledTimes(1);
});
+43 -4
View File
@@ -1,7 +1,8 @@
import { metadata as homeMetadata } from '@/app/page';
import { metadata as rankingsMetadata } from '@/app/rankings/page';
import { metadata as admissionsMetadata } from '@/app/admissions/page';
import { generateMetadata as compareMetadata } from '@/app/compare/page';
import { metadata as homeMetadata } from '@/app/(frontend)/page';
import { metadata as rankingsMetadata } from '@/app/(frontend)/rankings/page';
import { metadata as admissionsMetadata } from '@/app/(frontend)/admissions/page';
import { generateMetadata as compareMetadata } from '@/app/(frontend)/compare/page';
import { metadata as rootMetadata } from '@/app/(frontend)/layout';
describe('canonical URLs', () => {
it('the homepage canonicalises to the bare root', () => {
@@ -128,3 +129,41 @@ describe('C1 snippet copy', () => {
}
});
});
/**
* The share card must be declared, not inherited.
*
* `app/opengraph-image.tsx` is a metadata file convention, and it does attach
* to routes in the app root segment — `_not-found` gets an og:image from it.
* It does NOT attach to the site's pages, which live in the `(frontend)`
* route group whose own layout is a root layout. Staging served og:title,
* og:description, og:url, og:site_name and og:type and no og:image at all,
* so every link pasted into a chat rendered bare.
*
* The file stays at the app root, because /robots.txt and /icon.png depend on
* it being there. The site's root layout points at the route it generates.
*/
describe('the share card', () => {
it('declares an opengraph image on the site root layout', () => {
// No og:image means every link pasted into a chat renders bare.
const images = rootMetadata.openGraph?.images;
expect(images).toBeTruthy();
expect(JSON.stringify(images)).toContain('/opengraph-image');
});
it('declares a twitter image too', () => {
// twitter.card is summary_large_image. Claiming a large-image card and
// supplying no image is worse than claiming a summary card.
// Metadata['twitter'] is a union and `card` is not on every member, so
// this reads the serialised shape rather than narrowing the type.
const twitter = JSON.stringify(rootMetadata.twitter);
expect(twitter).toContain('summary_large_image');
expect(twitter).toContain('/opengraph-image');
});
it('resolves the card to an absolute url via metadataBase', () => {
// The e2e journey does `new URL(ogUrl)`, which throws on a relative path.
expect(rootMetadata.metadataBase?.toString())
.toBe('https://www.schoolcompare.co.uk/');
});
});
@@ -0,0 +1,66 @@
/**
* next.config.mjs carries the staging noindex rule. Breaking it turns
* stx.schoolcompare.co.uk into a fully crawlable duplicate of production,
* and nothing else in the suite would notice.
*
* The non-null assertions are deliberate: every key asserted here is optional
* on NextConfig, and a missing one is precisely the regression under test, so
* the assertion below should fail the test rather than the compile.
*/
import nextConfig from '@/next.config.mjs';
async function headerRules() {
return nextConfig.headers!();
}
describe('next.config.mjs', () => {
it('keeps the staging host out of the index', async () => {
const headers = await headerRules();
const stagingRule = headers.find((rule) =>
rule.has?.some(
(cond) => cond.type === 'host' && cond.value === 'stx.schoolcompare.co.uk',
),
);
expect(stagingRule).toBeDefined();
expect(stagingRule!.headers).toContainEqual({
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
});
});
it('still emits standalone output for the Docker runner', () => {
expect(nextConfig.output).toBe('standalone');
});
it('still traces the share-card fonts into the standalone bundle', () => {
expect(nextConfig.outputFileTracingIncludes!['/opengraph-image']).toEqual([
'./assets/**',
]);
});
it('still allows the analytics subdomain to frame the site', async () => {
const headers = await headerRules();
const csp = headers
.flatMap((rule) => rule.headers)
.find((header) => header.key === 'Content-Security-Policy');
expect(csp).toBeDefined();
expect(csp!.value).toContain('https://analytics.schoolcompare.co.uk');
});
});
describe('admin surface', () => {
it('serves noindex on the admin panel and the CMS API', async () => {
// robots.txt disallows these too, but a Disallow only blocks crawling — a
// URL found from an external link can still be indexed without ever being
// fetched. This header is what actually keeps them out.
const headers = await headerRules();
for (const source of ['/admin/:path*', '/cms-api/:path*']) {
const rule = headers.find((entry) => entry.source === source);
expect(rule).toBeDefined();
expect(rule!.headers).toContainEqual({
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
});
}
});
});
@@ -1,4 +1,4 @@
import { generateMetadata as placeMeta } from '@/app/schools/[place]/page';
import { generateMetadata as placeMeta } from '@/app/(frontend)/schools/[place]/page';
jest.mock('@/lib/places', () => ({
...jest.requireActual('@/lib/places'),
@@ -0,0 +1,30 @@
/** @jest-environment node */
import { GET } from '@/app/(frontend)/release.json/route';
import { readFile } from 'node:fs/promises';
jest.mock('node:fs/promises', () => ({ readFile: jest.fn() }));
const realFetch = global.fetch;
const identity = { sha: 'a'.repeat(40), build_id: 'b'.repeat(32) };
beforeEach(() => {
jest.mocked(readFile).mockResolvedValue(JSON.stringify(identity));
global.fetch = jest.fn(async () => Response.json(identity));
});
afterEach(() => { global.fetch = realFetch; jest.resetAllMocks(); });
test('reports immutable file identity and backend identity without caching', async () => {
const response = await GET();
expect(response.status).toBe(200);
expect(response.headers.get('Cache-Control')).toBe('no-store');
expect(await response.json()).toEqual({ frontend: identity, backend: identity });
expect(fetch).toHaveBeenCalledWith(expect.stringMatching(/\/api\/release$/), expect.objectContaining({ cache: 'no-store', signal: expect.anything() }));
});
test('missing build metadata cannot pass the release gate', async () => {
jest.mocked(readFile).mockRejectedValueOnce(new Error('missing file'));
expect((await GET()).status).toBe(503);
});
test('backend failure cannot pass the release gate', async () => {
jest.mocked(fetch).mockResolvedValueOnce(new Response('', { status: 503 }));
expect((await GET()).status).toBe(503);
});
+20
View File
@@ -0,0 +1,20 @@
import robots from '@/app/robots';
describe('robots.txt', () => {
it('disallows the admin panel and the CMS API', () => {
const rules = robots().rules;
const rule = Array.isArray(rules) ? rules[0] : rules;
expect(rule.disallow).toEqual(
expect.arrayContaining(['/api/', '/_next/', '/admin/', '/cms-api/']),
);
});
});
describe('sitemap discovery', () => {
it('lists both the proxied school sitemap and the Next-owned content sitemap', () => {
expect(robots().sitemap).toEqual([
'https://www.schoolcompare.co.uk/sitemap.xml',
'https://www.schoolcompare.co.uk/content-sitemap.xml',
]);
});
});
@@ -0,0 +1,149 @@
import { render, screen } from '@testing-library/react';
import { DestinationsSection } from '@/components/school/DestinationsSection';
import type { DestinationPhase } from '@/lib/types';
import type { DestinationCategory, DestinationStatus } from '@/lib/destinations';
const cell = (
category: DestinationCategory,
pupils: number | null,
status: DestinationStatus = 'published',
) => ({
category, pupils,
percentage: pupils === null ? null : (pupils / 180) * 100,
status,
});
const ALL_PUBLISHED = [
cell('school_sixth_form', 75), cell('sixth_form_college', 21),
cell('further_education', 55), cell('other_education', 6),
cell('apprenticeship', 8), cell('employment', 6),
cell('not_sustained', 5), cell('not_captured', 4),
];
const fullPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: { all: { cohort: 180, categories: ALL_PUBLISHED } },
};
const suppressedPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: [
cell('school_sixth_form', 75), cell('sixth_form_college', null, 'suppressed'),
cell('further_education', 55), cell('other_education', 6),
cell('apprenticeship', 8), cell('employment', 6),
cell('not_sustained', 5), cell('not_captured', 4),
],
},
},
};
describe('DestinationsSection', () => {
it('dates its own cohort so it is not read as stale next to the GCSE section', () => {
render(<DestinationsSection destinations={fullPhase} />);
expect(screen.getByText(/2022\/23/)).toBeInTheDocument();
});
it('renders one bar segment per published category', () => {
const { container } = render(<DestinationsSection destinations={fullPhase} />);
expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(8);
});
it('renders NO bar at all when a category is withheld', () => {
const { container } = render(<DestinationsSection destinations={suppressedPhase} />);
// R1: a bar with a gap in it publishes the withheld figure by its width.
expect(container.querySelectorAll('[data-destination-segment]')).toHaveLength(0);
expect(screen.getAllByText(/withheld/i).length).toBeGreaterThan(0);
});
it('never states the remainder for a partially suppressed group', () => {
const { container } = render(<DestinationsSection destinations={suppressedPhase} />);
// 180 cohort - 159 published = 21, the withheld figure. It must appear nowhere.
expect(container.textContent).not.toMatch(/\b21\b/);
});
it('shows a card value for a group whose components are all published', () => {
render(<DestinationsSection destinations={fullPhase} />);
// academic route = 75 + 21 = 96 of 180 = 53%
expect(screen.getByText('53%')).toBeInTheDocument();
});
it('refuses a card value when one of its components is withheld', () => {
render(<DestinationsSection destinations={suppressedPhase} />);
// academic route needs sixth_form_college, which is suppressed.
expect(screen.getByText(/not published/i)).toBeInTheDocument();
expect(screen.queryByText('53%')).not.toBeInTheDocument();
});
it('never claims a pupil stayed at this school', () => {
const { container } = render(<DestinationsSection destinations={fullPhase} />);
// The published file reports destination TYPE, never destination institution.
expect(container.textContent).not.toMatch(/stayed on (here|at this school)/i);
});
it('renders nothing when no group carries categories', () => {
const empty: DestinationPhase = { cohort_year: '2022/23', groups: {} };
const { container } = render(<DestinationsSection destinations={empty} />);
expect(container.firstChild).toBeNull();
});
});
describe('the detail table keeps the three statuses apart', () => {
// 'suppressed' and 'not_applicable' are different claims, and the mart, the
// SQLAlchemy model and the serialiser all preserve the difference. The table
// used to key its Share column off `percentage === null`, which is true for
// both, so a category that simply does not apply was labelled "withheld" —
// while the Pupils column beside it rendered blank.
const mixedPhase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: [
cell('school_sixth_form', 75),
cell('sixth_form_college', null, 'suppressed'),
cell('further_education', null, 'suppressed'),
cell('apprenticeship', null, 'not_applicable'),
],
},
},
};
const rowFor = (container: HTMLElement, category: string) =>
Array.from(container.querySelectorAll('tbody tr'))
.find(tr => tr.textContent?.includes(category));
it('never labels a not-applicable category as withheld', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'Apprenticeship');
expect(row).toBeTruthy();
expect(row!.textContent).not.toMatch(/withheld/i);
});
it('labels a genuinely suppressed category as withheld in both columns', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'Sixth-form college');
expect(row).toBeTruthy();
expect(row!.querySelectorAll('td')).toHaveLength(2);
Array.from(row!.querySelectorAll('td')).forEach(td =>
expect(td.textContent).toMatch(/withheld/i));
});
it('the two columns of a row never disagree about what the row is', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
Array.from(container.querySelectorAll('tbody tr')).forEach(tr => {
const cells = Array.from(tr.querySelectorAll('td'))
.map(td => /withheld/i.test(td.textContent ?? ''));
expect(new Set(cells).size).toBe(1);
});
});
it('shows a published category its real figures', () => {
const { container } = render(<DestinationsSection destinations={mixedPhase} />);
const row = rowFor(container, 'State-funded school sixth form');
expect(row!.textContent).toMatch(/75/);
expect(row!.textContent).toMatch(/42%/);
});
});
@@ -3,10 +3,11 @@ import userEvent from '@testing-library/user-event';
import { FilterBar } from '@/components/FilterBar';
const push = jest.fn();
let searchParams = new URLSearchParams();
jest.mock('next/navigation', () => ({
useRouter: () => ({ push, replace: jest.fn(), prefetch: jest.fn() }),
usePathname: () => '/',
useSearchParams: () => new URLSearchParams(),
useSearchParams: () => searchParams,
}));
const FILTERS = {
@@ -24,6 +25,7 @@ beforeEach(() => {
phase: 'Primary', school_type: 'Community school' }] }),
})) as unknown as typeof fetch;
push.mockClear();
searchParams = new URLSearchParams();
});
afterEach(() => { global.fetch = realFetch; });
@@ -74,3 +76,35 @@ describe('FilterBar autosuggest', () => {
expect.stringContaining('search=brecknock')));
});
});
describe('FilterBar autosuggest does not reopen over results', () => {
it('stays shut when the input arrives pre-filled from the URL', async () => {
/*
* The results-page bar renders with the search term already in the input.
* Opening on that would drop the dropdown on top of the results the search
* just produced — which is exactly what happened: the first result became
* unclickable, because the list sat over it and swallowed the pointer.
*
* Suggestions answer typing, not the presence of a value.
*/
searchParams = new URLSearchParams('search=brecknock');
render(<FilterBar filters={FILTERS} autosuggest />);
expect(screen.getByRole('combobox')).toHaveValue('brecknock');
await new Promise((r) => setTimeout(r, 300)); // past the 200ms debounce
expect(global.fetch).not.toHaveBeenCalled();
expect(screen.queryByRole('listbox')).not.toBeInTheDocument();
});
it('closes the dropdown when the search is submitted', async () => {
render(<FilterBar filters={FILTERS} autosuggest />);
const input = screen.getByRole('combobox');
await userEvent.type(input, 'brecknock');
expect(await screen.findByRole('listbox')).toBeInTheDocument();
await userEvent.type(input, '{Enter}');
await waitFor(() =>
expect(screen.queryByRole('listbox')).not.toBeInTheDocument());
});
});
@@ -0,0 +1,47 @@
/**
* The footer is the only navigational route to /about and /blog, so it is
* where a dark flag would otherwise leave a link into a 404.
*
* Both props default to false. A caller that forgets to pass them hides the
* links, which is the direction that cannot break a page — the same reasoning
* as backend/flags.py's "every flag defaults to False".
*/
import { render, screen } from '@testing-library/react';
import { Footer } from '@/components/Footer';
describe('footer feature links', () => {
it('links to both when both flags are on', () => {
render(<Footer aboutEnabled blogEnabled />);
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toHaveAttribute('href', '/about');
expect(screen.getByRole('link', { name: /^blog$/i }))
.toHaveAttribute('href', '/blog');
});
it('omits the about link when that flag is dark', () => {
render(<Footer blogEnabled />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.getByRole('link', { name: /^blog$/i })).toBeInTheDocument();
});
it('omits the blog link when that flag is dark', () => {
render(<Footer aboutEnabled />);
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
expect(screen.getByRole('link', { name: /who's behind this/i }))
.toBeInTheDocument();
});
it('drops the whole section when both are dark, not an empty heading', () => {
// Shipping dark means the footer renders as it did before the feature
// existed, not as a section with its contents removed.
render(<Footer />);
expect(screen.queryByRole('heading', { name: /^about$/i })).toBeNull();
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
expect(screen.queryByRole('link', { name: /^blog$/i })).toBeNull();
});
it('defaults to dark when a caller passes nothing', () => {
render(<Footer />);
expect(screen.queryByRole('link', { name: /who's behind this/i })).toBeNull();
});
});
@@ -0,0 +1,81 @@
import { act, fireEvent, render, screen } from '@testing-library/react';
import { HomeView } from '@/components/HomeView';
import { fetchSchools } from '@/lib/api';
import { primaryFixture } from '../support/schoolFixtures';
import type { SchoolsResponse, School } from '@/lib/types';
let params = new URLSearchParams('postcode=SW1A+1AA');
jest.mock('next/navigation', () => ({
useSearchParams: () => params,
usePathname: () => '/',
useRouter: () => ({ push: jest.fn(), replace: jest.fn() }),
}));
jest.mock('@/context/ComparisonContext', () => ({
useComparisonContext: () => ({ addSchool: jest.fn(), removeSchool: jest.fn(), selectedSchools: [] }),
}));
jest.mock('@/lib/api', () => ({
fetchSchools: jest.fn(),
fetchNationalAverages: jest.fn(async () => ({})),
fetchLAaverages: jest.fn(async () => ({ secondary: { attainment_8_by_la: {} } })),
}));
jest.mock('@/components/FilterBar', () => ({ FilterBar: () => null }));
jest.mock('@/components/SchoolRow', () => ({ SchoolRow: ({ school }: {school: School}) => <div>{school.school_name}</div> }));
jest.mock('@/components/SchoolMap', () => ({ SchoolMap: ({ schools }: {schools: School[]}) => <div data-testid="map">{schools.map(s => s.school_name).join(',')}</div> }));
const filters = { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
function response(name: string): SchoolsResponse {
return { schools: [{ ...primaryFixture.schoolInfo, school_name: name }],
total: 2, page: 1, page_size: 1, total_pages: 2 };
}
function deferred() {
let resolve!: (value: SchoolsResponse) => void;
let reject!: (error: Error) => void;
const promise = new Promise<SchoolsResponse>((yes, no) => { resolve = yes; reject = no; });
return { promise, resolve, reject };
}
beforeEach(() => {
params = new URLSearchParams('postcode=SW1A+1AA');
jest.mocked(fetchSchools).mockReset();
});
test('load-more results from an old search are discarded, even after returning to it', async () => {
const pending = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(pending.promise);
const view = render(<HomeView initialSchools={response('Initial A')} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Load more schools' }));
const signal = jest.mocked(fetchSchools).mock.calls[0][1]?.signal;
params = new URLSearchParams('postcode=SW2+1AA');
view.rerender(<HomeView initialSchools={response('Initial B')} filters={filters} />);
expect(signal?.aborted).toBe(true);
params = new URLSearchParams('postcode=SW1A+1AA');
view.rerender(<HomeView initialSchools={response('Fresh A')} filters={filters} />);
await act(async () => pending.resolve(response('Stale append')));
expect(screen.queryByText('Stale append')).not.toBeInTheDocument();
expect(screen.getByRole('button', { name: 'Load more schools' })).toBeEnabled();
});
test('an older map response cannot overwrite the current search', async () => {
const first = deferred(), second = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(first.promise).mockReturnValueOnce(second.promise);
const initial = response('Initial A');
const view = render(<HomeView initialSchools={initial} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Map' }));
params = new URLSearchParams('postcode=SW2+1AA');
view.rerender(<HomeView initialSchools={response('Initial B')} filters={filters} />);
await act(async () => second.resolve(response('Current map')));
await act(async () => first.resolve(response('Stale map')));
expect(screen.getByTestId('map')).toHaveTextContent('Current map');
expect(screen.getByTestId('map')).not.toHaveTextContent('Stale map');
});
test('failed map requests can be retried by reopening the map', async () => {
const pending = deferred();
jest.mocked(fetchSchools).mockReturnValueOnce(pending.promise).mockResolvedValue(response('Retry result'));
render(<HomeView initialSchools={response('Initial')} filters={filters} />);
fireEvent.click(screen.getByRole('button', { name: 'Map' }));
await act(async () => pending.reject(new Error('offline')));
fireEvent.click(screen.getByRole('button', { name: 'List' }));
await act(async () => fireEvent.click(screen.getByRole('button', { name: 'Map' })));
expect(fetchSchools).toHaveBeenCalledTimes(2);
expect(screen.getByTestId('map')).toHaveTextContent('Retry result');
});
@@ -0,0 +1,92 @@
/**
* The module that ends the stranding: before it, a school page's only anchor
* pointed at the school's own website, so ~27k pages sent authority off-site
* and none of it reached the location layer.
*/
import { render, screen } from '@testing-library/react';
import { NearbyPlaces } from '@/components/school/NearbyPlaces';
const essex = { kind: 'authority', slug: 'essex', name: 'Essex', count: 480, url: '/schools/authority/essex', phases: [] };
const brentwood = { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 37, url: '/schools/brentwood', phases: [] };
const cm15 = { kind: 'outcode', slug: 'cm15', name: 'CM15', count: 12, url: '/schools/near/cm15', phases: [] };
describe('NearbyPlaces', () => {
it('links to every place the school belongs to', () => {
render(<NearbyPlaces places={[essex, brentwood, cm15]} />);
expect(screen.getByRole('link', { name: /Brentwood/ }))
.toHaveAttribute('href', '/schools/brentwood');
expect(screen.getByRole('link', { name: /Essex/ }))
.toHaveAttribute('href', '/schools/authority/essex');
expect(screen.getByRole('link', { name: /CM15/ }))
.toHaveAttribute('href', '/schools/near/cm15');
});
it('says how many schools each link leads to', () => {
// An anchor that states its destination's size is worth more to a reader
// and to a crawler than "see more".
render(<NearbyPlaces places={[brentwood]} />);
expect(screen.getByRole('link', { name: /37 schools in Brentwood/ }))
.toBeInTheDocument();
});
it('renders nothing at all when the school has no published places', () => {
// Not an empty heading. A school whose town and authority both fall below
// the threshold has nowhere to point, and the page should look as it did
// before the module existed.
const { container } = render(<NearbyPlaces places={[]} />);
expect(container).toBeEmptyDOMElement();
});
it('puts the narrowest place first, which is the most useful link', () => {
// The API orders widest-first for the breadcrumb; a reader on a school
// page wants its town before its county.
render(<NearbyPlaces places={[essex, brentwood, cm15]} />);
const hrefs = screen.getAllByRole('link').map((a) => a.getAttribute('href'));
expect(hrefs.indexOf('/schools/brentwood'))
.toBeLessThan(hrefs.indexOf('/schools/authority/essex'));
});
it('handles a singular count without saying "1 schools"', () => {
render(<NearbyPlaces places={[{ ...brentwood, count: 1 }]} />);
expect(screen.getByRole('link', { name: /1 school in Brentwood/ }))
.toBeInTheDocument();
});
it('links the phase page the school appears on', () => {
// "primary schools in brentwood" is the query these pages exist for.
render(<NearbyPlaces places={[{
...brentwood,
phases: [{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' }],
}]} />);
expect(screen.getByRole('link', { name: /22 primary schools in Brentwood/ }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('links both phase pages for an all-through school', () => {
render(<NearbyPlaces places={[{
...brentwood,
phases: [
{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' },
{ phase: 'secondary', count: 9, url: '/schools/brentwood/secondary' },
],
}]} />);
expect(screen.getByRole('link', { name: /22 primary schools/ })).toBeInTheDocument();
expect(screen.getByRole('link', { name: /9 secondary schools/ })).toBeInTheDocument();
});
it('keeps a phase link next to the place it belongs to', () => {
// Grouping matters: "22 primary schools in Brentwood" directly after
// "37 schools in Brentwood" reads as one place, not two unrelated links.
render(<NearbyPlaces places={[essex, {
...brentwood,
phases: [{ phase: 'primary', count: 22, url: '/schools/brentwood/primary' }],
}]} />);
const hrefs = screen.getAllByRole('link').map((a) => a.getAttribute('href'));
expect(hrefs.indexOf('/schools/brentwood/primary'))
.toBe(hrefs.indexOf('/schools/brentwood') + 1);
});
});
@@ -0,0 +1,192 @@
/**
* The section's job is to be honest about what it is showing. These tests pin
* the ways it could mislead: rendering below the minimum, claiming a likeness
* it does not rank on, showing a missing figure as a number, or hiding a card
* behind an arrow where a crawler cannot reach it.
*/
import { render, screen } from '@testing-library/react';
import {
nearbyNoun,
NearbySchoolsSection,
shouldRenderNearby,
} from '@/components/school/NearbySchoolsSection';
import type { NearbySchool } from '@/lib/types';
jest.mock('@/components/school/AddToCompareButton', () => ({
AddToCompareButton: ({ school }: { school: NearbySchool }) => (
<button type="button">Add {school.school_name} to compare</button>
),
}));
jest.mock('@/components/school/NearbySchoolsCompareBar', () => ({
NearbySchoolsCompareBar: ({ thisUrn }: { thisUrn: number }) => (
<div data-testid="compare-bar">bar for {thisUrn}</div>
),
}));
function school(overrides: Partial<NearbySchool> = {}): NearbySchool {
return {
urn: 100002,
school_name: 'Willow Lane Primary School',
distance_miles: 0.6,
school_type: 'Community school',
age_range: '4-11',
shared: ['Mixed', 'No religious character'],
metric_value: 74,
metric_key: 'rwm_expected_pct',
metric_year: 202425,
...overrides,
};
}
function renderSection(nearby: NearbySchool[]) {
return render(
<NearbySchoolsSection
urn={100001}
schoolName="Meadowbrook Primary School"
phase="Primary"
thisMetricValue={72}
nearby={nearby}
/>,
);
}
describe('render gates', () => {
it.each([
['undefined', undefined],
['null', null],
['empty', []],
['a single school', [school()]],
])('renders nothing for %s', (_label, value) => {
expect(shouldRenderNearby(value as NearbySchool[] | null | undefined)).toBe(false);
});
it('renders for two or more schools', () => {
expect(shouldRenderNearby([school(), school({ urn: 100003 })])).toBe(true);
});
it('returns null rather than an empty shell below the minimum', () => {
const { container } = renderSection([school()]);
expect(container).toBeEmptyDOMElement();
});
});
describe('what the section claims', () => {
it('never claims a similar intake, because it does not rank on one', () => {
renderSection([school(), school({ urn: 100003, shared: [] })]);
expect(screen.queryByText(/similar intake/i)).not.toBeInTheDocument();
});
it('is headed "Other schools nearby", not "similar"', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByRole('heading', { name: 'Other schools nearby' })).toBeInTheDocument();
});
it('shows chips for what is shared', () => {
renderSection([school({ shared: ['Mixed', 'Roman Catholic'] }), school({ urn: 100003 })]);
expect(screen.getAllByText('Roman Catholic').length).toBe(1);
});
it('shows no chips at all when nothing is shared, rather than inventing one', () => {
const { container } = render(
<NearbySchoolsSection
urn={100001}
schoolName="Meadowbrook Primary School"
phase="Primary"
thisMetricValue={72}
nearby={[school({ shared: [] }), school({ urn: 100003, shared: [] })]}
/>,
);
// The card still carries its distance, name, type and figure — just no
// claim of likeness.
expect(container.querySelectorAll('li ul').length).toBe(0);
expect(screen.getAllByText(/miles away/).length).toBe(2);
});
});
describe('what the lede calls the set', () => {
it.each([
['Primary', 'primary schools'],
['Middle deemed primary', 'primary schools'],
['Secondary', 'secondary schools'],
['Middle deemed secondary', 'secondary schools'],
['All-through', 'all-through schools'],
// GIAS phase 6. Its candidates span the whole secondary group, so no
// single noun fits and it takes the honest general one.
['16 plus', 'schools and colleges'],
['', 'schools'],
[null, 'schools'],
])('calls a %s school\'s neighbours "%s"', (phase, expected) => {
expect(nearbyNoun(phase)).toBe(expected);
});
it('never calls a sixth form college\'s neighbours primary schools', () => {
render(
<NearbySchoolsSection
urn={100001}
schoolName="Barnet Sixth Form College"
phase="16 plus"
thisMetricValue={null}
nearby={[school(), school({ urn: 100003 })]}
/>,
);
expect(screen.getByText(/Other schools and colleges near Barnet Sixth Form College/)).toBeInTheDocument();
expect(screen.queryByText(/primary schools/)).not.toBeInTheDocument();
});
});
describe('cards', () => {
it('links each school to its canonical slug', () => {
renderSection([school(), school({ urn: 100003, school_name: 'Oakfield Primary School' })]);
const link = screen.getByRole('link', { name: /Willow Lane Primary School/ });
expect(link).toHaveAttribute('href', '/school/100002-willow-lane-primary-school');
});
it('shows the distance and the shared characteristics', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getAllByText('0.6 miles away').length).toBeGreaterThan(0);
expect(screen.getAllByText('Mixed').length).toBeGreaterThan(0);
});
it('renders a missing figure as "Not published", never as a number', () => {
renderSection([school({ metric_value: null }), school({ urn: 100003 })]);
expect(screen.getByText('Not published')).toBeInTheDocument();
});
it('anchors each figure against this school', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getAllByText('72% at this school').length).toBe(2);
});
it('offers the compare bar once, for this school', () => {
renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByTestId('compare-bar')).toHaveTextContent('bar for 100001');
});
it('keeps every card in the DOM, including the ones scrolled out of view', () => {
const six = Array.from({ length: 6 }, (_, n) =>
school({ urn: 100002 + n, school_name: `Peer ${n} School` }),
);
renderSection(six);
expect(screen.getAllByRole('link', { name: /Peer \d School/ })).toHaveLength(6);
});
it('offers no arrows when three cards fit the row', () => {
renderSection([school(), school({ urn: 100003 }), school({ urn: 100004 })]);
expect(screen.queryByRole('button', { name: /More schools/ })).not.toBeInTheDocument();
});
it('offers arrows once there is a fourth school', () => {
renderSection(Array.from({ length: 4 }, (_, n) => school({ urn: 100002 + n })));
expect(screen.getByRole('button', { name: /More schools/ })).toBeInTheDocument();
expect(screen.getByRole('button', { name: /Previous schools/ })).toBeInTheDocument();
});
it('says distances are straight-line, and offers no method panel', () => {
const { container } = renderSection([school(), school({ urn: 100003 })]);
expect(screen.getByText(/straight-line from this school/i)).toBeInTheDocument();
expect(container.querySelector('details')).toBeNull();
});
});
@@ -346,3 +346,134 @@ describe('PlaceView unlinkable authorities', () => {
.toContain('Isles Of Scilly');
});
});
describe('PlaceView school attributes', () => {
/*
* The table shipped with one column of scores, which answers "how did they
* do" and nothing about whether the school is one a family could use. Age
* range, faith, nursery and constituency are the four facts a parent
* filters on before they look at a number at all.
*/
const withAttributes: PlaceDetail = {
place: { kind: 'town', slug: 'chelmsford', name: 'Chelmsford', count: 3,
parent_authority: 'Essex', phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null,
age_range: '4-11', religious_denomination: 'Church of England',
nursery_provision: true,
parliamentary_constituency: 'Chelmsford' } as never,
{ urn: 2, school_name: 'Beta High', phase: 'Secondary',
rwm_expected_pct: null, attainment_8_score: 47,
age_range: '11-16', religious_denomination: 'Does not apply',
nursery_provision: false,
parliamentary_constituency: 'Witham' } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: 45 },
};
function headings(container: HTMLElement, table = 0): string[] {
return Array.from(container.querySelectorAll('table')[table]
.querySelectorAll('thead th')).map((th) => th.textContent ?? '');
}
it('heads a primary table with all four attributes', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
expect(headings(container)).toEqual([
'School', 'Reading, writing & maths',
'Ages', 'Religious character', 'Nursery', 'Constituency',
]);
});
it('omits nursery from a secondary table, where it does not apply', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
expect(headings(container, 1)).toEqual([
'School', 'Attainment 8', 'Ages', 'Religious character', 'Constituency',
]);
});
it('keeps the measure beside the school name, where a phone can see it', () => {
// Six columns overflow a phone and .tableWrap turns that into a swipe.
// With the measure last, the one number the page exists for is the one
// scrolled off the screen.
const { container } = render(<PlaceView detail={withAttributes}
phase="primary" englandAverage={61} neighbours={[]} />);
expect(headings(container)[1]).toBe('Reading, writing & maths');
});
it('shows the age range without repeating the column heading', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('4–11')).toBeInTheDocument();
expect(screen.queryByText('Ages 4–11')).not.toBeInTheDocument();
});
it('names the faith of a faith school', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('Church of England')).toBeInTheDocument();
});
it('reads "Does not apply" as no religious character, not as a value', () => {
// GIAS spells the absence of a faith as "Does not apply", which is a
// database answer rather than an English one. The school page already
// suppresses it; the two must not disagree about the same school.
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
const secondary = container.querySelectorAll('table')[1]
.querySelectorAll('tbody td');
expect(secondary[3].textContent).toBe('—');
expect(screen.queryByText(/Does not apply/)).not.toBeInTheDocument();
});
it('marks a nursery as such and a school without one as not', () => {
const { container } = render(<PlaceView detail={withAttributes}
englandAverage={61} neighbours={[]} />);
const cells = container.querySelectorAll('table')[0]
.querySelectorAll('tbody td');
expect(cells[4].textContent).toBe('Yes');
});
it('names the constituency of each school', () => {
render(<PlaceView detail={withAttributes} englandAverage={61}
neighbours={[]} />);
expect(screen.getByText('Chelmsford', { selector: 'td' })).toBeInTheDocument();
expect(screen.getByText('Witham', { selector: 'td' })).toBeInTheDocument();
});
it('dashes an attribute the data does not carry', () => {
// nursery_provision and parliamentary_constituency are absent from marts
// the pipeline has not rebuilt, and the API degrades them to null rather
// than failing. A row must survive that.
const bare: PlaceDetail = {
...withAttributes,
schools: [{ urn: 3, school_name: 'Gamma Primary', phase: 'Primary',
rwm_expected_pct: 70 } as never],
};
const { container } = render(<PlaceView detail={bare} phase="primary"
englandAverage={61} neighbours={[]} />);
const cells = Array.from(container.querySelectorAll('tbody td'))
.map((td) => td.textContent);
expect(cells.slice(2)).toEqual(['—', '—', '—', '—']);
});
it('gives an all-through school its nursery under primary only', () => {
// All-through schools render in both groups. Nursery belongs to the
// primary reading of the same school, not the secondary one.
const allThrough: PlaceDetail = {
...withAttributes,
schools: [{ urn: 4, school_name: 'Delta Academy', phase: 'All-through',
rwm_expected_pct: 66, attainment_8_score: 51,
age_range: '4-18', religious_denomination: 'None',
nursery_provision: true,
parliamentary_constituency: 'Chelmsford' } as never],
};
const { container } = render(<PlaceView detail={allThrough}
englandAverage={61} neighbours={[]} />);
const tables = container.querySelectorAll('table');
expect(tables[0].textContent).toContain('Yes');
expect(tables[1].textContent).not.toContain('Yes');
});
});
@@ -0,0 +1,44 @@
import { render, screen } from '@testing-library/react';
import { Post16DestinationsSection } from '@/components/school/Post16DestinationsSection';
import type { DestinationPhase } from '@/lib/types';
const phase: DestinationPhase = {
cohort_year: '2022/23',
groups: {
all: {
cohort: 96,
categories: [
{ category: 'higher_education', pupils: 56, percentage: 58.3, status: 'published' },
{ category: 'further_education', pupils: 12, percentage: 12.5, status: 'published' },
{ category: 'apprenticeship', pupils: 9, percentage: 9.4, status: 'published' },
{ category: 'employment', pupils: 13, percentage: 13.5, status: 'published' },
{ category: 'not_sustained', pupils: 6, percentage: 6.3, status: 'published' },
],
},
},
};
describe('Post16DestinationsSection', () => {
it('names the Year 13 cohort, not Year 11', () => {
const { container } = render(<Post16DestinationsSection destinations={phase} />);
expect(container.textContent).toMatch(/Year 13/);
expect(container.textContent).not.toMatch(/Year 11/);
});
it('reports higher education destinations', () => {
render(<Post16DestinationsSection destinations={phase} />);
expect(screen.getByText(/UK higher education/i)).toBeInTheDocument();
});
it('uses its own anchor so the nav does not collide with After Year 11', () => {
const { container } = render(<Post16DestinationsSection destinations={phase} />);
expect(container.querySelector('#post16-destinations')).toBeTruthy();
expect(container.querySelector('#destinations')).toBeNull();
});
it('renders nothing when no group carries categories', () => {
const empty: DestinationPhase = { cohort_year: '2022/23', groups: {} };
const { container } = render(<Post16DestinationsSection destinations={empty} />);
expect(container.firstChild).toBeNull();
});
});
@@ -0,0 +1,38 @@
/**
* The trail has to be written by something, and it has to be written on every
* route — not only the ones that happen to track an event.
*/
import { render } from '@testing-library/react';
const recordVisitedPath = jest.fn();
let pathname = '/schools/brentwood';
jest.mock('next/navigation', () => ({ usePathname: () => pathname }));
jest.mock('@/lib/analytics', () => ({
recordVisitedPath: (p: string) => recordVisitedPath(p),
}));
// eslint-disable-next-line @typescript-eslint/no-var-requires
const { RouteTrail } = require('@/components/RouteTrail');
describe('RouteTrail', () => {
beforeEach(() => recordVisitedPath.mockClear());
it('records the page it is mounted on', () => {
render(<RouteTrail />);
expect(recordVisitedPath).toHaveBeenCalledWith('/schools/brentwood');
});
it('records each new route as the user moves through the app', () => {
const { rerender } = render(<RouteTrail />);
pathname = '/school/115429-brentwood-school';
rerender(<RouteTrail />);
expect(recordVisitedPath).toHaveBeenLastCalledWith(
'/school/115429-brentwood-school');
});
it('renders nothing, so it can sit anywhere in the layout', () => {
const { container } = render(<RouteTrail />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -0,0 +1,45 @@
import { render } from '@testing-library/react';
import { TrackPlaceView } from '@/components/places/TrackPlaceView';
const trackMock = jest.fn();
jest.mock('@/lib/analytics', () => ({
track: (...args: unknown[]) => trackMock(...args),
getNavigationSource: () => 'search',
}));
describe('TrackPlaceView', () => {
beforeEach(() => trackMock.mockClear());
it('reports which kind of location page was viewed', () => {
/*
* `kind` is the reason this event exists. Whether to keep investing in the
* location layer turns on which *sort* of page earns engagement — towns,
* authorities or postcode districts — and a bare pageview cannot say,
* because all four families share the /schools/ prefix.
*/
render(<TrackPlaceView kind="authority" slug="kent" count={412} />);
expect(trackMock).toHaveBeenCalledWith('place_viewed', {
kind: 'authority', slug: 'kent', phase: 'all',
school_count: 412, from: 'search',
});
});
it('names the phase when the page is a phase variant', () => {
render(<TrackPlaceView kind="town" slug="brentwood" count={29} phase="primary" />);
expect(trackMock).toHaveBeenCalledWith('place_viewed',
expect.objectContaining({ phase: 'primary' }));
});
it('fires once, not once per render', () => {
const { rerender } = render(
<TrackPlaceView kind="town" slug="brentwood" count={29} />);
rerender(<TrackPlaceView kind="town" slug="brentwood" count={29} />);
expect(trackMock).toHaveBeenCalledTimes(1);
});
it('renders nothing', () => {
const { container } = render(
<TrackPlaceView kind="town" slug="brentwood" count={29} />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -32,10 +32,17 @@ function stylesheets(dir: string): string[] {
}
/** Innermost `selector { body }` pairs. Nested at-rules never match as rules,
* because their body contains braces. */
* because their body contains braces.
*
* Comments are stripped before matching rather than after, so that the whole
* selector survives. Taking only its last line — which is what stripping a
* leading comment used to require — silently discarded every selector in a
* grouped rule but the final one, and a safety guard that cannot see half its
* input fails open. */
function rules(css: string): Array<{ selector: string; body: string }> {
return Array.from(css.matchAll(/([^{}]+)\{([^{}]*)\}/g), (m) => ({
selector: m[1].trim().split('\n').pop()!.trim(),
const bare = css.replace(/\/\*[\s\S]*?\*\//g, '');
return Array.from(bare.matchAll(/([^{}]+)\{([^{}]*)\}/g), (m) => ({
selector: m[1].trim().replace(/\s*\n\s*/g, ' '),
body: m[2],
}));
}
@@ -45,6 +52,15 @@ const THEMED_COLOR = /(?:^|[^-])color:\s*var\(--/;
const files = stylesheets(COMPONENTS);
/** Component sources, for the third-party-surface rule below. */
function sources(dir: string): string[] {
return fs.readdirSync(dir, { withFileTypes: true }).flatMap((entry) => {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) return sources(full);
return entry.name.endsWith('.tsx') ? [full] : [];
});
}
describe('dark-theme safety', () => {
it('finds stylesheets to check', () => {
expect(files.length).toBeGreaterThan(0);
@@ -76,3 +92,101 @@ describe('dark-theme safety', () => {
expect(offenders).toEqual([]);
});
});
/**
* The same defect one stylesheet further out.
*
* The rules above scan our own CSS modules. They cannot see a surface painted
* by a third-party sheet: leaflet.css hardcodes `background: white` on
* `.leaflet-popup-content-wrapper` and `.leaflet-popup-tip`, and
* LeafletMapInner builds its popup as an HTML string with inline
* `color: var(--text-primary)`. Neither half lives in a .module.css, so the
* module scan passed while dark mode rendered #E9EEF0 on #FFFFFF — 1.17:1,
* with the school name and the headline figure effectively invisible.
*
* globals.css already pulls the rest of Leaflet's chrome onto the tokens (the
* attribution bar, the zoom controls) for exactly this reason. The popup was
* simply missed.
*/
describe('third-party surfaces under themed text', () => {
const GLOBALS = path.join(__dirname, '..', '..', 'app', '(frontend)', 'globals.css');
/** Leaflet surfaces our own code writes token-coloured text onto. */
const LEAFLET_POPUP_SURFACES = [
'.leaflet-popup-content-wrapper',
'.leaflet-popup-tip',
];
it('still finds a component painting themed text into a Leaflet popup', () => {
// Guards the rule below against passing vacuously if the popups are ever
// rewritten as React components rather than HTML strings.
const themed = sources(COMPONENTS).filter((file) => {
const src = fs.readFileSync(file, 'utf8');
return /bindPopup\(/.test(src) && /color:var\(--|color: var\(--/.test(src);
});
expect(themed.length).toBeGreaterThan(0);
});
it('themes the Leaflet popup surface, because the text on it is themed', () => {
const globals = rules(fs.readFileSync(GLOBALS, 'utf8'));
const unthemed = LEAFLET_POPUP_SURFACES.filter((surface) => {
const rule = globals.find((r) => r.selector.includes(surface));
return !rule || !/background[^;]*var\(--/.test(rule.body);
});
// Leaflet's white is not a colour this site owns. Either the surface
// follows the theme or the text on it must be literal — and the text is
// already themed.
expect(unthemed).toEqual([]);
});
it('never puts a literal white label on a themed fill', () => {
/*
* The mirror image of the module-CSS rule above, and the half of the popup
* that theming the card does not reach. "View Details" is
* `background:var(--status-above);color:white`; --status-above is #36743F
* in light but #7FCB8A in dark, so the label went from 5.63:1 to 1.94:1.
*
* --text-inverse is the token for ink on a saturated fill — #FFFFFF in
* light, #111A20 in dark — and the popup's Ofsted badge already uses it.
*/
const offenders = sources(COMPONENTS).flatMap((file) => {
const src = fs.readFileSync(file, 'utf8');
return Array.from(
src.matchAll(/background:\s*var\(--[^;"']*;[^"']*?color:\s*(white|#fff\b|#ffffff\b)/gi),
() => path.relative(COMPONENTS, file));
});
expect(offenders).toEqual([]);
});
});
/**
* Destination measures add the first new colour family since the palette was
* set. The tokens have to exist in both blocks or the section renders one
* theme's fills on the other theme's ground — the exact failure the suite
* above exists to catch, but for tokens rather than literals.
*/
describe('destination tokens', () => {
const css = fs.readFileSync(
path.join(__dirname, '..', '..', 'app', '(frontend)', 'globals.css'), 'utf8');
const TOKENS = [
'--dest-sixthform', '--dest-sfcollege', '--dest-fecollege',
'--dest-apprentice', '--dest-employment', '--dest-none', '--dest-none-hatch',
];
const DARK_AT = css.indexOf('@media (prefers-color-scheme: dark)');
it('defines every destination token in the light palette', () => {
const light = css.slice(0, DARK_AT);
expect(TOKENS.filter((t) => !light.includes(`${t}:`))).toEqual([]);
});
it('redefines every destination token for dark', () => {
const dark = css.slice(DARK_AT);
expect(TOKENS.filter((t) => !dark.includes(`${t}:`))).toEqual([]);
});
});
@@ -0,0 +1,108 @@
import fs from 'fs';
import path from 'path';
/**
* The hero search and the results filter bar are the same component in two
* costumes. `.filterBar` is the card — background, border, shadow, padding —
* and `.heroMode` strips all of it so the search sits directly on the hero
* panel.
*
* Both selectors have specificity (0,1,0), so **source order decides**, and
* `.heroMode` only wins because it is declared immediately after. Any later
* bare `.filterBar` rule — which in practice means one inside a media query —
* silently wins instead, and the hero grows a card's padding back.
*
* That is exactly what happened: `@media (max-width: 768px) { .filterBar {
* padding: 0.875rem } }` re-added 14px in hero mode, indenting the search box,
* the hint and the location link 14px past the headline above them and costing
* the search field 28px of width on a 390px screen. The two rules directly
* below it in the same block were correctly written as
* `.filterBar:not(.heroMode)`; this one was missed, and nothing caught it
* because the result is a plausible-looking layout rather than a broken one.
*/
const CSS = path.join(__dirname, '..', '..', 'components', 'FilterBar.module.css');
/** Properties `.heroMode` resets. A later bare `.filterBar` rule setting any
* of these puts the card back on the hero. */
const RESET_BY_HERO_MODE = [
'background', 'border', 'border-radius', 'box-shadow', 'padding',
];
/**
* Comments are stripped before anything is parsed.
*
* A `{` or `}` inside a comment would otherwise desynchronise the brace walk
* below and the rule regex alike, and the selector text captured for each rule
* would carry the preceding comment along with it.
*/
function withoutComments(css: string): string {
return css.replace(/\/\*[\s\S]*?\*\//g, '');
}
/**
* The individual selectors in a rule's prelude.
*
* Split on commas, because a selector list is a list: `.filterBar, .other { }`
* applies to `.filterBar` just as surely as `.filterBar { }` does, and an
* earlier version of this guard compared the whole prelude against the literal
* string '.filterBar' — so writing the regression as a comma list, or across
* two lines, would have walked straight past it.
*/
function selectorsOf(prelude: string): string[] {
return prelude.split(',').map((sel) => sel.trim().replace(/\s+/g, ' '))
.filter(Boolean);
}
function mediaQueryBodies(css: string): string[] {
const bodies: string[] = [];
const re = /@media[^{]*\{/g;
let m: RegExpExecArray | null;
while ((m = re.exec(css)) !== null) {
// Walk braces from the opening one to find this at-rule's whole body.
let depth = 1;
let i = m.index + m[0].length;
const start = i;
while (i < css.length && depth > 0) {
if (css[i] === '{') depth++;
else if (css[i] === '}') depth--;
i++;
}
bodies.push(css.slice(start, i - 1));
}
return bodies;
}
describe('FilterBar hero-mode scoping', () => {
const css = withoutComments(fs.readFileSync(CSS, 'utf8'));
it('confirms heroMode still resets the card, which is what makes this matter', () => {
const hero = css.match(/\.heroMode\s*\{([^}]*)\}/);
expect(hero).not.toBeNull();
expect(hero![1]).toMatch(/padding:\s*0/);
});
it('never re-applies card styling to the hero from inside a media query', () => {
const offenders: string[] = [];
for (const body of mediaQueryBodies(css)) {
for (const rule of body.matchAll(/([^{}]+)\{([^{}]*)\}/g)) {
// Only a *bare* .filterBar is dangerous, and it is dangerous wherever
// it appears in a selector list. Scoped variants
// (`.filterBar:not(.heroMode)`) and descendants are fine.
const selectors = selectorsOf(rule[1]);
if (!selectors.includes('.filterBar')) continue;
for (const prop of RESET_BY_HERO_MODE) {
if (new RegExp(`(^|[;\\s])${prop}\\s*:`).test(rule[2])) {
offenders.push(`${rule[1].trim()} sets ${prop}`);
}
}
}
}
// Fix by scoping the rule as `.filterBar:not(.heroMode)`, the way the
// neighbouring rules in the same block already are.
expect(offenders).toEqual([]);
});
});
@@ -98,6 +98,17 @@ describe('secondary detail page', () => {
expect(screen.getByText(/has not published a cut-off distance/)).toBeInTheDocument();
});
it('makes no claim about publication when the feature is switched off', () => {
// Absent, not null. The API omits the key entirely while the
// admission_distance flag is off, and "Islington has not published a
// cut-off distance" is then a statement about us, not about Islington —
// false wherever the authority does publish one.
renderSecondarySchoolDetail({ ...secondaryFixture, admissionDistance: undefined });
expect(screen.queryByText(/has not published a cut-off distance/)).not.toBeInTheDocument();
expect(screen.queryByText(/Contact the admissions authority/)).not.toBeInTheDocument();
});
});
// ── The Distance section ───────────────────────────────────────────────
+158
View File
@@ -0,0 +1,158 @@
import { getNavigationSource } from '@/lib/analytics';
/** jsdom's document.referrer is read-only; redefining it is the way in. */
function referrer(url: string) {
Object.defineProperty(document, 'referrer', { value: url, configurable: true });
}
const ORIGIN = 'http://localhost';
describe('getNavigationSource', () => {
afterEach(() => referrer(''));
it('attributes a visit from a location page to the place layer', () => {
/*
* The one this was added for.
*
* W2 published ~3,900 location pages whose entire purpose is to funnel
* search traffic onto school pages. Before this case existed they fell
* through to 'direct' — so the location layer's contribution was not
* merely missing from the funnel, it was being counted in the bucket you
* read as "typed the URL". The measurement that decides whether W2 worked
* was confidently reporting the wrong answer.
*/
referrer(`${ORIGIN}/schools/barnet`);
expect(getNavigationSource()).toBe('place');
});
it.each([
['/schools/authority/kent', 'authority'],
['/schools/near/sw11', 'outcode'],
['/schools/brentwood/primary', 'phase variant'],
])('covers %s (%s)', (path) => {
referrer(`${ORIGIN}${path}`);
expect(getNavigationSource()).toBe('place');
});
it('still calls a school page "detail", one character away', () => {
// /school/ and /schools/ differ by one letter and mean different things.
// A prefix test written in the wrong order silently merges them.
referrer(`${ORIGIN}/school/100010-brecknock-primary-school`);
expect(getNavigationSource()).toBe('detail');
});
it.each([
['/', 'search'],
['/rankings', 'rankings'],
['/compare?urns=1,2', 'compare'],
])('leaves %s attributed as %s', (path, expected) => {
referrer(`${ORIGIN}${path}`);
expect(getNavigationSource()).toBe(expected);
});
it('treats an external referrer as direct', () => {
// Umami records the real referrer on the pageview; this field is only
// about internal navigation.
referrer('https://www.google.com/search?q=schools+in+barnet');
expect(getNavigationSource()).toBe('direct');
});
it('treats no referrer as direct', () => {
referrer('');
expect(getNavigationSource()).toBe('direct');
});
});
/*
* The defect the existing suite could not see.
*
* Every test above sets document.referrer, which the browser writes only when
* a *document* loads. Every internal navigation in this app is an App Router
* soft navigation — history.pushState, no new document — so document.referrer
* keeps naming whatever opened the tab for the whole session. Verified on
* staging: /schools/brentwood → click a school → URL changes to /school/…
* and document.referrer is still "".
*
* So `from` reported 'direct' for essentially every in-app journey, and the
* suite passed because it only ever exercised the full-page-load path.
*/
function freshAnalytics() {
let mod!: typeof import('@/lib/analytics');
jest.isolateModules(() => {
mod = require('@/lib/analytics');
});
return mod;
}
function at(path: string) {
window.history.pushState({}, '', path);
}
describe('getNavigationSource across a soft navigation', () => {
afterEach(() => {
referrer('');
at('/');
});
it('attributes a school view to the place page the user actually came from', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/schools/brentwood');
recordVisitedPath('/schools/brentwood');
at('/school/115429-brentwood-school');
recordVisitedPath('/school/115429-brentwood-school');
expect(source()).toBe('place');
});
it('does not depend on whether the new path was recorded first', () => {
// The trail is written by a layout-level effect and read by a page-level
// one. React orders those by tree position, which is not a contract worth
// resting a measurement on, so the answer must be the same either way.
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
recordVisitedPath('/rankings');
at('/school/115429-brentwood-school');
expect(source()).toBe('rankings');
});
it('names the previous page, not the current one, when both are schools', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/school/100010-brecknock-primary-school');
recordVisitedPath('/school/100010-brecknock-primary-school');
at('/school/115429-brentwood-school');
recordVisitedPath('/school/115429-brentwood-school');
expect(source()).toBe('detail');
});
it('looks past a return visit to the page the user came back from', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
for (const p of ['/schools/brentwood', '/school/115429-brentwood-school',
'/schools/brentwood']) {
at(p);
recordVisitedPath(p);
}
expect(source()).toBe('detail');
});
it('falls back to the referrer on a real document load, where it is true', () => {
// A fresh module is a fresh document: nothing has been recorded, and
// document.referrer is meaningful again.
const { getNavigationSource: source } = freshAnalytics();
at('/school/115429-brentwood-school');
referrer(`${ORIGIN}/schools/barnet`);
expect(source()).toBe('place');
});
it('still reads an arrival from outside as direct', () => {
const { recordVisitedPath, getNavigationSource: source } = freshAnalytics();
at('/schools/brentwood');
recordVisitedPath('/schools/brentwood');
referrer('https://www.google.com/search?q=schools+in+brentwood');
expect(source()).toBe('direct');
});
});
@@ -0,0 +1,87 @@
import {
canAggregate, aggregateCells,
canRenderBar, toBarSegments, CARD_GROUPS,
type DestinationCell, type DestinationGroup, type DestinationCategory,
} from '@/lib/destinations';
const pub = (category: DestinationCategory, pupils: number, cohort: number): DestinationCell => ({
category, pupils, percentage: (pupils / cohort) * 100, status: 'published',
});
const sup = (category: DestinationCategory): DestinationCell => ({
category, pupils: null, percentage: null, status: 'suppressed',
});
const fullGroup = (): DestinationGroup => ({
cohort: 180,
cells: [
pub('school_sixth_form', 75, 180), pub('sixth_form_college', 21, 180),
pub('further_education', 55, 180), pub('other_education', 6, 180),
pub('apprenticeship', 8, 180), pub('employment', 6, 180),
pub('not_sustained', 5, 180), pub('not_captured', 4, 180),
],
});
describe('canAggregate — R2, computing from components', () => {
it('allows a sum when every component is published', () => {
expect(canAggregate([pub('apprenticeship', 8, 180), pub('employment', 6, 180)])).toBe(true);
});
it('refuses a sum when any component is suppressed', () => {
expect(canAggregate([pub('apprenticeship', 8, 180), sup('employment')])).toBe(false);
});
it('refuses a sum when every component is suppressed', () => {
expect(canAggregate([sup('apprenticeship'), sup('employment')])).toBe(false);
});
});
describe('aggregateCells', () => {
it('sums published cells and derives a percentage from the cohort', () => {
expect(aggregateCells([pub('apprenticeship', 8, 180), pub('employment', 6, 180)], 180))
.toEqual({ pupils: 14, percentage: (14 / 180) * 100 });
});
it('returns null rather than a partial sum when a component is suppressed', () => {
expect(aggregateCells([pub('apprenticeship', 8, 180), sup('employment')], 180)).toBeNull();
});
});
describe('canRenderBar — R1', () => {
it('allows a bar when the whole group is published', () => {
expect(canRenderBar(fullGroup())).toBe(true);
});
it('refuses a bar when a single category is suppressed', () => {
const g = fullGroup();
g.cells[1] = sup('sixth_form_college');
expect(canRenderBar(g)).toBe(false);
});
});
describe('toBarSegments', () => {
it('derives widths from counts, not from rounded percentages', () => {
const segs = toBarSegments(fullGroup());
expect(segs).toHaveLength(8);
expect(segs[0].widthPct).toBeCloseTo((75 / 180) * 100, 10);
expect(segs.reduce((a, s) => a + s.widthPct, 0)).toBeCloseTo(100, 6);
});
it('throws rather than silently leaving a gap when the group is suppressed', () => {
const g = fullGroup();
g.cells[1] = sup('sixth_form_college');
expect(() => toBarSegments(g)).toThrow(/suppressed/i);
});
});
describe('CARD_GROUPS', () => {
it('partitions every destination category exactly once, plus the absence', () => {
const grouped = Object.values(CARD_GROUPS).flat();
expect(new Set(grouped).size).toBe(grouped.length);
expect(grouped).toEqual(expect.arrayContaining([
'school_sixth_form', 'sixth_form_college', 'further_education',
'other_education', 'apprenticeship', 'employment',
]));
expect(grouped).not.toContain('not_sustained');
expect(grouped).not.toContain('not_captured');
});
});
+25 -1
View File
@@ -1,4 +1,4 @@
import { getFlags } from '@/lib/flags';
import { getFlags, FLAGS_REVALIDATE } from '@/lib/flags';
// jsdom provides no global fetch, so there is nothing for jest.spyOn to attach
// to — assign it and restore the original afterwards. This is the first test
@@ -31,4 +31,28 @@ describe('getFlags', () => {
mockFetch(async () => ({ ok: false, status: 503 }));
await expect(getFlags()).resolves.toEqual({});
});
/*
* Reading a flag pins the calling route's ISR floor: Next uses the LOWEST
* revalidate among a route's fetches for the whole route. That is why the
* revalidate is an argument rather than the constant.
*
* Every SEO route here declares `revalidate = 604800`. A gate that read
* flags at the 300s default would drop the whole school and place corpus
* from a weekly cache to a 5-minute one, which is a large origin-load
* regression to pay for a feature flag.
*/
it('reads at the 300s floor by default', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags();
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: FLAGS_REVALIDATE } });
});
it('lets a caller pass its own route floor instead', async () => {
mockFetch(async () => ({ ok: true, json: async () => ({}) }));
await getFlags(604800);
expect((global.fetch as jest.Mock).mock.calls[0][1])
.toEqual({ next: { revalidate: 604800 } });
});
});
@@ -0,0 +1,68 @@
/**
* School pages had no BreadcrumbList and no links into the location layer.
* Both are fixed by the same data — the `places` array the API now returns —
* so they are tested together.
*/
import { schoolBreadcrumbJsonLd } from '@/lib/jsonld';
const essex = { kind: 'authority', slug: 'essex', name: 'Essex', count: 480, url: '/schools/authority/essex', phases: [] };
const brentwood = { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 37, url: '/schools/brentwood', phases: [] };
const outcode = { kind: 'outcode', slug: 'cm15', name: 'CM15', count: 12, url: '/schools/near/cm15', phases: [] };
describe('school breadcrumbs', () => {
it('reads home to authority to town to school', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood],
});
expect(ld['@type']).toBe('BreadcrumbList');
expect(ld.itemListElement.map((i) => i.name))
.toEqual(['schoolcompare', 'Essex', 'Brentwood', 'Brentwood School']);
expect(ld.itemListElement.map((i) => i.position)).toEqual([1, 2, 3, 4]);
});
it('skips a level the school has no published place for', () => {
// A school whose town falls below the publish threshold has no town page.
// The trail closes over the gap rather than linking to a 404.
const ld = schoolBreadcrumbJsonLd({
name: 'Lone School', url: '/school/1-lone-school', places: [essex],
});
expect(ld.itemListElement.map((i) => i.name))
.toEqual(['schoolcompare', 'Essex', 'Lone School']);
expect(ld.itemListElement.map((i) => i.position)).toEqual([1, 2, 3]);
});
it('omits outcodes, which are not a place a breadcrumb reads through', () => {
// CM15 is a useful link in the module but nonsense in a trail: nobody
// navigates Essex → CM15 → school.
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood, outcode],
});
expect(JSON.stringify(ld)).not.toContain('cm15');
});
it('still produces a valid trail when the school has no places at all', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Orphan School', url: '/school/2-orphan-school', places: [],
});
expect(ld.itemListElement.map((i) => i.name)).toEqual(['schoolcompare', 'Orphan School']);
});
it('uses absolute urls, as every other entity on the site does', () => {
const ld = schoolBreadcrumbJsonLd({
name: 'Brentwood School', url: '/school/100000-brentwood-school',
places: [essex, brentwood],
});
for (const item of ld.itemListElement) {
expect(item.item).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\//);
}
// The root is the homepage: there is no /schools index page to link to.
expect(ld.itemListElement[0].item).toBe('https://www.schoolcompare.co.uk/');
});
});
@@ -0,0 +1,83 @@
import { computeSecondaryFlags, buildSecondaryNavItems } from '@/lib/schoolSections';
import type { School, SchoolDestinations } from '@/lib/types';
const schoolInfo = {
urn: 137083, school_name: 'Northbrook Academy', phase: 'Secondary',
has_sixth_form: true,
} as unknown as School;
const base = { schoolInfo, yearlyData: [], deprivation: null, finance: null };
const phase = (categories = 1) => ({
cohort_year: '2022/23',
groups: {
all: {
cohort: 180,
categories: Array.from({ length: categories }, () => ({
category: 'school_sixth_form' as const,
pupils: 75, percentage: 41.7, status: 'published' as const,
})),
},
},
});
const ks4Only: SchoolDestinations = { ks4: phase(), ks5: null };
const both: SchoolDestinations = { ks4: phase(), ks5: phase() };
describe('computeSecondaryFlags — destinations', () => {
it('flags KS4 destinations when the block carries categories', () => {
const flags = computeSecondaryFlags({ ...base, destinations: ks4Only });
expect(flags.hasKs4Destinations).toBe(true);
expect(flags.hasKs5Destinations).toBe(false);
});
it('flags both phases when both are present', () => {
const flags = computeSecondaryFlags({ ...base, destinations: both });
expect(flags.hasKs4Destinations).toBe(true);
expect(flags.hasKs5Destinations).toBe(true);
});
it('flags neither when the block is absent', () => {
const flags = computeSecondaryFlags({ ...base, destinations: null });
expect(flags.hasKs4Destinations).toBe(false);
expect(flags.hasKs5Destinations).toBe(false);
});
it('does not flag a phase whose groups carry no categories', () => {
const empty: SchoolDestinations = {
ks4: { cohort_year: '2022/23', groups: {} }, ks5: null,
};
expect(computeSecondaryFlags({ ...base, destinations: empty }).hasKs4Destinations)
.toBe(false);
});
it('does not flag a phase whose only group has an empty category list', () => {
const empty: SchoolDestinations = { ks4: phase(0), ks5: null };
expect(computeSecondaryFlags({ ...base, destinations: empty }).hasKs4Destinations)
.toBe(false);
});
});
describe('buildSecondaryNavItems — destinations', () => {
const navInput = {
ofsted: null, admissions: null, admissionDistance: null,
hasLocation: false, yearlyDataLength: 0,
};
it('adds both entries, after GCSEs', () => {
const flags = computeSecondaryFlags({ ...base, destinations: both });
const ids = buildSecondaryNavItems({ ...flags, hasResults: true }, navInput)
.map(i => i.id);
expect(ids).toContain('destinations');
expect(ids).toContain('post16-destinations');
expect(ids.indexOf('destinations')).toBeGreaterThan(ids.indexOf('gcse'));
expect(ids.indexOf('post16-destinations')).toBe(ids.indexOf('destinations') + 1);
});
it('adds no entry for a phase that will not render — the nav must not link to a missing anchor', () => {
const flags = computeSecondaryFlags({ ...base, destinations: null });
const ids = buildSecondaryNavItems(flags, navInput).map(i => i.id);
expect(ids).not.toContain('destinations');
expect(ids).not.toContain('post16-destinations');
});
});
@@ -172,3 +172,39 @@ describe('buildSecondaryNavItems', () => {
expect(ids).not.toContain('history');
});
});
describe('the nearby-schools nav item', () => {
const navInput = {
ofsted: null, admissions: null, admissionDistance: null,
hasLocation: true, yearlyDataLength: 1,
};
it('appears on both templates when the section renders', () => {
const primary = computeSchoolFlags(primaryFixture);
const secondary = computeSecondaryFlags(secondaryFixture);
const input = { ...navInput, hasNearbySchools: true };
expect(buildNavItems(primary, input).map((i) => i.id)).toContain('nearby');
expect(buildSecondaryNavItems(secondary, input).map((i) => i.id)).toContain('nearby');
});
it('is absent when the section does not render', () => {
const primary = computeSchoolFlags(primaryFixture);
const secondary = computeSecondaryFlags(secondaryFixture);
const input = { ...navInput, hasNearbySchools: false };
expect(buildNavItems(primary, input).map((i) => i.id)).not.toContain('nearby');
expect(buildSecondaryNavItems(secondary, input).map((i) => i.id)).not.toContain('nearby');
});
it('is absent when nothing says either way', () => {
const primary = computeSchoolFlags(primaryFixture);
expect(buildNavItems(primary, navInput).map((i) => i.id)).not.toContain('nearby');
});
it('comes last, because the section renders last', () => {
const primary = computeSchoolFlags(primaryFixture);
const ids = buildNavItems(primary, { ...navInput, hasNearbySchools: true }).map((i) => i.id);
expect(ids[ids.length - 1]).toBe('nearby');
});
});
+26
View File
@@ -13,6 +13,8 @@ import {
metricKind,
shortName,
computeYBounds,
formatAgeRange,
formatAgeSpan,
} from '@/lib/utils';
describe('formatPercentage', () => {
@@ -320,3 +322,27 @@ describe('shortName', () => {
expect(shortName('A'.repeat(30), 10)).toBe('AAAAAAAAA…');
});
});
describe('formatAgeSpan', () => {
it('normalises a hyphenated range to an en dash, without a label', () => {
// The place table carries "Ages" in the column heading, so repeating it
// in every cell is noise. formatAgeRange keeps the label for the contexts
// that have no heading to hang it on.
expect(formatAgeSpan('4-11')).toBe('4–11');
});
it('leaves a range it does not recognise alone rather than mangling it', () => {
expect(formatAgeSpan('3-19 (SEN)')).toBe('3-19 (SEN)');
});
it('returns an empty string for a missing range', () => {
expect(formatAgeSpan(null)).toBe('');
expect(formatAgeSpan(undefined)).toBe('');
});
});
describe('formatAgeRange', () => {
it('keeps its label, so the two helpers stay distinguishable', () => {
expect(formatAgeRange('4-11')).toBe('Ages 4–11');
});
});
@@ -0,0 +1,70 @@
/**
* Payload is ESM-only and next/jest will not transform it, so the collections
* cannot be imported and their sanitised config inspected here (see
* lib/payloadRoutes.ts for the full reasoning). These assert the source of the
* collection definitions instead — enough to catch the settings whose loss is
* silent, and cheap. Behaviour is proved by the e2e journeys against staging.
*/
import fs from 'fs';
import path from 'path';
const read = (file: string) =>
fs.readFileSync(path.join(__dirname, '..', '..', 'collections', file), 'utf8');
const POSTS = read('Posts.ts');
const MEDIA = read('Media.ts');
const CONFIG = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'),
'utf8',
);
describe('posts collection', () => {
it('supports drafts, so saving is not publishing', () => {
expect(POSTS).toMatch(/drafts:\s*true/);
});
it('has a unique, indexed slug for stable URLs', () => {
const slugField = POSTS.slice(POSTS.indexOf("name: 'slug'"));
expect(slugField).toMatch(/unique:\s*true/);
expect(slugField).toMatch(/index:\s*true/);
});
it('hides drafts from anonymous readers at the access layer', () => {
// Payload's docs are explicit: "The `draft` argument alone does not
// restrict documents with _status: 'draft' from being returned by the
// API." The blog pages' where-clause is not enforcement — a direct GET
// /cms-api/posts would return unpublished drafts to anyone. Access
// control returning a query constraint is the only thing that stops it.
expect(POSTS).toMatch(/_status:\s*\{\s*equals:\s*'published'\s*\}/);
expect(POSTS).toMatch(/if\s*\(req\.user\)\s*return true/);
});
it('revalidates the post page when a post changes or is deleted', () => {
// /blog/[slug] is ISR — generated on first request and cached — so an edit
// to an already-published post would otherwise not appear until the
// revalidate window expired, up to an hour of a writer concluding that
// saving is broken. The index and feeds are force-dynamic and need no hook.
expect(POSTS).toContain('afterChange');
expect(POSTS).toContain('afterDelete');
expect(POSTS).toMatch(/revalidatePath\(`\/blog\/\$\{[^}]+\}`\)/);
});
});
describe('media collection', () => {
it('writes uploads to the mounted volume, by absolute path', () => {
// Must match the payload_media mount in docker-compose.portainer.yml.
// Payload 3 requires staticDir to be absolute.
expect(MEDIA).toMatch(/staticDir:\s*'\/app\/media'/);
});
it('requires alt text on every upload', () => {
const altField = MEDIA.slice(MEDIA.indexOf("name: 'alt'"));
expect(altField).toMatch(/required:\s*true/);
});
});
describe('payload config', () => {
it('registers every collection', () => {
expect(CONFIG).toMatch(/collections:\s*\[Users,\s*Posts,\s*Media\]/);
});
});
@@ -0,0 +1,67 @@
/**
* The admin panel does not import field components directly. Payload sends the
* client a *path* for each one — a richText field's is
* `@payloadcms/richtext-lexical/rsc#RscEntryLexicalField` — and resolves it
* through this generated map. An entry that is missing from the map is not an
* error the panel reports: the field simply does not render.
*
* That failure is quietly awful, because `required: true` is enforced on the
* server regardless. A writer gets a new-post form with no Content editor and
* a save that refuses on a field they were never shown.
*
* The map is generated by `npx payload generate:importmap`, so it drifts every
* time a field or a lexical feature is added and nobody re-runs it. These
* assert the entries the current config needs.
*/
import fs from 'fs';
import path from 'path';
const MAP = fs.readFileSync(
path.join(__dirname, '..', '..', 'app', '(payload)', 'admin', 'importMap.js'),
'utf8',
);
const POSTS = fs.readFileSync(
path.join(__dirname, '..', '..', 'collections', 'Posts.ts'),
'utf8',
);
describe('admin import map', () => {
it('resolves the richText field, so Content renders in the editor', () => {
// Guarded because Posts.content is required: without this entry the field
// is invisible and the post is unsaveable.
expect(POSTS).toMatch(/type:\s*'richText'/);
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#RscEntryLexicalField');
});
it('resolves the richText cell, so the list view can render the column', () => {
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#RscEntryLexicalCell');
});
it('resolves the diff component, which the drafts UI needs', () => {
// versions.drafts is on, so the panel offers version comparison.
expect(POSTS).toMatch(/drafts:\s*true/);
expect(MAP).toContain('@payloadcms/richtext-lexical/rsc#LexicalDiffComponent');
});
it('resolves BlocksFeature, so the Callout block is insertable', () => {
expect(POSTS).toContain('BlocksFeature');
expect(MAP).toContain('@payloadcms/richtext-lexical/client#BlocksFeatureClient');
});
it('resolves the default toolbar features the editor is built with', () => {
// defaultFeatures is spread into the editor config; each one contributes a
// client component the toolbar cannot render without.
for (const feature of [
'BoldFeatureClient',
'ItalicFeatureClient',
'HeadingFeatureClient',
'LinkFeatureClient',
'UploadFeatureClient',
'UnorderedListFeatureClient',
'OrderedListFeatureClient',
'InlineToolbarFeatureClient',
]) {
expect(MAP).toContain(`@payloadcms/richtext-lexical/client#${feature}`);
}
});
});
@@ -0,0 +1,52 @@
/**
* The generated migration is schema-qualified to "payload" throughout but does
* not create that schema — `schemaName` says where tables go, it does not
* create anything. On staging and production, which have never run it, the
* whole migration fails with `schema "payload" does not exist`.
*
* The CREATE SCHEMA is therefore hand-added, which makes it exactly the kind
* of edit a regeneration silently discards. This is the guard.
*/
import fs from 'fs';
import path from 'path';
const DIR = path.join(__dirname, '..', '..', 'migrations');
function migrationFiles() {
return fs
.readdirSync(DIR)
.filter((f) => f.endsWith('.ts') && f !== 'index.ts');
}
describe('payload migrations', () => {
it('ships at least one migration, so a container has tables to find', () => {
expect(migrationFiles().length).toBeGreaterThan(0);
});
it('creates the payload schema before creating anything in it', () => {
const initial = migrationFiles().find((f) => f.includes('initial'))!;
const sql = fs.readFileSync(path.join(DIR, initial), 'utf8');
expect(sql).toMatch(/CREATE SCHEMA IF NOT EXISTS "payload"/);
// Ordering matters: the schema must be created before the first object
// that lives in it, or the migration fails on its first statement.
expect(sql.indexOf('CREATE SCHEMA IF NOT EXISTS "payload"'))
.toBeLessThan(sql.indexOf('CREATE TABLE "payload"'));
});
it('creates the tables the app queries on boot', () => {
const initial = migrationFiles().find((f) => f.includes('initial'))!;
const sql = fs.readFileSync(path.join(DIR, initial), 'utf8');
for (const table of ['users', 'posts', '_posts_v', 'media', 'payload_migrations']) {
expect(sql).toContain(`CREATE TABLE "payload"."${table}"`);
}
});
it('is wired into the adapter, so it runs on server init', () => {
const config = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'), 'utf8',
);
expect(config).toMatch(/prodMigrations:\s*migrations/);
});
});
@@ -0,0 +1,44 @@
/**
* Guards the one thing about Payload's mounting that fails silently.
*
* payload.config.ts itself cannot be imported here — Payload is ESM-only and
* next/jest will not transform it — so this asserts the shared constants and
* that the config actually wires them in, by reading its source. The live
* proof that /api still reaches FastAPI is the e2e journeys, which call
* /api/schools against the running app.
*/
import fs from 'fs';
import path from 'path';
import { PAYLOAD_API_ROUTE, PAYLOAD_ADMIN_ROUTE } from '@/lib/payloadRoutes';
const CONFIG = fs.readFileSync(
path.join(__dirname, '..', '..', 'payload.config.ts'),
'utf8',
);
describe('payload mount points', () => {
it('serves the CMS API from /cms-api, never /api', () => {
// /api is the FastAPI proxy's catch-all. Payload's default would be
// swallowed by it and forwarded to the backend, silently.
expect(PAYLOAD_API_ROUTE).toBe('/cms-api');
expect(PAYLOAD_API_ROUTE).not.toBe('/api');
});
it('serves the admin panel from /admin', () => {
expect(PAYLOAD_ADMIN_ROUTE).toBe('/admin');
});
it('wires both constants into the Payload config', () => {
expect(CONFIG).toContain('PAYLOAD_API_ROUTE');
expect(CONFIG).toContain('PAYLOAD_ADMIN_ROUTE');
});
it('never hardcodes a routes block that could drift from the constants', () => {
expect(CONFIG).not.toMatch(/routes:\s*\{[^}]*api:\s*['"]/);
});
it('isolates CMS tables in their own postgres schema', () => {
// Blog content must stay separate from school marts and Airflow metadata.
expect(CONFIG).toMatch(/schemaName:\s*['"]payload['"]/);
});
});
@@ -21,7 +21,7 @@ import {
import { nationalAveragesFixture } from './schoolFixtures';
// The shell calls useComparison(), which throws outside the provider. In the
// app this wrapper comes from app/layout.tsx.
// app this wrapper comes from app/(frontend)/layout.tsx.
function withProviders(ui: ReactNode) {
return <ComparisonProvider>{ui}</ComparisonProvider>;
}
@@ -0,0 +1,82 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.header {
display: flex;
align-items: center;
gap: 1.25rem;
margin-bottom: 2rem;
}
.portrait {
border-radius: 50%;
border: 2px solid var(--border);
object-fit: cover;
flex-shrink: 0;
}
.kicker {
font-family: var(--font-ui);
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--brand);
margin: 0 0 0.35rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.5rem, 4vw, 2rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0;
}
.subheading {
font-family: var(--font-display);
font-size: 1.15rem;
font-weight: 600;
color: var(--text-primary);
margin: 2.25rem 0 0.75rem;
}
.prose p {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
margin: 0 0 1.1rem;
}
/* The opening paragraph carries the page. Larger, and in the primary ink
rather than the secondary, so it reads as a voice rather than as body copy.
Must stay in the descendant form: `.prose p` scores (0,1,1) and would beat a
bare `.lede` at (0,1,0), so simplifying this selector silently reverts the
lede to ordinary body copy. */
.prose .lede {
font-size: 1.125rem;
color: var(--text-primary);
}
.link {
color: var(--brand);
font-weight: 600;
}
.link:hover {
color: var(--brand-strong);
}
@media (max-width: 480px) {
.header {
flex-direction: column;
align-items: flex-start;
gap: 1rem;
}
}
+143
View File
@@ -0,0 +1,143 @@
import type { Metadata } from 'next';
import Image from 'next/image';
import { notFound } from 'next/navigation';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import { personJsonLd, organizationJsonLd } from '@/lib/jsonld';
import styles from './About.module.css';
export const metadata: Metadata = {
title: 'About',
description:
'Who builds schoolcompare, why it exists, and where its numbers come from.',
alternates: { canonical: absoluteUrl('/about') },
};
/*
* Gated on about_page. The default 300s read is the right floor here: this
* page declares no revalidate of its own, so nothing is lost by it, and a flip
* lands within five minutes.
*
* notFound(), not a redirect: while the flag is dark this URL does not exist,
* and a 404 is what tells a crawler not to keep it.
*/
export default async function AboutPage() {
const flags = await getFlags();
if (flags.about_page !== true) notFound();
const jsonLd = {
'@context': 'https://schema.org',
'@graph': [personJsonLd(), organizationJsonLd()],
};
return (
<div className={styles.page}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<header className={styles.header}>
<Image
src="/brand/tudor.jpg"
alt="Tudor, who builds schoolcompare"
width={96}
height={96}
className={styles.portrait}
priority
/>
<div>
<p className={styles.kicker}>Who&apos;s behind this</p>
<h1 className={styles.heading}>I&apos;m Tudor. I built this site.</h1>
</div>
</header>
<div className={styles.prose}>
<p className={styles.lede}>
I&apos;m a parent in south-west London. When we started looking at
primary schools, I found the information I needed was all published,
and almost impossible to hold in one place.
</p>
<p>
SATs results were in one government table. Ofsted judgements were in a
separate service, in a format that had just changed. Admissions
distances were buried in council PDFs, a different one per borough,
each with its own layout. I ended up building a spreadsheet, and then
I got tired of the spreadsheet.
</p>
<p>
So I built this instead. It pulls the official figures into one place
and puts them side by side, which is what I wanted and could not find.
</p>
<h2 className={styles.subheading}>I&apos;m not an education expert</h2>
<p>
I want to be straightforward about that. I&apos;m not a teacher, a
governor, or an education researcher. I have no qualification that
makes my opinion about a school worth more than yours.
</p>
<p>
What I do have is the problem itself. I&apos;m going through primary
admissions right now, and I work with data for a living. That
combination is enough to take published figures and present them
honestly. It is not enough to tell you which school is right for your
child, and this site never tries to.
</p>
<h2 className={styles.subheading}>Where the numbers come from</h2>
<p>
Everything here is official published data: Key Stage 2 and Key Stage
4 results and school characteristics from the Department for
Education, inspection outcomes from Ofsted, and admissions data from
local authorities. Nothing is estimated, modelled or filled in. Where
a figure is missing, the page says so rather than showing a guess.
</p>
<p>
This is an independent site. It is not affiliated with the Department
for Education or with Ofsted, and nobody pays to appear on it or to
rank higher.
</p>
<h2 className={styles.subheading}>What the data can&apos;t tell you</h2>
<p>
A school is not its results. The figures here describe one year group,
on a handful of days, measured in a way that suits national statistics
rather than your child. A small cohort makes percentages swing wildly.
In a class of thirty, one pupil is worth more than three points.
Results say nothing at all about whether a child will be happy
somewhere.
</p>
<p>
I try to build that honesty into the site rather than just say it
here. Special schools and pupil referral units are never compared
against a mainstream national average, because that comparison is
meaningless and makes good schools look like failing ones. Where a
number is unreliable, the aim is for the page to tell you before you
draw a conclusion from it.
</p>
<h2 className={styles.subheading}>If something&apos;s wrong</h2>
<p>
Tell me and I&apos;ll fix it. If a figure looks wrong, or a page gives
a misleading impression of a school, I genuinely want to know.
It&apos;s the fastest way this gets better.
</p>
<p>
<a href="mailto:contact@schoolcompare.co.uk" className={styles.link}>
contact@schoolcompare.co.uk
</a>
</p>
</div>
</div>
);
}
@@ -0,0 +1,76 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.header { margin-bottom: 2.5rem; }
.kicker {
font-family: var(--font-ui);
font-size: 0.75rem;
font-weight: 600;
text-transform: uppercase;
letter-spacing: 0.06em;
color: var(--brand);
margin: 0 0 0.35rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.5rem, 4vw, 2rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0 0 0.75rem;
}
.standfirst {
font-family: var(--font-ui);
font-size: 1.05rem;
line-height: 1.65;
color: var(--text-secondary);
margin: 0;
}
.list { list-style: none; padding: 0; margin: 0; }
.item {
padding: 1.5rem 0;
border-top: 1px solid var(--border);
}
.date {
font-family: var(--font-ui);
font-size: 0.8rem;
color: var(--text-muted);
/* Inter's tabular numerals keep a column of dates aligned. */
font-variant-numeric: tabular-nums;
}
.itemTitle {
font-family: var(--font-display);
font-size: 1.25rem;
font-weight: 600;
line-height: 1.3;
margin: 0.35rem 0 0.5rem;
}
.itemLink { color: var(--text-primary); text-decoration: none; }
.itemLink:hover { color: var(--brand); }
.excerpt {
font-family: var(--font-ui);
font-size: 0.95rem;
line-height: 1.65;
color: var(--text-secondary);
margin: 0;
}
.empty {
font-family: var(--font-ui);
color: var(--text-muted);
}
.link { color: var(--brand); font-weight: 600; }
.link:hover { color: var(--brand-strong); }
@@ -0,0 +1,87 @@
.page {
max-width: 42rem;
margin: 0 auto;
padding: 2.5rem 1.25rem 4rem;
}
.crumb {
font-family: var(--font-ui);
font-size: 0.85rem;
margin-bottom: 1.25rem;
}
.heading {
font-family: var(--font-display);
font-size: clamp(1.6rem, 5vw, 2.25rem);
font-weight: 700;
line-height: 1.2;
color: var(--text-primary);
margin: 0 0 0.75rem;
}
.byline {
font-family: var(--font-ui);
font-size: 0.9rem;
color: var(--text-muted);
margin: 0 0 2rem;
}
.hero {
width: 100%;
height: auto;
border-radius: 10px;
border: 1px solid var(--border);
margin-bottom: 2rem;
}
/* Rich-text output: the editor emits plain elements, so these are styled by
descendant selector rather than by class. */
.prose p {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
margin: 0 0 1.1rem;
}
.prose h2 {
font-family: var(--font-display);
font-size: 1.25rem;
font-weight: 600;
color: var(--text-primary);
margin: 2.25rem 0 0.75rem;
}
.prose h3 {
font-family: var(--font-display);
font-size: 1.05rem;
font-weight: 600;
color: var(--text-primary);
margin: 1.75rem 0 0.6rem;
}
.prose ul,
.prose ol {
font-family: var(--font-ui);
font-size: 1rem;
line-height: 1.7;
color: var(--text-secondary);
padding-left: 1.35rem;
margin: 0 0 1.1rem;
}
.prose li { margin-bottom: 0.4rem; }
.prose a { color: var(--brand); font-weight: 500; }
.prose a:hover { color: var(--brand-strong); }
.prose blockquote {
border-left: 3px solid var(--border-strong);
padding-left: 1rem;
margin: 1.5rem 0;
color: var(--text-muted);
font-style: italic;
}
.link { color: var(--brand); font-weight: 600; }
.link:hover { color: var(--brand-strong); }
@@ -0,0 +1,194 @@
import { cache } from 'react';
import type { Metadata } from 'next';
import Link from 'next/link';
import { notFound } from 'next/navigation';
import { RichText } from '@payloadcms/richtext-lexical/react';
import type { JSXConvertersFunction } from '@payloadcms/richtext-lexical/react';
import { getCachedPayload } from '@/lib/payload';
import type { Post, Media } from '@/payload-types';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import {
blogPostingJsonLd,
breadcrumbJsonLd,
personJsonLd,
organizationJsonLd,
} from '@/lib/jsonld';
import { CalloutBlock } from '@/components/blog/CalloutBlock';
import styles from './Post.module.css';
/*
* ISR. Unlike the index, this route has a dynamic param and no
* generateStaticParams, so there is nothing for the build to prerender: each
* post is generated on first request and cached until the collection's
* afterChange hook revalidates it. That hook is what makes an edit to an
* already-published post appear immediately.
*/
export const revalidate = 3600;
/**
* heroImage is `number | Media | null`: an id when the query is shallow, the
* populated document at depth 1. Both pages query at depth 1, but narrowing
* rather than asserting keeps it correct if that ever changes.
*/
function heroOf(post: Post): Media | null {
return typeof post.heroImage === 'object' && post.heroImage !== null
? post.heroImage
: null;
}
/**
* Spreads the default converters and adds the one custom block.
*
* Without the spread, every default node type — paragraphs, headings, links —
* loses its renderer and the post body comes out empty.
*/
const calloutConverters: JSXConvertersFunction = ({ defaultConverters }) => ({
...defaultConverters,
blocks: {
// Annotated because the generic block converter cannot infer a custom
// block's field shape; String() guards the values regardless.
callout: ({ node }: { node: { fields: Record<string, unknown> } }) => (
<CalloutBlock
tone={String(node.fields.tone ?? 'caveat')}
body={String(node.fields.body ?? '')}
/>
),
},
});
/**
* Wrapped in React's cache() because Next calls generateMetadata and the page
* component separately for the same request — without it, every post view runs
* this query against Postgres twice. cache() dedupes within a single request
* only, so it never serves one visitor's request from another's.
*/
const findPost = cache(async (slug: string) => {
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { slug: { equals: slug }, _status: { equals: 'published' } },
limit: 1,
depth: 1,
});
return docs[0] ?? null;
});
function summarise(post: Post) {
return {
title: post.title,
slug: post.slug,
excerpt: post.excerpt,
publishedAt: post.publishedAt,
};
}
export async function generateMetadata(
{ params }: { params: Promise<{ slug: string }> },
): Promise<Metadata> {
const { slug } = await params;
const post = await findPost(slug);
if (!post) return { title: 'Not found' };
const hero = heroOf(post);
return {
title: post.title,
description: post.excerpt,
alternates: { canonical: absoluteUrl(`/blog/${post.slug}`) },
openGraph: {
type: 'article',
title: post.title,
description: post.excerpt,
url: absoluteUrl(`/blog/${post.slug}`),
publishedTime: post.publishedAt,
// A post with a hero image shares that; one without falls through to the
// generated share card at app/opengraph-image.tsx.
...(hero?.url ? { images: [{ url: hero.url }] } : {}),
},
};
}
export default async function PostPage(
{ params }: { params: Promise<{ slug: string }> },
) {
const { slug } = await params;
/*
* Flags read at this route's own declared floor, so gating costs it nothing.
* Checked before the post is fetched: a dark blog should not query Payload.
*/
const flags = await getFlags(3600);
if (flags.blog !== true) notFound();
const namedAuthor = flags.about_page === true;
const post = await findPost(slug);
if (!post) notFound();
const summary = summarise(post);
const hero = heroOf(post);
const jsonLd = {
'@context': 'https://schema.org',
/*
* The Person entity is anchored at /about#tudor, so it is declared only
* when that page exists. Claiming an author whose URL 404s is a worse
* signal than attributing the post to the publisher.
*/
'@graph': [
blogPostingJsonLd(summary, { namedAuthor }),
breadcrumbJsonLd(summary),
...(namedAuthor ? [personJsonLd()] : []),
organizationJsonLd(),
],
};
return (
<article className={styles.page}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<nav className={styles.crumb}>
<Link href="/blog" className={styles.link}>Blog</Link>
</nav>
<h1 className={styles.heading}>{summary.title}</h1>
<p className={styles.byline}>
{/* Unlinked while about_page is dark; the flags are independent. */}
By {namedAuthor
? <Link href="/about" className={styles.link}>Tudor</Link>
: 'Tudor'}
{' · '}
<time dateTime={summary.publishedAt}>
{new Date(summary.publishedAt).toLocaleDateString('en-GB', {
day: 'numeric',
month: 'long',
year: 'numeric',
})}
</time>
</p>
{/*
A plain <img>, not next/image: Payload already generated the sized
derivatives on upload (Media's imageSizes), so routing it through the
optimizer would resize an image that is already the right size.
*/}
{hero?.url && (
<img
className={styles.hero}
src={hero.url}
alt={hero.alt ?? ''}
width={hero.width ?? undefined}
height={hero.height ?? undefined}
/>
)}
<div className={styles.prose}>
<RichText data={post.content} converters={calloutConverters} />
</div>
</article>
);
}
+85
View File
@@ -0,0 +1,85 @@
import type { Metadata } from 'next';
import Link from 'next/link';
import { notFound } from 'next/navigation';
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import styles from './Blog.module.css';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
export const metadata: Metadata = {
title: 'Blog',
description:
'Notes on what school performance data shows, and what it does not.',
alternates: { canonical: absoluteUrl('/blog') },
};
function formatDate(value: string) {
return new Date(value).toLocaleDateString('en-GB', {
day: 'numeric',
month: 'long',
year: 'numeric',
});
}
export default async function BlogIndexPage() {
const flags = await getFlags();
if (flags.blog !== true) notFound();
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 50,
depth: 0,
});
return (
<div className={styles.page}>
<header className={styles.header}>
<p className={styles.kicker}>Blog</p>
<h1 className={styles.heading}>Notes on the numbers</h1>
<p className={styles.standfirst}>
What school performance data shows, what it doesn&apos;t, and how to
read it without being misled. Written by{' '}
{/* Plain text when about_page is dark: the two flags are
independent, so this link would otherwise point at a 404. */}
{flags.about_page === true
? <Link href="/about" className={styles.link}>Tudor</Link>
: 'Tudor'}.
</p>
</header>
{docs.length === 0 ? (
<p className={styles.empty}>No posts yet.</p>
) : (
<ul className={styles.list}>
{docs.map((post) => (
<li key={post.id} className={styles.item}>
<time className={styles.date} dateTime={String(post.publishedAt)}>
{formatDate(String(post.publishedAt))}
</time>
<h2 className={styles.itemTitle}>
<Link href={`/blog/${post.slug}`} className={styles.itemLink}>
{post.title}
</Link>
</h2>
<p className={styles.excerpt}>{post.excerpt}</p>
</li>
))}
</ul>
)}
</div>
);
}
@@ -0,0 +1,58 @@
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
function escapeXml(value: string): string {
return value.replace(/[<>&'"]/g, (char) =>
({ '<': '&lt;', '>': '&gt;', '&': '&amp;', "'": '&apos;', '"': '&quot;' }[char]!));
}
export async function GET() {
// A dark blog has no feed. 404 rather than an empty channel: an empty feed
// is a live feed with nothing in it, which a reader would keep polling.
const flags = await getFlags();
if (flags.blog !== true) return new Response('Not found', { status: 404 });
const payload = await getCachedPayload();
const { docs } = await payload.find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 50,
depth: 0,
});
const items = docs.map((post) => `
<item>
<title>${escapeXml(String(post.title))}</title>
<link>${absoluteUrl(`/blog/${post.slug}`)}</link>
<guid isPermaLink="true">${absoluteUrl(`/blog/${post.slug}`)}</guid>
<description>${escapeXml(String(post.excerpt))}</description>
<pubDate>${new Date(String(post.publishedAt)).toUTCString()}</pubDate>
</item>`).join('');
const xml = `<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
<channel>
<title>schoolcompare blog</title>
<link>${absoluteUrl('/blog')}</link>
<description>What school performance data shows, and what it does not.</description>
<language>en-GB</language>${items}
</channel>
</rss>`;
return new Response(xml, {
headers: { 'Content-Type': 'application/rss+xml; charset=utf-8' },
});
}
File renamed without changes.
@@ -0,0 +1,68 @@
/*
* A second sitemap for the URLs Next owns.
*
* /sitemap.xml is proxied from FastAPI (app/(frontend)/sitemap.xml), which
* knows nothing about Payload — the backend and frontend ship as separate
* images. Rather than teach it, the Next-owned URLs get their own sitemap and
* robots.txt lists both.
*/
import { getCachedPayload } from '@/lib/payload';
import { absoluteUrl } from '@/lib/site';
import { getFlags } from '@/lib/flags';
/*
* Dynamic, not ISR.
*
* This route has no dynamic params, so Next prerenders it at build time — and
* CI builds the image with no database reachable, which fails the build. It is
* a single indexed query against Postgres on the same Docker network, so
* rendering per request is cheap, and it means a newly published post appears
* here immediately rather than waiting on a revalidation.
*/
export const dynamic = 'force-dynamic';
export async function GET() {
/*
* A dark page must not be advertised. Submitting a URL that 404s is the one
* thing a sitemap is not allowed to do, so each entry is gated on the same
* flag that gates the page itself.
*
* With both flags dark this emits a valid, empty <urlset> rather than a 404:
* robots.txt names this sitemap unconditionally, and an empty sitemap is a
* well-formed statement that there is nothing here yet.
*/
const flags = await getFlags();
const aboutEnabled = flags.about_page === true;
const blogEnabled = flags.blog === true;
// Only query Payload when the blog is actually being advertised.
const docs = blogEnabled
? (await (await getCachedPayload()).find({
collection: 'posts',
where: { _status: { equals: 'published' } },
sort: '-publishedAt',
limit: 500,
depth: 0,
})).docs
: [];
const urls: Array<{ loc: string; lastmod: string | null }> = [
...(aboutEnabled ? [{ loc: absoluteUrl('/about'), lastmod: null }] : []),
...(blogEnabled ? [{ loc: absoluteUrl('/blog'), lastmod: null }] : []),
...docs.map((post) => ({
loc: absoluteUrl(`/blog/${post.slug}`),
lastmod: new Date(String(post.updatedAt ?? post.publishedAt)).toISOString(),
})),
];
const xml = `<?xml version="1.0" encoding="UTF-8"?>
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
${urls.map(({ loc, lastmod }) =>
` <url><loc>${loc}</loc>${lastmod ? `<lastmod>${lastmod}</lastmod>` : ''}</url>`,
).join('\n')}
</urlset>`;
return new Response(xml, {
headers: { 'Content-Type': 'application/xml; charset=utf-8' },
});
}
+12
View File
@@ -0,0 +1,12 @@
'use client';
export default function ErrorPage({ reset }: { error: Error & { digest?: string }; reset: () => void }) {
return (
<main style={{ maxWidth: '48rem', margin: '4rem auto', padding: '1.5rem' }}>
<h1>We couldn’t load this page</h1>
<p>School information is temporarily unavailable. Please try again.</p>
<button type="button" onClick={reset}>Try again</button>
<p><a href="/">Return to school search</a></p>
</main>
);
}
@@ -105,6 +105,23 @@
--series-7: #0E7A86;
--series-8: #8A4A6B;
/* ── Destination measures ───────────────────────────────────────────
Education is one hue in three steps (school-like -> college-like) so the
education destinations read as one family; apprenticeship and employment
are separate hues. The absence is neutral and HATCHED, never a colour:
"activity not captured" covers independent schools, moving abroad and
training DfE holds no data on, so rendering it as a bad outcome would be
a factual error. The hatch is also the secondary encoding that rescues
the neutral/blue pair, which separates at only dE 7.6 as flat fills.
Every other adjacent pair clears dE 10.9 under protanopia. */
--dest-sixthform: #0F766E;
--dest-sfcollege: #4A9E96;
--dest-fecollege: #7CBFB8;
--dest-apprentice: #806200;
--dest-employment: #2F6F8F;
--dest-none: #6B7580;
--dest-none-hatch: rgba(107, 117, 128, 0.34);
/* ── Phase: category, desaturated so it stays under the status hues ── */
--phase-primary: #0F766E;
--phase-primary-bg: rgba(167, 215, 197, 0.40);
@@ -293,6 +310,17 @@
--series-7: #6FD0DC;
--series-8: #D99BB8;
/* Destinations. Not a naive inversion: the education ramp reverses
direction so its darkest step stays the one furthest from the
school, and each step is re-checked against the dark card. */
--dest-sixthform: #5FC7BB;
--dest-sfcollege: #3E9B92;
--dest-fecollege: #2A716B;
--dest-apprentice: #EFC658;
--dest-employment: #8FB4D9;
--dest-none: #8B9AA1;
--dest-none-hatch: rgba(139, 154, 161, 0.34);
--phase-primary: #5FC7BB;
--phase-primary-bg: rgba(95, 199, 187, 0.16);
--phase-primary-text: #8ADACF;
@@ -588,6 +616,35 @@ html .leaflet-bar a:hover {
color: var(--text-primary);
}
/*
* The popup, which leaflet.css paints `background: white; color: #333` on both
* the card and its tip. The content LeafletMapInner binds into it is themed —
* the school name and the headline figure are `var(--text-primary)` — so in
* dark mode that was #E9EEF0 on #FFFFFF, a contrast ratio of 1.17:1. The name
* and the number were the two least readable things on the page.
*
* Moving the surface onto --bg-card fixes every foreground at once rather than
* one at a time: the muted phase line goes 2.90:1 -> 5.45:1, the vs-national
* delta 1.94:1 -> 8.14:1, the Ofsted badge 1.74:1 -> 9.11:1. In light mode
* --bg-card is #FFFFFF, so the popup looks as it always did.
*/
html .leaflet-popup-content-wrapper,
html .leaflet-popup-tip {
background: var(--bg-card);
color: var(--text-primary);
}
/* Leaflet's own selector is `.leaflet-container a.leaflet-popup-close-button`
at 0,2,1 — an `html` prefix alone would lose to it. */
html .leaflet-container a.leaflet-popup-close-button {
color: var(--text-muted);
}
html .leaflet-container a.leaflet-popup-close-button:hover,
html .leaflet-container a.leaflet-popup-close-button:focus {
color: var(--text-primary);
}
/* Main content column */
.main {
max-width: 1400px;
@@ -4,8 +4,10 @@ import Script from 'next/script';
import { Navigation } from '@/components/Navigation';
import { Footer } from '@/components/Footer';
import { ComparisonToast } from '@/components/ComparisonToast';
import { RouteTrail } from '@/components/RouteTrail';
import { ComparisonProvider } from '@/context/ComparisonProvider';
import { SITE_URL } from '@/lib/site';
import { getFlags } from '@/lib/flags';
import './globals.css';
// Manrope carries headings and key messaging — the guideline's "friendly,
@@ -57,14 +59,32 @@ export const metadata: Metadata = {
authors: [{ name: 'schoolcompare' }],
manifest: '/manifest.json',
// No `icons` key on purpose: setting it here would override the file
// conventions. app/icon.svg and app/apple-icon.tsx are the source, and
// app/opengraph-image.tsx supplies og:image and twitter:image.
// conventions. app/icon.png and app/apple-icon.png are the source.
//
// og:image and twitter:image are NOT inherited from
// app/opengraph-image.tsx — see the note on openGraph.images below. The
// icon conventions do reach these pages; the opengraph-image one does not.
metadataBase: new URL(SITE_URL),
openGraph: {
type: 'website',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
/*
* Declared, not inherited.
*
* app/opengraph-image.tsx is a metadata file convention, and it does
* attach to routes in the app root segment — _not-found gets an og:image
* from it. It does not reach the site's pages, which live in the
* (frontend) route group whose own layout.tsx is a root layout. Staging
* served og:title, og:description, og:url, og:site_name and og:type with
* no og:image at all, so every link pasted into a chat rendered bare.
*
* The file stays at the app root: /robots.txt and /icon.png depend on it
* being there, and moving it is what broke those before. This points at
* the route it generates instead. metadataBase makes it absolute.
*/
images: ['/opengraph-image'],
url: SITE_URL,
siteName: 'schoolcompare',
},
@@ -74,14 +94,34 @@ export const metadata: Metadata = {
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
// The card is summary_large_image; claiming that and supplying no image
// is worse than claiming a summary card.
images: ['/opengraph-image'],
},
};
export default function RootLayout({
/*
* The footer's About and Blog links are flagged, which makes this the one
* place on the site that reads a flag on every route.
*
* 604800 is deliberate and load-bearing: it is the revalidate every SEO route
* here already declares. Next pins a route to the LOWEST revalidate among its
* fetches, so reading flags at the 300s default would drop the whole school
* and place corpus from a weekly cache to a 5-minute one — a large origin-load
* regression to hide two footer links.
*
* The cost is latency in one direction only. The pages themselves read the
* same flags at their own floors and flip within minutes; the footer links
* follow within a week. Turning a feature on early therefore shows the page
* before its footer link, which is harmless. Turning one off leaves a link to
* a 404 until the cache turns over, so a rollback that matters wants a purge.
*/
export default async function RootLayout({
children,
}: Readonly<{
children: React.ReactNode;
}>) {
const flags = await getFlags(604800);
return (
// The font variable classes must sit on <html>, not <body>. globals.css
// declares --font-display on :root as var(--font-manrope) and --font-ui as
@@ -114,6 +154,10 @@ export default function RootLayout({
/>
</head>
<body>
{/* Records every route so funnel attribution has a previous page to
name. document.referrer cannot: a soft navigation creates no
document, so the browser never updates it. */}
<RouteTrail />
<ComparisonProvider>
<a href="#main-content" className="skip-link">Skip to main content</a>
<Navigation />
@@ -121,7 +165,10 @@ export default function RootLayout({
{children}
</main>
<ComparisonToast />
<Footer />
<Footer
aboutEnabled={flags.about_page === true}
blogEnabled={flags.blog === true}
/>
</ComparisonProvider>
</body>
</html>
@@ -85,73 +85,50 @@ export default async function HomePage({ searchParams }: HomePageProps) {
params.has_sixth_form
);
// Fetch data on server with error handling
try {
const [filtersData, dataInfo] = await Promise.all([fetchFilters(), fetchDataInfo().catch(() => null)]);
// Failures propagate to the retryable error boundary.
const [filtersData, dataInfo] = await Promise.all([fetchFilters(), fetchDataInfo().catch(() => null)]);
// Only fetch schools if there are search parameters
let schoolsData;
if (hasSearchParams) {
schoolsData = await fetchSchools({
search: params.search,
local_authority: params.local_authority,
school_type: params.school_type,
phase: params.phase,
postcode: params.postcode,
radius,
page,
page_size: 50,
gender: params.gender,
admissions_policy: params.admissions_policy,
has_sixth_form: params.has_sixth_form,
});
} else {
// Empty state by default
schoolsData = { schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 };
}
const resolvedFilters = filtersData || { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
// `unique_schools`, not `total_schools` — the latter is not a field this
// endpoint returns, and reading it silently yielded null on every request.
const total = dataInfo?.unique_schools ?? null;
const years = dataInfo?.years_available ?? [];
return (
<HomeView
autosuggest={autosuggest}
initialSchools={schoolsData}
filters={resolvedFilters}
totalSchools={total}
howItWorks={hasSearchParams ? null : <HowItWorksSection />}
editorial={hasSearchParams ? null : (
<EditorialSection
totalSchools={total}
localAuthorityCount={resolvedFilters.local_authorities.length}
earliestYearLabel={years.length ? formatAcademicYear(years[0]) : null}
latestYearLabel={years.length ? formatAcademicYear(years[years.length - 1]) : null}
/>
)}
/>
);
} catch (error) {
console.error('Error fetching data for home page:', error);
const emptyFilters = { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
return (
<HomeView
autosuggest={autosuggest}
initialSchools={{ schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 }}
filters={emptyFilters}
totalSchools={null}
howItWorks={hasSearchParams ? null : <HowItWorksSection />}
editorial={hasSearchParams ? null : (
<EditorialSection
totalSchools={null}
localAuthorityCount={0}
earliestYearLabel={null}
latestYearLabel={null}
/>
)}
/>
);
// Only fetch schools if there are search parameters
let schoolsData;
if (hasSearchParams) {
schoolsData = await fetchSchools({
search: params.search,
local_authority: params.local_authority,
school_type: params.school_type,
phase: params.phase,
postcode: params.postcode,
radius,
page,
page_size: 50,
gender: params.gender,
admissions_policy: params.admissions_policy,
has_sixth_form: params.has_sixth_form,
});
} else {
// Empty state by default
schoolsData = { schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 };
}
const resolvedFilters = filtersData || { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
// `unique_schools`, not `total_schools` — the latter is not a field this
// endpoint returns, and reading it silently yielded null on every request.
const total = dataInfo?.unique_schools ?? null;
const years = dataInfo?.years_available ?? [];
return (
<HomeView
autosuggest={autosuggest}
initialSchools={schoolsData}
filters={resolvedFilters}
totalSchools={total}
howItWorks={hasSearchParams ? null : <HowItWorksSection />}
editorial={hasSearchParams ? null : (
<EditorialSection
totalSchools={total}
localAuthorityCount={resolvedFilters.local_authorities.length}
earliestYearLabel={years.length ? formatAcademicYear(years[0]) : null}
latestYearLabel={years.length ? formatAcademicYear(years[years.length - 1]) : null}
/>
)}
/>
);
}
File renamed without changes.
@@ -0,0 +1,17 @@
import { readFile } from 'node:fs/promises';
import path from 'node:path';
export const dynamic = 'force-dynamic';
export const runtime = 'nodejs';
export async function GET() {
try {
const frontend = JSON.parse(await readFile(path.join(process.cwd(), 'build-info.json'), 'utf8'));
const base = process.env.FASTAPI_URL || 'http://localhost:8000/api';
const res = await fetch(`${base}/release`, { cache: 'no-store', signal: AbortSignal.timeout(5000) });
if (!res.ok) throw new Error('Backend identity unavailable');
return Response.json({ frontend, backend: await res.json() }, { headers: { 'Cache-Control': 'no-store', 'X-Robots-Tag': 'noindex' } });
} catch {
return Response.json({ detail: 'Release identity unavailable' }, { status: 503, headers: { 'Cache-Control': 'no-store', 'X-Robots-Tag': 'noindex' } });
}
}
@@ -4,9 +4,12 @@
* URL format: /school/138267-school-name-here
*/
import { fetchSchoolDetails, fetchSchools, fetchNationalAverages } from '@/lib/api';
import { APIFetchError, fetchSchoolDetails, fetchSchools, fetchNationalAverages } from '@/lib/api';
import { notFound, redirect } from 'next/navigation';
import { SchoolDetailShell } from '@/components/school/SchoolDetailShell';
import { NearbyPlaces } from '@/components/school/NearbyPlaces';
import { shouldRenderNearby } from '@/components/school/NearbySchoolsSection';
import { schoolBreadcrumbJsonLd, type SchoolPlace } from '@/lib/jsonld';
import { PrimarySchoolSections } from '@/components/school/PrimarySchoolSections';
import { SecondarySchoolSections } from '@/components/school/SecondarySchoolSections';
import {
@@ -144,11 +147,17 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
fetchNationalAverages().catch(() => null),
]);
} catch (error) {
console.error(`Failed to fetch school ${urn}:`, error);
notFound();
if (error instanceof APIFetchError && error.status === 404) notFound();
throw error;
}
const { school_info, yearly_data, absence_data, ofsted, census, admissions, admissions_history, admission_distance, deprivation, finance } = data;
const { school_info, yearly_data, absence_data, ofsted, census, admissions, admissions_history, admission_distance, deprivation, finance, destinations } = data;
// Absent on an older API build; the module and the trail both degrade to
// nothing rather than throwing, which is how this shipped without a
// lockstep deploy of the two images.
const places: SchoolPlace[] = data.places ?? [];
// Absent on an older API build, exactly like `places` above.
const nearbySchools = data.nearby_schools ?? [];
// Redirect bare URN to canonical slug URL
const canonicalSlug = schoolUrl(urn, school_info.school_name).replace('/school/', '');
@@ -171,6 +180,7 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
schoolInfo: school_info, yearlyData: yearly_data,
absenceData: absence_data, census: census ?? null,
deprivation: deprivation ?? null, finance: finance ?? null,
destinations: destinations ?? null,
};
const primaryFlags = computeSchoolFlags(sectionInput);
const secondaryFlags = computeSecondaryFlags(sectionInput);
@@ -179,15 +189,25 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
admissions: admissions ?? null,
admissionDistance: admission_distance ?? null,
hasLocation: school_info.latitude != null && school_info.longitude != null,
hasNearbySchools: shouldRenderNearby(nearbySchools),
yearlyDataLength: yearly_data.length,
};
const primaryNavItems = buildNavItems(primaryFlags, navInput);
const secondaryNavItems = buildSecondaryNavItems(secondaryFlags, navInput);
// Generate JSON-LD structured data for SEO
/*
* `School`, not `EducationalOrganization`.
*
* Both are valid, but EducationalOrganization is the parent type covering
* universities, training providers and nurseries alike. School is the
* specific one, and a type that says what the page is about is the whole
* point of declaring it. Google's own guidance treats the narrower type as
* the correct choice where it applies.
*/
const structuredData = {
'@context': 'https://schema.org',
'@type': 'EducationalOrganization',
'@graph': [{
'@type': 'School',
name: school_info.school_name,
identifier: school_info.urn.toString(),
...(school_info.address && {
@@ -209,6 +229,15 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
...(school_info.school_type && {
additionalType: school_info.school_type,
}),
},
// The trail the page sits at the end of. School pages carried no
// breadcrumb at all, while every place page already emitted one.
schoolBreadcrumbJsonLd({
name: school_info.school_name,
url: `/school/${slug}`,
places,
}),
],
};
return (
@@ -232,10 +261,12 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
census={census ?? null}
admissions={admissions ?? null}
admissionsHistory={admissions_history ?? []}
admissionDistance={admission_distance ?? null}
admissionDistance={admission_distance}
deprivation={deprivation ?? null}
finance={finance ?? null}
nationalAvg={nationalAvg}
destinations={destinations ?? null}
nearbySchools={nearbySchools}
flags={secondaryFlags}
/>
</SchoolDetailShell>
@@ -258,10 +289,12 @@ export default async function SchoolPage({ params }: SchoolPageProps) {
deprivation={deprivation ?? null}
finance={finance ?? null}
nationalAvg={nationalAvg}
nearbySchools={nearbySchools}
flags={primaryFlags}
/>
</SchoolDetailShell>
)}
<NearbyPlaces places={places} />
</>
);
}
Loaded 100 of 205 files, more files were not shown because too many files have changed in this diff. Show more