Compare commits

..
Author SHA1 Message Date
TudorandClaude Opus 5 3236efa846 fix(map): the hero map's fade to the header was hardcoded white
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m7s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 10s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 49s
The fade between the map band and the school header ramped through
rgba(255,255,255,...) and landed on var(--bg-card). In the light theme
that is white into white and invisible, as designed. In the dark theme
it climbed to 95% WHITE and then met a near-black card, putting a bright
band across the full width exactly where the map should dissolve into
the title.

Fading to the colour the gradient lands on is the whole trick, and it
only works if that colour is a token — so --bg-card-rgb now exists in
both theme blocks, matching the --hero-ground-rgb precedent.

Two more defects in the same file, same cause, found while in there:

The controls floating over the map paired a hardcoded white background
with color: var(--text-primary), which resolves to #E9EEF0 in dark —
near-white text on a near-white button. These deliberately do NOT follow
the theme, because the map tiles are light in both, so the ink is now
literal too and says why. A themed token is the wrong tool for a surface
that never changes.

The loading skeleton swept 50% white across var(--bg-secondary), which
is a bright flash every 1.4s on a dark page. It now sweeps toward the
card colour, a shade lighter than the ground in both themes.

The guard is a stylesheet test rather than a render test, because the
bug is invisible in the theme it was written for. Verified by reverting
each fix in turn: it names .fade and .openHint exactly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 21:20:57 +01:00
tudor d8ccb5b733 Merge pull request 'feat(suggest): school autosuggest, and the rate-limit fix it needed first' (#127) from feat/school-autosuggest into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 5m49s
Reviewed-on: #127
2026-08-26 19:58:57 +00:00
TudorandClaude Opus 5 0fa1a292c7 fix(api): bound what a forged CF-Connecting-IP can buy
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Code review, both findings valid.

The design doc claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the X-Forwarded-For fallback was
forgeable. That is true only for traffic that actually passed through
Cloudflare, and nothing in this process can verify that it did. Reaching
the origin directly, both headers are equally attacker-controlled — and
rotating CF-Connecting-IP mints a fresh rate-limit bucket per request,
defeating per-client limits on every endpoint including the
DataFrame-heavy /api/schools. Against abuse that is worse than the
shared bucket it replaced, which at least capped everyone together.

So the ceiling comes back. I dropped it earlier arguing it belonged at
Cloudflare; that argument assumed the keying was sound, and it is not.
GlobalRateLimitMiddleware counts all /api/ traffic in a fixed window
against a total, independent of client identity, outermost so it refuses
before any work happens. Written by hand because slowapi cannot express
a global cap: default_limits and application_limits are both keyed by
key_func, and the latter needs middleware this app does not install.

It does not make the header trustworthy — it makes trusting it
survivable. The real fix is Authenticated Origin Pulls or an origin
firewall, now documented in DEPLOY.md as the open gap it is.

127.0.0.1 is exempt: the healthcheck curls localhost from inside the
container, and starving it would restart the container and turn a load
spike into an outage loop. Keyed on the peer address, never the Host
header, which the caller sets.

Second finding: suggest_schools_typesense promised "never raises" while
the parsing loop sat outside the try, so int(None) on a malformed
document would have made a keystroke a 500. The loop now skips bad rows
rather than dropping the whole list — and a hit with no document no
longer becomes a suggestion pointing at /school/0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:56:32 +01:00
tudor c3f044bd65 Merge pull request 'docs(flags): Unleash does not create flags by itself' (#126) from fix/flags-runbook into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 44s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 53s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 2m10s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m46s
Reviewed-on: #126
2026-08-26 19:48:58 +00:00
TudorandClaude Opus 5 d2115364ae test(e2e): autosuggest journeys, gated on the flag
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 31s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m13s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m40s
Feature state is read from its observable effect — whether the search
box is a combobox — because /api/flags is denied to the public on
purpose. Same approach as the distance journeys.

The three endpoint tests are ungated: /api/suggest is live whether or
not the UI is, which is what lets it be smoke-tested in an environment
where the feature is still dark.

The flag-off journey asserts the plain search still works, not just that
the combobox is absent. Verified against staging, where the flag is off:
it passes and the flag-on journey correctly skips.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:39:33 +01:00
TudorandClaude Opus 5 28cf0a342c feat(suggest): wire autosuggest into the search box behind a flag
Off means off — no combobox role, no listener, no fetch. A test asserts
the absence of the request, not just the absence of the dropdown,
because a hidden-but-fetching control would still be spending the rate
limit on a feature nobody can see.

Enter with no active option falls through to the form's submit handler
and searches the typed text exactly as before. The existing behaviour is
preserved, not replaced, and that has its own test.

Suppressed once the value parses as a postcode: the box takes a name OR
a postcode, and suggesting schools during postcode entry fights the user.

.omniBoxContainer gains position: relative — the dropdown is absolutely
positioned and without it would have anchored to the page instead.

Four render sites, all wired: page.tsx renders HomeView in the success
path AND the catch fallback, and HomeView renders FilterBar as hero AND
sticky. Missing any one would make the flag silently do nothing
somewhere.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:38:52 +01:00
TudorandClaude Opus 5 d88e77f459 feat(suggest): the dropdown, with combobox ARIA
Presentational only — it fetches nothing and owns no state, so the
fetching rules and the ARIA rules can be read separately.

onMouseDown, not onClick. The input's blur handler closes the list and
blur fires before click, so a click handler never runs: the classic bug
where a dropdown works perfectly by keyboard and is dead to the mouse.

The plan's CSS guessed at token names like --color-surface. The real
tokens are --bg-card, --border, --text-muted, --bg-secondary and
--shadow-soft, and all five are redefined in the dark theme — invented
names would have silently fallen back to hardcoded light values and
broken dark mode.

Local authority is rendered because there are many schools called
'St Mary's'; a list without it is unusable for exactly the query
autosuggest exists to serve, which is what the test asserts.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:36:39 +01:00
TudorandClaude Opus 5 06eb433db5 feat(suggest): debounced, abortable suggestion hook
The AbortController is correctness, not economy. Without it a slow
response for 'st' can land after the fast one for 'st marys' and replace
a correct list with a stale one — the classic autosuggest race.

No cache: 'no-store', unlike the compare modal's search. This is the one
endpoint where prefix queries repeat most across users, so discarding
the browser cache and the backend's ETag 304s would be throwing away the
cheapest win available.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:35:27 +01:00
TudorandClaude Opus 5 1a6d349dad feat(suggest): GET /api/suggest, cacheable and DataFrame-free
A dedicated endpoint rather than a mode of /api/schools, because that
path filters and sorts 25,000 pandas rows per query while holding the
GIL — affordable once per search, not once per keystroke. A test asserts
the distinction directly by making load_school_data raise and requiring
the endpoint to answer anyway.

Nothing errors on ordinary input: a short query, no matches, or
Typesense being down are all 200 with an empty list.

Cached deliberately. Prefix queries repeat enormously across users and
school names change once a year, so s-maxage plus the existing ETag
middleware turns most keystrokes into 304s.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:34:34 +01:00
TudorandClaude Opus 5 75d3534d82 feat(suggest): Typesense rows for autosuggest, no DataFrame
search_schools_typesense returns URNs, which forces the caller to
hydrate from the 25,000-row in-memory frame. Every field a suggestion
needs is already in the Typesense document, so this returns documents
and the caller needs no pandas at all — the difference between a query
that can run per keystroke and one that cannot.

Never raises. Typesense unreachable or erroring gives an empty list,
because a dropdown that quietly stops appearing is the right failure for
a keystroke path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:33:42 +01:00
TudorandClaude Opus 5 ff041544f2 fix(api): rate-limit per caller, not per proxy
The limiter keyed on request.client.host, which in staging and prod is
the Next container — the backend has no published ports and nothing else
can reach it. So every browser user on the site shared one 60/minute
bucket per route. Measured against staging: 70 concurrent requests to
/api/schools returned exactly 60 OK and 10 refused, from one machine.

CF-Connecting-IP first. Cloudflare fronts both environments and
overwrites any client-supplied value, which a parsed X-Forwarded-For
chain does not guarantee. The XFF fallback is forgeable only from inside
the Docker network.

Named rather than hidden: the shared bucket was an accidental global
throttle on a single-process backend, and correct per-user keying
removes it. A real global ceiling belongs at Cloudflare, which is
already in the path; slowapi cannot express one without a second Limiter
and middleware this app does not install.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:32:58 +01:00
TudorandClaude Opus 5 59265f78b6 docs(flags): Unleash does not create flags by itself
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m6s
PR Checks / Backend Smoke (pull_request) Successful in 9s
PR Checks / Build Backend (no push) (pull_request) Successful in 35s
PR Checks / Build Frontend (no push) (pull_request) Successful in 46s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 1m15s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 35s
The runbook said a flag 'appears in the Unleash UI after the backend has
evaluated it once'. That is wrong. SDKs read definitions from the server
and never register anything, and metrics for an unknown flag are
discarded — so a declared flag is evaluated on every request, stays
False forever, and never shows up until someone creates it by hand.

Found the way these things usually are: staging had been running the
flag code for a while and the UI was still empty.

Also names the environment trap while here — each stack's token is
scoped to one environment, so toggling the other does nothing visible
and looks like the flag is broken.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 20:11:45 +01:00
TudorandClaude Opus 5 6e0a278340 docs(suggest): implementation plan, eight tasks
Self-review caught three defects in the plan. .omniBoxContainer, the
wrapper the dropdown positions against, does not declare position:
relative — without it the list anchors to the page. The postcode
suppression test typed character by character, so it would have asserted
no request while 'NW1' legitimately fires one; it now sets the value in
one go. And the Enter-submits-search assertion needed waitFor, because
updateURL pushes inside startTransition.

Task 1 is the one to review hardest: it is the only unflagged change and
it alters rate limiting for every endpoint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:48:39 +01:00
TudorandClaude Opus 5 e651dd0d65 docs(suggest): the in-app global ceiling would not have worked
Reading slowapi rather than assuming: default_limits and
application_limits are both evaluated with the same key_func, so they
are per-client across routes, not global. And application_limits only
apply 'if in_middleware' — this app installs no SlowAPIMiddleware, so
they would never have fired at all.

A genuine global cap would need a second Limiter with a constant key
plus that middleware. Cloudflare is already in the path on both
environments and does this at the right layer, so the ceiling is named
as a follow-up there rather than built badly here.

The risk that leaves is stated plainly in the risks section instead of
being papered over with a mechanism that does not do the job.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:44:21 +01:00
TudorandClaude Opus 5 22c113fc29 docs(suggest): design for school autosuggest
The load-bearing finding is not about autosuggest. The rate limiter keys
on request.client.host, which in staging and prod is the Next container
— so all browser users share one 60/min bucket per route. Measured
against staging: 70 concurrent requests gave exactly 60 x 200 and
10 x 429. Eight concurrent searchers would 429 the site once each
keystroke costs a request, so the keying fix is part of this work.

Both environments are behind Cloudflare, which sets CF-Connecting-IP and
overwrites any client-supplied value — trustworthy in a way a parsed
X-Forwarded-For chain is not, and the backend is unreachable except
through the Next proxy.

Named honestly: the shared bucket has been an accidental global throttle
on a single-process backend, so correct per-user keying removes a
protection. A global ceiling ships with it rather than instead of it.

Suggestions come from Typesense alone. The existing search path filters
a 25,000-row DataFrame per query, which is exactly the cost a keystroke
endpoint cannot pay, so there is deliberately no DataFrame fallback.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-26 19:41:00 +01:00
tudor e953ee7c5f Merge pull request 'feat(flags): ship-dark feature flags, with last-distance-offered behind the first one' (#125) from feat/feature-flags into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 40s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m34s
Reviewed-on: #125
2026-08-23 11:34:56 +00:00
TudorandClaude Opus 5 413d86cc3c chore(flags): wire UNLEASH_URL and the SDK cache volume into the stacks
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 27s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m59s
Both variables default to empty, so an environment without Unleash has
every flag off — the correct dark state rather than a boot failure.

The cache volume is the mitigation for the one real regression risk in
this design: the SDK evaluates everything False until it syncs, so a
backend cold-starting with an empty cache while Unleash is unreachable
would make a *released* feature disappear. On a named volume the disk
cache survives a restart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:58:29 +01:00
TudorandClaude Opus 5 4f01fbdedb test(e2e): make the distance journeys fail loudly, not skip quietly
The existing distance journeys all skip when no school has a published
figure, which is right when the feature is off — and wrong when it is
supposed to be on and is silently broken, because that shows up as a
green run full of skips. The new gate fails in exactly that case.

Feature state is read from the data, not from /api/flags: the public
proxy denies that path on purpose, since it names unreleased features.
Presence of the admission_distance key is the observable effect.

Verified against staging, where the feature is currently on: the on-gate
passes, the off-gate skips, the existing eight distance journeys are
unaffected.

One honest caveat — the /api/flags check passes on staging today because
that image predates the endpoint, not because the denylist works. The
denylist itself is covered by the jest unit test; this is defence in
depth and becomes a real assertion once deployed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:57:34 +01:00
TudorandClaude Opus 5 c3ba7aae0d feat(flags): ship last-distance-offered dark behind a flag
One gate, at the source. The frontend needs no change: DistanceSection
already returns null when distance_m is missing, and the admissions
block already conditions on (admissions || admissionDistance). Only 57
local authorities publish cut-offs, so the off-path is the commonest
path on the site and is well covered already.

Absent, not null. /api/schools/ is public and unauthenticated, so a
field left in the payload is a published field — the reasoning already
recorded in c9a1892 when history was withheld. The two are also
different claims: null says this school has no cut-off, absent says
cut-offs are not being published at all. The frontend type now says so.

The feature is on main and live on staging and has never reached
production, which is what makes it the right first consumer: the flag
lets the code promote without the feature appearing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:56:33 +01:00
TudorandClaude Opus 5 54a30de0d8 feat(flags): server-side getFlags for the frontend
Ships without a consumer, deliberately. The first flag needs none — the
backend withholds the field and the page follows — but 'UI elements on
existing pages' is one of the three surfaces this capability exists for,
and a flag layer that cannot gate one is incomplete.

Never throws: an unreadable flag is a dark one, which matches the
backend's fail-closed default. A page that 500s because the flags
endpoint blinked would be a worse outcome than a hidden feature.

Reading flags pins the calling route to a 300s ISR floor, since Next
takes the lowest revalidate among a route's fetches. That matches what
/school/[slug] already sits at, and it is the same property that makes a
flip propagate without a webhook.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:55:32 +01:00
TudorandClaude Opus 5 c30ad1db07 feat(flags): serve /api/flags, and keep the public proxy off it
The endpoint and its exposure control ship together on purpose. The
moment /api/flags exists, app/api/[...path] forwards it — and the
response names every unreleased feature the codebase knows about, along
with whether it is on. Publishing that is the opposite of shipping dark.

Denied on an exact first-segment match, not a prefix, so /api/flagship
does not go down with /api/flags. Next reads the endpoint server-side
over the Docker network, which never transits the public proxy.

jest.setup.js now guards its browser globals. It runs for every suite,
including the one that declares @jest-environment node to exercise the
route handler — NextRequest needs Fetch API globals jsdom lacks, and
there is no window there to define matchMedia on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:54:26 +01:00
TudorandClaude Opus 5 7424cef7c6 feat(flags): the registry and a fail-closed Unleash client
Unleash holds flag state; it does not hold the list of flags. REGISTRY
is that list, because the SDK evaluates an unknown flag to False and
without a registry that is an undeclared False — indistinguishable from
a typo in a flag name.

Fail-closed throughout, and never raises: an unset UNLEASH_URL, an
unreachable server, a client that throws, an undeclared name — all
False. A flag layer that can 500 a request path or stop the API booting
is worse than one that is switched off.

Every flag defaults to False, with no per-flag override, because a flag
that defaults on is a kill switch and this is deliberately not one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:52:55 +01:00
TudorandClaude Opus 5 01ccbb8e82 feat(flags): add the Unleash stack and its runbook
Its own Portainer stack, belonging to neither application stack: a
staging redeploy must not be able to disturb production's flag state.

One instance serves both. OSS Unleash ships development and production
environments with environment-scoped client tokens, so the same flag
holds independent state in each — which is what lets a feature be on in
staging, where the E2E journeys exercise it, while production stays dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:51:35 +01:00
TudorandClaude Opus 5 c339c2f1a1 docs(flags): implementation plan, eight tasks
Task 1 is the Unleash stack and ends with a human step — the Portainer
deploy and the token generation cannot be automated from here. Nothing
else blocks on it: an unset UNLEASH_URL means every flag is False, which
is what local development and CI get, so the whole suite runs without a
flag server existing.

Self-review caught three defects in the plan itself. get_supplementary_data
takes (db, urn), not (urn), and the test DataFrame was minimised to the
point where the endpoint would have failed for reasons unrelated to
flags — both now copy the known-good shape from test_school_details.py.
The proxy test needs the node jest environment, since NextRequest wants
Fetch API globals jsdom does not provide. And the e2e off-state check
hardcoded a URN, so a 404 page would have satisfied it without proving
anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:46:15 +01:00
TudorandClaude Opus 5 e2ca3d79f9 docs(flags): drop the webhook — the seven-day premise was wrong
Next uses the LOWEST revalidate among a route's fetches, not the segment
value. School pages fetch school details at 300s and place pages fetch
national averages at 3600s, so the effective ISR period is five minutes
and one hour respectively — not the seven days the segment declares.

A flag flip therefore propagates on its own, well inside the monthly,
by-hand cadence these flags are for. That deletes two webhook
integrations, a revalidate route, a secret-in-query-string scheme, an
idempotency requirement, and the rule that every fetch carry a cache
tag — which was the part most likely to rot as fetches are added.

Two constraints survive: a flag must never gate content on a
force-static page, because app/admissions never revalidates; and a
route-family flag must rebuild the sitemap, deferred with the route case
since no flag in scope touches it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:40:35 +01:00
TudorandClaude Opus 5 c2364bf09e docs(flags): design for a ship-dark feature flag layer
Unleash self-hosted in its own Portainer stack, with FastAPI holding the
only SDK and Next reading flags through a tagged fetch.

The two hard parts are consequences of putting flag state in a service
rather than the repo: main stops being the whole truth about what is on,
and a flag can now change without the deploy that would have cleared the
caches. A code-declared registry bounds the first; webhook-driven
revalidateTag handles the second.

Cache tagging is deliberately coarse — every server fetch carries the
flags tag, not just the flags fetch itself. The first consumer proves
why: admission_distance changes the shape of /api/schools/{urn}, so a
narrow purge would leave ~25,000 school pages serving the pre-flip
render for a week, invisibly.

First consumer is the last-distance-offered feature, which is on main
and staging and has never reached production. It needs one gate, at the
API, because the frontend already no-ops on a missing field.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 09:47:58 +01:00
tudor 43e0621728 Merge pull request 'fix(places): phase links must stay in their own namespace' (#124) from fix/place-phase-links into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m31s
Reviewed-on: #124
2026-08-22 17:33:12 +00:00
TudorandClaude Opus 5 d1358cc00f fix(places): phase links must stay in their own namespace
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m37s
Every place page built its phase links as /schools/[slug]/[phase], the
shape that belongs to towns alone.

On an authority page that pointed into the town namespace. For 87 of the
151 authorities the target does not exist and the link 404s; for the
other 64 it resolves to the town of the same name — a different set of
schools, which is precisely the near-duplicate the two namespaces were
introduced to prevent. On an outcode page it 404s outright.

Two causes behind it, both a rule written twice and inherited by only
one of the places that needed it.

The authority phase route was in the spec and dropped by the plan, which
built the three bare routes and no fourth. The sitemap is generated from
the place registry, which was right about them all along, so 302
authority phase URLs have been submitted to Google and every one 404s.
Adding the route makes the sitemap true and serves a real query —
admissions are authority-run, so "primary schools in Kent" is how a
parent searches before they have settled on a town.

The outcode variants were the opposite: the registry computed phases for
outcodes although the spec gives them no route, and the sitemap knew to
skip them while the API did not. The registry now decides alone, and the
sitemap's duplicate of that rule is gone.

Also: an authority under the five-school threshold has no page, so the
API sends a null slug for it and the page names it without linking.
Two English authorities are in that position. It was unreachable in
today's data — verified across the EC and TR outcodes — but the thin
place redirect would have sent a reader to a 404 the year it isn't.

The e2e journey now walks every /schools link a page of each family
emits and requires a 200, which is the check that was missing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 17:22:11 +01:00
tudor 865a69b54d Merge pull request 'feat(places): list schools alphabetically on place pages' (#123) from feat/place-alphabetical-sort into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m27s
Reviewed-on: #123
2026-08-21 23:22:49 +00:00
TudorandClaude Opus 5 9cc87c41bb fix(places): a phase page needs results, not merely publishable schools
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m4s
Asked where schools with no results should sit in an alphabetical list, and
found that some pages were almost entirely made of them.

The per-phase threshold counted schools that were publishable — a result OR
an Ofsted grade — while a phase page exists for its results column.
/schools/kent/primary published with none of its five rows carrying a result;
Minehead had one of seven, Buntingford one of five. Forty-four phase pages
were majority-blank.

It is the same rule as "no page without a local average", which was written
into the spec as a thin-page control and never extended per phase.

The threshold now counts schools with a result for that phase. It gates
whether the page exists; it does not filter rows — a page that publishes still
lists every school of the phase, because someone looking up a school by name
has to find it whether or not it published results.

126 of 1,012 variant pages stop publishing: 62 primary, 64 secondary. Every
one of them was a table with too little in it to be worth a page.

The ordering itself is unchanged: pure A-Z, blanks interleaved. A school sits
where its name says it does, and at roughly a tenth of rows that reads fine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:14:31 +01:00
TudorandClaude Opus 5 8967966eef feat(places): list schools alphabetically on place pages
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m21s
Someone on a place page is usually looking for a school they can name, so the
order should serve scanning for it rather than ranking. /api/rankings keeps
its league-table ordering; this is a place-page decision, not a site-wide one.
Sorted case-insensitively, or a capitalised name would sort ahead of every
lowercase one.

The change made five pieces of copy untrue, so they go with it. The phase
variant titled itself "— Ranked", and all four route families described
themselves as "ranked by SATs and GCSE results". A page that opens by claiming
an order it does not keep is worse than one that claims nothing.

The ItemList markup carried `position` with no declared order, which reads as
a ranking. It now declares ItemListOrderAscending, so the structured data says
what the table does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:09:58 +01:00
tudor 4a9a5c734b Merge pull request 'fix(e2e): three assertions that were wrong about correct behaviour' (#122) from fix/e2e-canonical-and-robots into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m26s
Reviewed-on: #122
2026-08-21 23:07:57 +00:00
TudorandClaude Opus 5 4e82e6c916 fix(e2e): three assertions that were wrong about correct behaviour
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 41s
The staging gate was red on three journeys. All three were faults in the
tests; the site was behaving correctly in each case.

Next normalises canonical URLs against trailingSlash:false, so the homepage
ships "https://www.schoolcompare.co.uk" with no slash while every other route
keeps its path. Both address the same document. The test hardcoded the slash
and so failed only on the root — /rankings and /admissions passed throughout,
which is what made it look like a homepage bug rather than a test bug.
Compared with trailing slashes stripped from both sides.

The robots.txt assertion matched "Disallow: /" anywhere in the file and
tripped over the AI-crawler groups Cloudflare injects — ClaudeBot, GPTBot,
Amazonbot and six others all carry a blanket disallow, deliberately, and none
of them is Googlebot. It now parses the file into user-agent groups and checks
only the "*" group, which is also the thing the test was always trying to say:
Google may crawl the page, so it can see the noindex header.

Both were the same mistake as the doubled brand: asserting a naive string
rather than the semantics, and asserting against what the code assembles
rather than what the page renders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 23:51:34 +01:00
tudor d4340a8fdd Merge pull request 'feat(places): name every authority a place sits in' (#121) from feat/place-multiple-authorities into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #121
2026-08-21 21:56:50 +00:00
TudorandClaude Opus 5 bb2f7a5841 fix(places): address review, and merge places GIAS spells more than one way
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m34s
Two findings from review on #121, plus a third the review prompted.

The cap at three authorities silently dropped the fourth in exactly the case
where the information matters most — a genuinely fragmented place — and
contradicted the stated goal of naming every authority a place sits in. It is
gone. The share rule was always the real limit and already bounds the list at
ten. Measured against the live corpus, one town would have been truncated
today: LONDON, split evenly between Hackney, Lambeth, Westminster and
Lewisham.

parent_authority used mode() while authorities used value_counts(), and on an
exact tie pandas does not guarantee the two pick the same name, so the 301
could have pointed somewhere other than the authority named first on the page.
The parent is now derived from authorities[0]: one computation, one answer.
It also inherits the sentinel filter, so a place can no longer redirect to
/schools/authority/does-not-apply.

Chasing the truncation case surfaced a worse bug. Places were grouped by raw
town value, but the registry is keyed by slug, and GIAS spells the same place
several ways. Five town slugs come from more than one spelling: "London"
(1,819 schools) and "LONDON" (12) both slugify to `london`, so the later group
simply overwrote the earlier one — /schools/london could have shown twelve
schools, silently, depending on row order. Weston-super-Mare was split 14/19
across two spellings and Newcastle-under-Lyme across three. Grouping is now by
slug, and the display name is the most common spelling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:42:18 +01:00
TudorandClaude Opus 5 1cb5314c53 feat(places): name every authority a place sits in
SW19 is mostly Merton but partly Wandsworth, and the page said only Merton.
The cause was one field doing two jobs: _parent_authority takes the modal
authority, which is right for a 301 target and wrong as a statement about
where a place is.

This is not a corner case. A quarter of viable outcodes (425 of 1,760) and a
third of viable towns (263 of 783) cross an authority boundary — Bedford the
town spans Bedford and Central Bedfordshire.

Place now carries `authorities`, every authority holding at least a tenth of
the schools and at least two of them, largest first. parent_authority stays
single and unchanged, because a redirect still needs one target.

The share threshold exists because GIAS carries postcode errors: EN6 lists two
Shropshire schools among fourteen in Hertfordshire, and a bare "any authority
present" rule would print those as though they were real. A place too small or
too fragmented to clear the threshold still names its largest, so the page
never goes silent about where it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:40:58 +01:00
tudor 4cea26b813 Merge pull request 'fix(places): align the measure column's heading with its values' (#120) from fix/place-table-alignment into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #120
2026-08-21 21:21:03 +00:00
TudorandClaude Opus 5 dbb74d9b60 fix(places): align the measure column's heading with its values
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 14s
The heading sat on the right edge of the column and every value on the left.
A specificity collision, not a layout problem: the two were aligned by
different selectors and only one of them won.

  .table td            (0,1,1)  text-align: left    <- won for the value
  .num                 (0,1,0)  text-align: right   <- lost
  .table th:last-child (0,2,1)  text-align: right   <- won for the heading

The heading and the value cell now share one class and one rule, so they
cannot drift apart again whatever else changes around them.

The column also stretched to half the table. It now hugs its content with
width:1% and nowrap, so the school name takes the remaining width — which is
what made the gap read as misalignment on a wide screen, and what crowded the
name column on a narrow one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:15:02 +01:00
tudor 9545aec7f4 Merge pull request 'fix(places): phase-grouped tables, plain-English measures, styled links' (#119) from fix/place-presentation-to-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 56s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #119
2026-08-21 20:57:13 +00:00
Tudor 3365ebcb3a fix(places): phase-grouped tables, plain-English measures, styled links
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Three presentation faults on the place pages, all found by looking at a
rendered page rather than at a test.

An unphased place page showed one primary-only measure for a list holding both
phases: 8 of 27 rows on /schools/brentwood were blank, because secondaries
have no reading-writing-maths score. Picking the other measure would only have
inverted which rows were empty, and putting both in one column would have
mixed a percentage with a 0-90 score. Each phase now gets its own table, so a
blank cell means the school genuinely has no published result — which is worth
saying, and now says "Not published" rather than a bare dash.

"RWM expected" was invented here. The site already names the measure in
METRIC_DEFINITIONS, surfaced at /api/metrics: "Reading, Writing & Maths
Combined %". The heading now reads "Reading, writing & maths" with the full
definition in the tooltip.

Links carried no class at all, so they rendered as default blue underlined
browser links beside a site that styles table links as body colour with a
brand hover. They now follow RankingsView's convention, and running-copy links
take the brand colour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 21:56:06 +01:00
tudor 6d79bd3331 Merge pull request 'fix(places): stop the place titles doubling the brand' (#117) from fix/place-title-brand-doubling into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #117
2026-08-21 20:53:14 +00:00
TudorandClaude Opus 5 24e114dee7 fix(places): stop the place titles doubling the brand
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 27s
Every place page shipped as 'Schools in Brentwood - Compare 27 Schools |
schoolcompare | schoolcompare'. The root layout's title template appends
'| schoolcompare' to any plain-string title, and all four place routes already
carried the brand. W8 opted the other routes out with an absolute title; the
place routes were written afterwards and did not inherit the lesson.

~2,600 titles affected, and the repetition pushed them past Google's
truncation point, so the doubled brand displaced real words in the result.

An e2e journey now asserts no title repeats the brand, across the static
routes and a place page, so this cannot come back on a route added later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 21:43:32 +01:00
tudor 6c5db0c266 Merge pull request 'fix(places): submit and link the phase variants' (#116) from fix/place-phase-variants into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m29s
Reviewed-on: #116
2026-08-21 19:46:55 +00:00
TudorandClaude Opus 5 6f749ed21f fix(places): submit and link the phase variants
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m39s
/schools/[place]/[phase] shipped as routes but reached nothing. The sitemap
emitted one URL per registry entry and the registry had no phase dimension, so
~950 pages were absent from every sitemap — and PlaceView did not link them
either, leaving them reachable by nothing at all.

That is the query shape the baseline actually showed: 'primary schools in
beccles', 'secondary schools in brentwood'. Publishing the routes without a
path in meant building for the demand and then hiding from it.

Place now carries phase_urns so the per-phase threshold can be applied without
re-querying, the sitemap emits a variant wherever a phase clears the threshold
on its own, and the API exposes the qualifying phases so the place page links
only variants that exist. Outcodes are excluded: nobody searches 'primary
schools in SW11' and those routes do not exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 20:43:07 +01:00
tudor d423826840 Merge pull request 'fix(places): a locality collision must not break the sitemap' (#115) from fix/locality-collision-skip into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m38s
Reviewed-on: #115
2026-08-21 19:26:02 +00:00
TudorandClaude Opus 5 d3c63ccc6d fix(places): a locality collision must not break the sitemap
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 34s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 34s
Sitemap regeneration failed on staging. 'richmond' in the curated locality
list collides with the GIAS town Richmond in North Yorkshire (37 schools), the
registry raised, and the admin endpoint 500d — taking down sitemap generation
for all 25,000 school pages over one bad row of curated data.

The guard now skips the colliding locality and logs an error. Skipping still
achieves what the guard was for — a locality never silently shadows a town —
without letting curated data break the site. That matters beyond this bug:
GIAS town names change with no code change here, so a raise could fire
spontaneously in production later.

Also removes four localities that were London boroughs rather than districts.
Hackney, Islington, Greenwich and Ealing are local authorities with 104, 72,
108 and 115 schools and already have authority pages; a locality defined by
two or three outcodes would have been a partial near-duplicate of one — the
thin-content failure the two-namespace design exists to avoid. A test now
guards the whole borough list.

Validated against the live corpus: 15 localities, no town collisions, no
authority duplicates, all 15 clear the threshold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 20:20:52 +01:00
tudor b93eb3a691 Merge pull request 'feat(seo): the location layer — town, locality, authority and outcode pages (W2)' (#114) from feat/w2-location-layer into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m18s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 4s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 11m8s
Reviewed-on: #114
2026-08-21 17:56:05 +00:00
TudorandClaude Opus 5 6b871ce1e9 feat(places): ItemList and BreadcrumbList, and the e2e gate
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 4m33s
ItemList tells Google the page is a ranked set rather than prose;
BreadcrumbList puts the place in a hierarchy. School URLs in the markup are
absolute on the canonical host, since a relative URL in JSON-LD is ambiguous.

Eight journeys covering all four families, the two-namespace guarantee, the
threshold, the canonical, the sitemap and the local-versus-England line — the
last because that comparison is the reason these pages are not a name dropped
into a template.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:19:08 +01:00
TudorandClaude Opus 5 c981d89137 feat(places): town, locality, authority and outcode routes
Every generateStaticParams is gated behind PRERENDER_PLACES and wrapped in the
same try/catch the school route uses. The plan claimed authority pages were
'few enough to always prebuild' — but few enough still means the API must be
reachable at build time, and in CI it is not: the build failed with
ECONNREFUSED rather than degrading to ISR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:17:57 +01:00
TudorandClaude Opus 5 de5e790112 feat(places): place page client and view component
One component for all four families: they differ in what fills the registry,
not in what the page shows, so a second would be a second place to forget the
same change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:15:10 +01:00
TudorandClaude Opus 5 42138fc402 feat(places): submit place and outcode sitemaps
Separate children per family so Search Console reports the location layer's
indexation apart from the school pages' — which is the point of the index
built in W1, and the number the stop condition watches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:13:55 +01:00
TudorandClaude Opus 5 c5af476213 feat(places): /api/places registry and place detail endpoints
The registry is cached for the process and reset by the same admin endpoint
that rebuilds the sitemaps, so places and sitemap always describe the same
corpus rather than drifting apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:13:05 +01:00
TudorandClaude Opus 5 de853b90b3 feat(places): London localities and postcode districts
The GIAS town field puts 1,819 London schools under the single value
'London', so it cannot answer 'schools in Battersea' — a query that appears in
the baseline. No single field can: parliamentary constituency gives Battersea
but not Canary Wharf, admin_ward gives Canary Wharf but not Battersea, and
neither gives Clapham or Shoreditch. So a locality is curated, defined by the
postcode districts it covers, which needs no new ingestion.

A locality may not shadow a published town: the registry raises rather than
silently costing a page that carries real demand. One below the threshold is
logged rather than raising, because a locality can legitimately be too small.

The pipeline seed mirrors the module, with a test guarding the drift — the
same arrangement gias_codes has, and for the same reason: the backend image
does not contain pipeline/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:11:56 +01:00
TudorandClaude Opus 5 759d9f5cea feat(places): registry of towns and authorities
Two namespaces because 67 town names collide with an authority name and
neither set contains the other — postal towns cross authority boundaries, so
Bedford the town holds 104 schools against the authority's 86.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:10:37 +01:00
TudorandClaude Opus 5 555d3f0a7d docs(seo): implementation plan for the W2 location layer
Seven tasks: the place registry, London localities and outcodes, the places
API, per-family sitemaps, the shared place view, the four route families, and
structured data plus the e2e gate.

Two things the plan corrects against the spec. The backend image does not
contain pipeline/, so the curated locality list cannot live only in a dbt
seed — it follows the gias_codes.py precedent instead, canonical in backend
with the seed as a mirror. And NationalAverages is nested by phase rather than
flat, which the first draft read wrongly and would have rendered every page
without its England comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:09:35 +01:00
TudorandClaude Opus 5 ecc847091c docs(seo): design for the W2 location layer
Supersedes the original spec's W2. The Search Console baseline inverted its
ordering: every measured location query is town or district level, none is an
administrative area, and phase is part of the query rather than a filter.

Two problems the original design did not anticipate. 67 viable towns share a
name with a local authority, and the authority is the larger set in only 43 of
them — postal towns cross authority boundaries, so neither can absorb the
other. Two namespaces resolve it by construction. And the GIAS town field
collapses 1,819 London schools into one value, which a curated
locality-to-outcode seed solves without new ingestion.

Sizing is measured against the live 25,185-school corpus rather than
estimated: 783 viable towns, 1,760 outcodes, 154 authorities.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:09:35 +01:00
tudor b187a478c9 Merge pull request 'feat(seo): rewrite the C1 snippets to earn the click (W8)' (#113) from feat/seo-metadata-c1 into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #113
2026-08-20 23:26:12 +00:00
TudorandClaude Opus 5 c0547c45e5 feat(seo): rewrite the C1 snippets to earn the click (W8)
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 11s
The baseline says these pages already rank and are not clicked. 'compare
school performance' sits at position 6.1 with 0.43% CTR; 'compare schools' at
7.2 with 0.87%. The brand query 'school compare' draws 9.16% from the same
neighbourhood of the same results page, which rules out a ranking explanation
— when the snippet gives a reason to click, it gets clicked.

These SERPs are owned by the DfE's own 'Compare school performance' service.
The old title put a lowercase brand nobody searches for in the most valuable
pixels, then a near-paraphrase of that service's name. Beside the government's
own result it read as a lookalike.

Intent in the title, differentiator in the description. Titles now match what
people type, and the descriptions carry the one fact gov.uk does not publish:
how close you had to live to get a place.

/compare deliberately takes the tool phrasing rather than the homepage's, so
the two pages stop competing for one phrase. The root layout's default and
Open Graph copy were saying something different again; they now agree.

No hard school counts in any of it. The corpus moves with every data refresh
and this repo has already shipped one copy bug of that kind.

Tests guard the mechanics — SERP length, intent keyword, the differentiator,
no brand-first title — and leave the wording free to iterate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:21:08 +01:00
tudor 4a3928df9f Merge pull request 'fix(seo): a school is publishable on any year's results, not the latest' (#112) from fix/sitemap-any-year-data into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #112
2026-08-20 23:15:19 +00:00
TudorandClaude Opus 5 07c97a46c5 fix(seo): a school is publishable on any year's results, not the latest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 53s
_school_sitemap_rows tested only the latest year's row, which quietly dropped
every school with results in its history but a null row for the most recent
year — a school that stopped reporting, or whose figures were suppressed for
small-cohort disclosure.

The Mallard Academy (150367) is the case that caught it: real KS2 results for
2015-16 through 2018-19, then null rows from 2022-23 on. Its detail page shows
all four years; the sitemap omitted it. Sampling 40 of the 2,206 excluded
schools found 4 like this, so roughly 220 real pages were being withheld.

Publishable is now a property of the school, computed across every row, while
lastmod still comes from the latest row so the most recent Ofsted date wins.
The field list is a module constant shared with _has_publishable_data so the
two checks cannot drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:04:32 +01:00
tudor bb81337aba Merge pull request 'fix(seo): keep staging out of the search index' (#111) from fix/staging-noindex into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m1s
Reviewed-on: #111
2026-08-20 22:39:38 +00:00
TudorandClaude Opus 5 b34511e459 chore: record the branch cleanup manifest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m30s
79 remote branches deleted: 77 fully merged into main, plus
feat/seo-crawl-hygiene and feat/england-only-corpus, whose content is
preserved on feat/seo-crawl-hygiene-main (PR #110).

Each line carries the SHA, so any branch can be restored with
  git push origin <sha>:refs/heads/<name>

The 14 branches left standing all carry content that differs from main and
none of them is mine to judge abandoned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:23:30 +01:00
tudor f928a15c1e Merge pull request 'fix(seo): crawl hygiene and a per-family sitemap index (W1)' (#110) from feat/seo-crawl-hygiene-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m21s
Reviewed-on: #110
2026-08-20 22:23:10 +00:00
TudorandClaude Opus 5 1fc1e07d21 fix(seo): keep staging out of the search index
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 13s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
Staging serves the same image as production off stx., with robots.txt saying
Allow: / and no noindex — a fully crawlable duplicate of the site. Nothing
appears indexed today, most likely because the pages canonicalise across to
production, but that is a side effect rather than a control.

X-Robots-Tag, not a robots.txt Disallow. Disallow blocks crawling, which is
not the same as blocking indexing: a disallowed URL can still be indexed from
external links, and blocking the crawl means Google never fetches the page and
so never sees a noindex at all. Staging stays crawlable and answers noindex.

Matched on the staging host explicitly rather than 'any host that is not
production'. The inverted form would cover future environments automatically,
but its failure mode is deindexing production if the Host header ever arrives
rewritten by a proxy — which cannot be verified from here. This form's failure
mode is a new environment being indexable until someone adds it, which is
recoverable. Any new non-production hostname must be added.

The journeys only ever run against staging (deploy.yml passes
STAGING_BASE_URL; promote.yml smoke-polls production without Playwright), so
asserting the header there is safe. The two assertions live in one test
because the halves only work together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:21:58 +01:00
64 changed files with 10163 additions and 36 deletions

No files matched your search

+330 -7
View File
@@ -6,6 +6,7 @@ Uses real data from UK Government Compare School Performance downloads.
import hashlib
import re
import time
from contextlib import asynccontextmanager
from datetime import datetime, timezone
from typing import Optional
@@ -15,7 +16,7 @@ import pandas as pd
from fastapi import FastAPI, HTTPException, Query, Request, Depends, Header
from fastapi.middleware.cors import CORSMiddleware
from fastapi.middleware.gzip import GZipMiddleware
from fastapi.responses import FileResponse, Response
from fastapi.responses import FileResponse, JSONResponse, Response
from fastapi.staticfiles import StaticFiles
from slowapi import Limiter, _rate_limit_exceeded_handler
from slowapi.util import get_remote_address
@@ -33,8 +34,11 @@ from .data_loader import (
get_supplementary_data,
get_supplementary_data_batch,
search_schools_typesense,
suggest_schools_typesense,
)
from .data_loader import get_data_info as get_db_info
from . import flags
from .places import build_place_registry
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .utils import clean_for_json, convert_to_native
@@ -58,6 +62,12 @@ MAX_SLUG_LENGTH = 60
# regenerate endpoint after a pipeline run.
_sitemaps: dict[str, str] | None = None
# Built from the same DataFrame the sitemap uses, so places and sitemap can
# never describe different corpora. Reset by the same admin endpoint.
_place_registry: dict | None = None
VALID_PLACE_KINDS = ("town", "locality", "authority", "outcode")
def _slugify(text: str) -> str:
text = text.lower()
@@ -79,6 +89,12 @@ def _school_url(urn: int, school_name: str) -> str:
STATIC_SITEMAP_PATHS = ("/", "/rankings", "/compare", "/admissions")
# A page has something a search result could state if any of these is present
# in any year. Shared by _has_publishable_data and the per-school check in
# _school_sitemap_rows so the two can never drift.
_PUBLISHABLE_FIELDS = ("rwm_expected_pct", "attainment_8_score", "ofsted_grade")
def _has_publishable_data(row) -> bool:
"""True when a school page has something a search result could state.
@@ -87,7 +103,7 @@ def _has_publishable_data(row) -> bool:
signal down, so it stays out of the sitemap. The page itself still resolves
for anyone who has the URL.
"""
for field in ("rwm_expected_pct", "attainment_8_score", "ofsted_grade"):
for field in _PUBLISHABLE_FIELDS:
value = row.get(field)
if value is not None and not pd.isna(value):
return True
@@ -115,6 +131,21 @@ def _school_sitemap_rows(df) -> list[str]:
rows: list[str] = []
seen: set[int] = set()
# Publishable is a property of the SCHOOL, not of its latest row.
#
# The first cut tested the latest year's row alone, which quietly dropped
# every school that has results in its history but a null row for the most
# recent year — a school that stopped reporting, or whose figures were
# suppressed for small-cohort disclosure. The Mallard Academy (150367) is
# the case that caught it: real KS2 results for 2015-16 through 2018-19,
# then null rows for 2022-23 onward. Its page shows all four years; the
# sitemap omitted it. Roughly 220 schools were affected.
publishable_cols = [c for c in _PUBLISHABLE_FIELDS if c in df.columns]
publishable: set[int] = (
set(df.loc[df[publishable_cols].notna().any(axis=1), "urn"].astype(int))
if publishable_cols else set()
)
# Latest row per URN first, so a school's most recent Ofsted date wins.
ordered = df.sort_values("year", ascending=False) if "year" in df.columns else df
@@ -123,7 +154,7 @@ def _school_sitemap_rows(df) -> list[str]:
if urn in seen:
continue
seen.add(urn)
if not _has_publishable_data(row):
if urn not in publishable:
continue
lastmod = None
@@ -149,6 +180,14 @@ SITEMAP_CHUNK_SIZE = 10_000
SITEMAP_CHILD_PREFIX = "/sitemaps"
def get_place_registry() -> dict:
"""The place registry, built once and cached for the process."""
global _place_registry
if _place_registry is None:
_place_registry = build_place_registry(load_school_data())
return _place_registry
def _urlset(rows: list[str]) -> str:
return "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
@@ -158,6 +197,44 @@ def _urlset(rows: list[str]) -> str:
])
def _place_url(place) -> str:
"""The canonical path for a place. Two namespaces, per the spec.
Towns and localities share /schools/[place]; authorities take their own
prefix because 67 town names collide with an authority name and neither
set contains the other.
"""
if place.kind == "authority":
return f"/schools/authority/{place.slug}"
if place.kind == "outcode":
return f"/schools/near/{place.slug}"
return f"/schools/{place.slug}"
def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
Phase is part of the query — "primary schools in beccles" — so each
variant is its own indexable page. Submitting only the bare place URL left
~950 of them reachable by nothing: absent from every sitemap, and not
linked from the place page either.
"""
rows: list[str] = []
for p in sorted(get_place_registry().values(), key=lambda p: (p.kind, p.slug)):
if p.kind not in kinds:
continue
rows.append(_url_element(BASE_URL + _place_url(p)))
# Which phases a place publishes is the registry's decision alone —
# outcodes report none, because the spec gives them no phase route.
# Repeating that rule here was how the page and the sitemap came to
# disagree about which URLs exist.
for phase in ("primary", "secondary"):
if p.publishes_phase(phase):
rows.append(_url_element(f"{BASE_URL}{_place_url(p)}/{phase}"))
return rows
def build_sitemaps() -> dict[str, str]:
"""Build the sitemap index and every child, keyed by name."""
df = load_school_data()
@@ -175,6 +252,17 @@ def build_sitemaps() -> dict[str, str]:
for n, chunk in enumerate(chunks, start=1):
children[f"schools-{n}.xml"] = _urlset(chunk)
# Separate children per family: Search Console reports coverage per
# submitted sitemap, which is how the location layer's indexation is
# measured apart from the school pages'.
for label, kinds in (("places", ("town", "locality", "authority")),
("outcodes", ("outcode",))):
rows = _place_sitemap_rows(kinds)
chunks = [rows[i:i + SITEMAP_CHUNK_SIZE]
for i in range(0, len(rows), SITEMAP_CHUNK_SIZE)] or [[]]
for n, chunk in enumerate(chunks, start=1):
children[f"{label}-{n}.xml"] = _urlset(chunk)
# On a sitemap index, lastmod means "when this sitemap file last changed",
# so generation time is the correct value here — unlike on a <url>, where
# it would be a claim about content we cannot support.
@@ -210,8 +298,101 @@ def clean_filter_values(series: pd.Series) -> list[str]:
# SECURITY MIDDLEWARE & HELPERS
# =============================================================================
# Rate limiter
limiter = Limiter(key_func=get_remote_address)
def client_key(request: Request) -> str:
"""The rate-limit bucket: the real caller, not the proxy in front of them.
`get_remote_address` reads request.client.host. In staging and production
the backend has no published ports and sits on the internal network, so its
only caller is the Next proxy — meaning every browser user on the site
shared one bucket. Measured before this fix: 70 concurrent requests to
/api/schools returned 60 OK and 10 refused.
CF-Connecting-IP first, because Cloudflare (in front of both environments)
sets it on every origin request and *overwrites* any client-supplied value,
which a parsed X-Forwarded-For chain does not guarantee. The XFF fallback is
forgeable, but only by a caller already inside the Docker network, which is
the one place nothing untrusted can reach.
"""
cf = request.headers.get("cf-connecting-ip")
if cf:
return cf.strip()
xff = request.headers.get("x-forwarded-for")
if xff:
return xff.split(",")[0].strip()
return get_remote_address(request)
# Per-client limiter. Paired with the global ceiling below — the two do
# different jobs and neither substitutes for the other.
limiter = Limiter(key_func=client_key)
# --- The ceiling no header can raise ----------------------------------------
#
# client_key trusts CF-Connecting-IP, and nothing in this process can tell an
# edge-set header from an attacker-set one. That distinction can only be made
# at Cloudflare, with Authenticated Origin Pulls or an origin firewall. A
# caller reaching the origin directly could otherwise mint a fresh rate-limit
# bucket per request and evade per-client limits entirely — which would make
# correct keying a net regression against abuse, since the single shared bucket
# it replaced at least capped everyone at 60/minute together.
#
# So per-client limits give fairness, and this gives the origin a hard total.
# It does not make the header trustworthy; it bounds what trusting it can cost.
# The header problem itself is closed at Cloudflare, not here.
#
# [window_start_monotonic, count], or None before the first request. A fixed
# window is crude, which is right for a backstop: it has to be obviously
# correct rather than fair.
_global_window: Optional[list] = None
# The container healthcheck runs `curl http://localhost:80/api/data-info` from
# inside the container. Starving it would fail the check, restart the
# container, and turn a load spike into an outage loop — the ceiling exists to
# protect the origin, not to kill it.
_LOCAL_HOSTS = frozenset({"127.0.0.1", "::1", "localhost"})
def exempt_from_ceiling(request: Request) -> bool:
"""Whether the ceiling should ignore this request.
Its own function so the rule is testable without standing up a server —
and so the healthcheck exemption is somewhere a reader can find it.
"""
if not request.url.path.startswith("/api/"):
return True
# The peer address, never the Host header: Host is set by the caller and
# would hand every attacker an exemption.
return (request.client.host if request.client else "") in _LOCAL_HOSTS
class GlobalRateLimitMiddleware(BaseHTTPMiddleware):
"""A cap on total /api/ traffic, independent of any client identity."""
async def dispatch(self, request: Request, call_next):
global _global_window
if exempt_from_ceiling(request):
return await call_next(request)
now = time.monotonic()
# One event loop, and no await between the read and the write, so this
# sequence is atomic without a lock.
if _global_window is None or now - _global_window[0] >= 60:
_global_window = [now, 0]
_global_window[1] += 1
if _global_window[1] > settings.global_rate_limit_per_minute:
return JSONResponse(
# Distinguishable from slowapi's per-client 429: an operator
# reading logs has to be able to tell "one noisy client" from
# "the origin is saturated".
{"detail": "The service is at capacity. Please retry shortly."},
status_code=429,
headers={"Retry-After":
str(max(1, int(60 - (now - _global_window[0]))))},
)
return await call_next(request)
class SecurityHeadersMiddleware(BaseHTTPMiddleware):
@@ -269,6 +450,7 @@ CACHE_RULES: list[tuple[str, tuple[int, int, int]]] = [
("/api/schools/", (300, 3600, 86400)), # /api/schools/{urn}
("/api/rankings", (60, 600, 3600)),
("/api/compare", (60, 600, 3600)),
("/api/suggest", (60, 3600, 86400)), # autosuggest
("/api/schools", (30, 300, 1800)), # search list
]
@@ -374,6 +556,7 @@ def validate_postcode(postcode: Optional[str]) -> Optional[str]:
async def lifespan(app: FastAPI):
"""Application lifespan - startup and shutdown events."""
global _sitemaps
flags.init()
print("Loading school data from marts...")
df = load_school_data()
if df.empty:
@@ -415,6 +598,10 @@ app.add_middleware(CacheAndETagMiddleware)
app.add_middleware(SecurityHeadersMiddleware)
app.add_middleware(RequestSizeLimitMiddleware)
app.add_middleware(GZipMiddleware, minimum_size=512)
# Added last, so it is outermost and refuses before anything downstream does
# work. A ceiling that only applies after the expensive part has run is not a
# ceiling.
app.add_middleware(GlobalRateLimitMiddleware)
# CORS middleware - restricted for production
app.add_middleware(
@@ -721,7 +908,13 @@ async def get_school_details(request: Request, urn: int):
"census": supplementary.get("census"),
"admissions": supplementary.get("admissions"),
"admissions_history": supplementary.get("admissions_history") or [],
"admission_distance": supplementary.get("admission_distance"),
# Behind a flag, and withheld at the source rather than rendered-but-
# hidden: this endpoint is public and unauthenticated, so a field left
# in the payload is a published field. The key is absent, not null —
# null would state that this school has no cut-off, which is a
# different claim from "we are not publishing cut-offs".
**({"admission_distance": supplementary.get("admission_distance")}
if flags.is_enabled("admission_distance") else {}),
"sen_detail": supplementary.get("sen_detail"),
"phonics": supplementary.get("phonics"),
"deprivation": supplementary.get("deprivation"),
@@ -1096,6 +1289,132 @@ async def get_rankings(
}
@app.get("/api/places")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def list_places(request: Request):
"""Every published place. The sitemap and the link modules read this."""
registry = get_place_registry()
return {"places": [
{"kind": p.kind, "slug": p.slug, "name": p.name, "count": len(p.urns)}
for p in sorted(registry.values(), key=lambda p: (p.kind, p.slug))
]}
@app.get("/api/places/{kind}/{slug}")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_place(request: Request, kind: str, slug: str,
phase: Optional[str] = None):
"""One place: its schools ranked, and its local averages."""
if kind not in VALID_PLACE_KINDS:
raise HTTPException(status_code=404, detail="No such place")
registry = get_place_registry()
place = registry.get(f"{kind}:{slug}")
if place is None:
raise HTTPException(status_code=404, detail="No such place")
df = load_latest_school_data()
rows = df[df["urn"].isin(place.urns)]
if phase:
wanted = PHASE_GROUPS.get(phase.lower())
if wanted and "phase" in rows.columns:
rows = rows[rows["phase"].fillna("").str.lower().isin(wanted)]
# The metric the page shows, and averages.
metric = "attainment_8_score" if phase == "secondary" else "rwm_expected_pct"
# Alphabetical, not by score. A place page is read by someone looking for
# a school they can name, and scanning for it is what the order should
# serve. /rankings is where the league-table ordering lives, and it keeps
# sorting by metric.
if "school_name" in rows.columns:
rows = rows.sort_values("school_name", key=lambda c: c.str.lower())
averages = {
m: (None if m not in rows.columns or rows[m].dropna().empty
else float(rows[m].dropna().mean()))
for m in ("rwm_expected_pct", "attainment_8_score")
}
cols = [c for c in SCHOOL_COLUMNS + ["latitude", "longitude", "phase",
"rwm_expected_pct", "attainment_8_score",
"total_pupils"]
if c in rows.columns]
return {
"place": {"kind": place.kind, "slug": place.slug, "name": place.name,
"count": len(place.urns),
"parent_authority": place.parent_authority,
# Every authority the place meaningfully sits in. SW19 is
# mostly Merton but partly Wandsworth; naming one asserts
# something false.
#
# The slug is null where that authority has no page of its
# own: City of London and the Isles of Scilly hold fewer
# schools than the threshold. Naming them is still right;
# linking them would be a 404.
"authorities": [
{"name": name,
"slug": (_slugify(name)
if f"authority:{_slugify(name)}" in registry
else None),
"count": n}
for name, n in place.authorities
],
# Only phases that clear the threshold, so the page links
# variants that exist rather than 404s.
"phases": [ph for ph in ("primary", "secondary")
if place.publishes_phase(ph)]},
"schools": clean_for_json(rows[cols]),
"averages": averages,
}
# Two characters. One is not a query — it matches thousands of schools and the
# response is useless, so it is not worth a round trip.
SUGGEST_MIN_QUERY = 2
@app.get("/api/suggest")
@limiter.limit("120/minute")
async def suggest_schools(
request: Request,
q: str = Query("", max_length=100),
limit: int = Query(8, ge=1, le=20),
):
"""School name suggestions, from Typesense alone.
Deliberately not a mode of /api/schools: that path filters and sorts the
full in-memory DataFrame, which is far too expensive to run per keystroke.
Nothing here returns an error for ordinary input. A short query, no
matches, or Typesense being unreachable are all 200 with an empty list —
a dropdown that quietly does not appear is the right failure for a
keystroke path, and there is no DataFrame fallback because the 25,000-row
substring scan is precisely what this endpoint exists to avoid.
120/minute rather than the default 60: a 200 ms debounce makes typing
legitimately bursty.
"""
query = q.strip()
if len(query) < SUGGEST_MIN_QUERY:
return {"suggestions": []}
return {"suggestions": suggest_schools_typesense(query, limit)}
@app.get("/api/flags")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_feature_flags(request: Request):
"""Every declared flag and its current value.
Internal only. The Next proxy denies this path, because the response names
every unreleased feature the codebase knows about — which is exactly what
shipping dark is meant to keep quiet.
"""
return flags.all_flags()
@app.get("/api/data-info")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_data_info(request: Request):
@@ -1207,7 +1526,11 @@ async def regenerate_sitemap(
_: bool = Depends(verify_admin_api_key),
):
"""Rebuild and cache the sitemap from current school data. Called by Airflow after data updates."""
global _sitemaps
global _sitemaps, _place_registry
# Places and sitemap are rebuilt together — they read the same marts, and
# letting them drift apart would submit URLs for places that no longer
# exist.
_place_registry = None
_sitemaps = build_sitemaps()
n = sum(x.count("<url>") for x in _sitemaps.values())
return {"status": "ok", "urls": n, "sitemaps": len(_sitemaps)}
+15
View File
@@ -35,6 +35,11 @@ class Settings(BaseSettings):
# Security
admin_api_key: str = Field(default_factory=lambda: secrets.token_urlsafe(32))
rate_limit_per_minute: int = 60 # Requests per minute per IP
# A ceiling on total /api/ traffic, independent of any client identity.
# client_key trusts headers only Cloudflare can vouch for, so a caller
# reaching the origin directly could otherwise mint a fresh bucket per
# request. See GlobalRateLimitMiddleware in backend/app.py.
global_rate_limit_per_minute: int = 3000
rate_limit_burst: int = 10 # Allow burst of requests
max_request_size: int = 1024 * 1024 # 1MB max request size
@@ -42,6 +47,16 @@ class Settings(BaseSettings):
typesense_url: str = "http://localhost:8108"
typesense_api_key: str = ""
# Feature flags (Unleash). An empty unleash_url disables flags entirely and
# every flag evaluates False — the correct behaviour for local development
# and CI, and the reason no test needs a running Unleash.
unleash_url: str = ""
unleash_api_token: str = ""
unleash_app_name: str = "schoolcompare-backend"
# On a named volume, so a restart during an Unleash outage keeps
# last-known state instead of reverting a released feature to dark.
unleash_cache_directory: str = "/app/.unleash"
# Analytics
ga_measurement_id: Optional[str] = "G-J0PCVT14NY" # Google Analytics 4 Measurement ID
+52
View File
@@ -100,6 +100,58 @@ def search_schools_typesense(query: str, limit: int = 250) -> List[int]:
return []
# The most a public endpoint will return in one response.
SUGGEST_MAX_LIMIT = 20
# Fields a suggestion row carries, and the default when the document omits an
# optional one. phase and school_type are optional in the Typesense schema.
_SUGGEST_FIELDS = ("school_name", "local_authority", "postcode",
"phase", "school_type")
def suggest_schools_typesense(query: str, limit: int = 8) -> List[dict]:
"""Autosuggest rows straight from Typesense. Never raises.
Returns documents rather than URNs, unlike search_schools_typesense, so the
caller needs no DataFrame. Every field below is already in the index — see
pipeline/scripts/sync_typesense.py — which is what makes this cheap enough
to run per keystroke.
"""
client = _get_typesense_client()
if client is None:
return []
try:
result = client.collections["schools"].documents.search({
"q": query,
"query_by": "school_name,local_authority",
"per_page": max(1, min(limit, SUGGEST_MAX_LIMIT)),
"typo_tokens_threshold": 1,
})
except Exception:
# A dropdown that quietly stops appearing is the right failure here.
return []
rows = []
for hit in result.get("hits", []) or []:
doc = (hit or {}).get("document") or {}
try:
urn = int(doc["urn"])
except (KeyError, TypeError, ValueError):
# Skip the row, keep the rest. Typesense declares urn as int32 so
# this should be unreachable, but the index is a separate system
# that something other than this code can reindex — and "never
# raises" is a promise the keystroke path actually depends on.
# Dropping one malformed document is right; blanking the whole
# dropdown, or serving a suggestion pointing at /school/0, is not.
logging.getLogger(__name__).warning(
"skipping malformed suggestion document: %r", doc)
continue
row = {"urn": urn}
row.update({f: str(doc.get(f, "") or "") for f in _SUGGEST_FIELDS})
rows.append(row)
return rows
def normalize_school_type(school_type: Optional[str]) -> Optional[str]:
"""Convert cryptic school type codes to user-friendly names."""
if not school_type:
+118
View File
@@ -0,0 +1,118 @@
"""Feature flags: what can be switched, and what is switched right now.
Ship-dark, not a kill switch. Flags let work merge and deploy without becoming
visible; they are expected to flip about monthly, by a person, deliberately.
Nothing here does percentage rollouts or user targeting — the site has no user
identity to target.
Unleash holds the state. It does not hold the list. REGISTRY below is that
list, and it exists for three reasons: the SDK evaluates an unknown flag to
False, so without a registry that is an *undeclared* False, indistinguishable
from a typo; /api/flags needs a key set to return when Unleash is unreachable;
and a flag in the UI but not in the registry is orphaned and should be visibly
so rather than quietly authoritative.
Every flag defaults to False. There is no per-flag default, because a flag that
defaults on is a kill switch, and this is not one.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from datetime import date
from .config import settings
logger = logging.getLogger(__name__)
# A flag is temporary scaffolding. See test_a_flag_older_than_the_limit.
MAX_FLAG_AGE_DAYS = 90
@dataclass(frozen=True)
class Flag:
# One string: the registry key, the Unleash flag name, and the JSON key in
# /api/flags. snake_case, matching the API's existing convention. No case
# transformation anywhere, so there is no mapping layer to get wrong.
name: str
description: str # one line: what turning this on reveals
added: date # for the staleness tripwire
REGISTRY: dict[str, Flag] = {
f.name: f for f in (
Flag(
name="admission_distance",
description=(
"The last-distance-offered figure on the Admissions tile and "
"the 'How far away are you?' section on school pages."
),
added=date(2026, 8, 23),
),
Flag(
name="school_autosuggest",
description=(
"School name suggestions as you type in the main search box."
),
added=date(2026, 8, 26),
),
)
}
_client = None
def init() -> None:
"""Start the Unleash client, or log why flags are all off.
Called once from the app lifespan. Never raises: a flag system that can
stop the API from booting is worse than one that is switched off.
"""
global _client
if not settings.unleash_url or not settings.unleash_api_token:
logger.warning(
"Unleash is not configured (UNLEASH_URL / UNLEASH_API_TOKEN); "
"every feature flag evaluates to False.")
return
try:
from UnleashClient import UnleashClient
_client = UnleashClient(
url=settings.unleash_url,
app_name=settings.unleash_app_name,
custom_headers={"Authorization": settings.unleash_api_token},
cache_directory=settings.unleash_cache_directory,
refresh_interval=15,
)
_client.initialize_client()
logger.info("Unleash client initialised against %s", settings.unleash_url)
except Exception:
# Fail closed and keep serving. The SDK also evaluates everything False
# until its first successful sync, so this is the same direction.
_client = None
logger.exception("Unleash client failed to start; flags are all False.")
def is_enabled(name: str) -> bool:
"""Whether `name` is on. False for anything unknown, unreachable or broken."""
if name not in REGISTRY:
logger.error(
"undeclared feature flag %r was evaluated; returning False. "
"Add it to backend/flags.py REGISTRY or fix the name.", name)
return False
if _client is None:
return False
try:
return bool(_client.is_enabled(
name, fallback_function=lambda feature_name, context: False))
except Exception:
logger.exception("flag %r failed to evaluate; returning False", name)
return False
def all_flags() -> dict[str, bool]:
"""Every declared flag and its current value. Serves /api/flags."""
return {name: is_enabled(name) for name in REGISTRY}
+54
View File
@@ -0,0 +1,54 @@
"""Curated London localities, defined by the postcode districts they cover.
The GIAS `town` field puts 1,819 London schools under the single value
"London", so it cannot answer "schools in Battersea" — a query that appears in
the Search Console baseline. No single field can: parliamentary constituency
gives Battersea but not Canary Wharf; postcodes.io's admin_ward gives Canary
Wharf but not Battersea; neither gives Clapham or Shoreditch, which are postal
and colloquial rather than administrative.
So this is curated. Where a locality ends is a judgement, not a fact, and a
reviewable file is the honest place for a judgement. No new ingestion is
needed — the corpus already carries postcodes.
This is the canonical copy. `pipeline/transform/seeds/locality_outcodes.csv`
mirrors it for anyone querying the warehouse directly; the backend image does
not contain `pipeline/`, which is why the module rather than the seed is
canonical. Same arrangement as `backend/gias_codes.py`.
A locality whose outcodes hold fewer than MIN_SCHOOLS schools is not
published, so a typo produces no page rather than an empty one. Places that
fail that check are logged at startup, because a locality you meant to publish
quietly not appearing is the failure worth hearing about.
Two rules for anything added here.
**Sub-borough districts only.** A London borough is a local authority and
already has a page at /schools/authority/[la] covering all of its schools; a
locality defined by two or three outcodes would be a partial, near-duplicate
subset of it. Hackney, Islington, Greenwich and Ealing were all in the first
draft for that reason and have been removed.
**The slug must not match a GIAS town.** "Richmond" did — GIAS has a Richmond
in North Yorkshire with 37 schools — so the London one could never publish.
The registry skips any locality that collides and logs it.
"""
# slug -> (display name, outcodes)
LOCALITY_OUTCODES: dict[str, tuple[str, tuple[str, ...]]] = {
"battersea": ("Battersea", ("SW11",)),
"canary-wharf": ("Canary Wharf", ("E14",)),
"clapham": ("Clapham", ("SW4",)),
"shoreditch": ("Shoreditch", ("EC2A", "E1")),
"peckham": ("Peckham", ("SE15",)),
"brixton": ("Brixton", ("SW2", "SW9")),
"camden-town": ("Camden Town", ("NW1",)),
"wimbledon": ("Wimbledon", ("SW19",)),
"putney": ("Putney", ("SW15",)),
"fulham": ("Fulham", ("SW6",)),
"chiswick": ("Chiswick", ("W4",)),
"stratford": ("Stratford", ("E15",)),
"walthamstow": ("Walthamstow", ("E17",)),
"tooting": ("Tooting", ("SW17",)),
"dulwich": ("Dulwich", ("SE21", "SE22")),
}
+314
View File
@@ -0,0 +1,314 @@
"""The place registry: what places the site publishes, and what is in each.
One module owns this question. The pages, the sitemap and the internal-link
modules all read from here, so the threshold and the collision rules exist in
exactly one place and are testable without a browser or a database.
Two namespaces, never one. 67 viable town names collide with a local
authority name, and the authority is the larger set in only 43 of them —
postal towns cross authority boundaries, so neither can absorb the other.
Keys are "<kind>:<slug>" so the collision cannot reappear in the dict.
"""
from __future__ import annotations
import logging
import re
from dataclasses import dataclass, field
logger = logging.getLogger(__name__)
# Five schools with publishable data. Below this a place has nothing to say
# that a list of schools does not, and publishing it is index bloat.
MIN_SCHOOLS = 5
@dataclass(frozen=True)
class Place:
kind: str # "town" | "locality" | "authority" | "outcode"
slug: str
name: str
urns: tuple[int, ...]
parent_authority: str | None # authority NAME, for the 301 target
# Every authority the place meaningfully sits in, largest first. A quarter
# of outcodes and a third of towns straddle a boundary — SW19 is mostly
# Merton but partly Wandsworth — so naming only one asserts something
# false. parent_authority stays single because a redirect needs one
# target; this is what the page shows.
authorities: tuple[tuple[str, int], ...] = ()
# URNs per phase, so the per-phase threshold can be applied without
# re-querying. A place with 30 primaries and 2 secondaries publishes a
# primary variant and no secondary one.
phase_urns: dict[str, tuple[int, ...]] = field(default_factory=dict)
def publishes_phase(self, phase: str) -> bool:
return len(self.phase_urns.get(phase, ())) >= MIN_SCHOOLS
@property
def key(self) -> str:
return f"{self.kind}:{self.slug}"
def _publishable_urns(df) -> set[int]:
"""URNs with something a page could state, deduplicated across years."""
from backend.app import _PUBLISHABLE_FIELDS
cols = [c for c in _PUBLISHABLE_FIELDS if c in df.columns]
if not cols:
return set()
return set(df.loc[df[cols].notna().any(axis=1), "urn"].astype(int))
# The measure a phase page is built around. A page with no results in this
# column has nothing a list of school names does not already give.
_PHASE_METRIC = {
"primary": "rwm_expected_pct",
"secondary": "attainment_8_score",
}
def _phase_urns(group, publishable: set[int]) -> dict[str, tuple[int, ...]]:
"""URNs per phase, counting only schools with a result for that phase.
Not merely "publishable". A school with an Ofsted grade and no results is
worth a page of its own and belongs in the place list, but it cannot
populate a phase page's results column — and the threshold is there to ask
whether that column will have anything in it.
Counting publishable schools instead let /schools/kent/primary publish
with none of its five rows carrying a result, and left 44 phase pages
majority-blank. It is the same rule as "no page without a local average",
which was never extended per phase.
All-through schools count toward both phases, matching the PHASE_GROUPS
mapping the search filters already use.
"""
from backend.app import PHASE_GROUPS
if "phase" not in group.columns:
return {}
lowered = group["phase"].fillna("").str.lower()
out: dict[str, tuple[int, ...]] = {}
for phase in ("primary", "secondary"):
wanted = PHASE_GROUPS.get(phase, set())
subset = group[lowered.isin(wanted)]
# The page lists every school of the phase; the threshold counts only
# those carrying a result, so a mostly-empty table never publishes.
metric = _PHASE_METRIC[phase]
with_result = (
{int(u) for u in subset.loc[subset[metric].notna(), "urn"]}
if metric in subset.columns else set()
)
if len(with_result & publishable) < MIN_SCHOOLS:
continue
urns = tuple(sorted({int(u) for u in subset["urn"]} & publishable))
if urns:
out[phase] = urns
return out
# A place is described by an authority when it holds at least a tenth of the
# schools, and at least two. GIAS carries occasional postcode errors — EN6
# lists two Shropshire schools among fourteen in Hertfordshire — and a bare
# "any authority present" rule would print those as though they were real.
# There is deliberately no cap on how many are named. An earlier cut stopped
# at three, which silently dropped the fourth in exactly the case where the
# information matters most — a genuinely fragmented place. The share rule is
# the only limit, and it already bounds the list at ten.
_AUTHORITY_MIN_SHARE = 0.10
_AUTHORITY_MIN_SCHOOLS = 2
def _authorities(group) -> tuple[tuple[str, int], ...]:
"""Authorities this place meaningfully sits in, largest first."""
from backend.app import EXCLUDED_FILTER_VALUES
if "local_authority" not in group.columns:
return ()
counts = group["local_authority"].dropna().value_counts()
total = int(counts.sum())
if not total:
return ()
kept = [
(str(name), int(n)) for name, n in counts.items()
if str(name) not in EXCLUDED_FILTER_VALUES
and n >= _AUTHORITY_MIN_SCHOOLS
and n / total >= _AUTHORITY_MIN_SHARE
]
# A place too small or too fragmented for the share rule still names its
# largest authority, or the page would say nothing about where it is.
if not kept:
for name, n in counts.items():
if str(name) not in EXCLUDED_FILTER_VALUES:
return ((str(name), int(n)),)
return ()
return tuple(kept)
def _parent_authority(authorities: tuple[tuple[str, int], ...]) -> str | None:
"""The 301 target: the largest authority a place sits in.
Derived from `authorities` rather than computed separately. The first cut
used `mode()` here while `authorities` used `value_counts()`, and on an
exact tie pandas does not guarantee the two pick the same name — so the
redirect could have pointed somewhere other than the authority the page
named first. One computation, one answer.
Deriving it also inherits the sentinel filter, so a place can no longer
redirect to /schools/authority/does-not-apply.
"""
return authorities[0][0] if authorities else None
def _group(df, column: str, kind: str, publishable: set[int]) -> dict[str, Place]:
"""One Place per distinct SLUG in `column` that clears the threshold.
Grouped by slug, not by raw value, because GIAS spells the same place
several ways and they all resolve to one URL. Five town slugs come from
more than one spelling: "London" (1,819 schools) and "LONDON" (12) both
slugify to `london`; Weston-super-Mare is split 14/19 across two
spellings; Newcastle-under-Lyme across three.
Grouping by raw value meant the later group simply overwrote the earlier
one in this dict — so /schools/london could have shown twelve schools
instead of 1,819, silently and depending on row order.
The display name is the most common spelling, which is the one a reader
expects to see.
"""
from backend.app import _slugify
if column not in df.columns:
return {}
working = df.assign(_slug=df[column].map(
lambda v: _slugify(str(v).strip()) if isinstance(v, str) and v.strip() else None))
working = working[working["_slug"].notna() & (working["_slug"] != "")]
out: dict[str, Place] = {}
for slug, group in working.groupby("_slug"):
slug = str(slug)
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
continue
spellings = group[column].dropna().value_counts()
if spellings.empty:
continue
name = str(spellings.index[0]).strip()
authorities = () if kind == "authority" else _authorities(group)
place = Place(
kind=kind, slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable),
)
out[place.key] = place
return out
# "SW11 2AA" -> "SW11". Two letters max, one or two digits, optional letter.
_OUTCODE_RE = re.compile(r"^([A-Z]{1,2}\d{1,2}[A-Z]?)\s")
def _outcode(postcode) -> str | None:
if not isinstance(postcode, str):
return None
m = _OUTCODE_RE.match(postcode.upper().strip())
return m.group(1) if m else None
def _outcode_places(df, publishable: set[int]) -> dict[str, Place]:
"""One Place per postcode district clearing the threshold.
These carry no phase variants: nobody searches "primary schools in SW11",
so the spec gives them no /primary or /secondary route. `phase_urns` is
left empty rather than computed and then filtered downstream — the
registry is the one place that decides which phases a place publishes,
and the page links whatever it reports.
Computing them here put a link to a route that does not exist on every one
of the 1,720 outcode pages.
"""
if "postcode" not in df.columns:
return {}
working = df.assign(_oc=df["postcode"].map(_outcode))
working = working[working["_oc"].notna()]
out: dict[str, Place] = {}
for oc, group in working.groupby("_oc"):
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
continue
authorities = _authorities(group)
place = Place(kind="outcode", slug=str(oc).lower(), name=str(oc),
urns=urns, parent_authority=_parent_authority(authorities),
authorities=authorities)
out[place.key] = place
return out
def _locality_places(df, publishable: set[int],
town_slugs: set[str]) -> dict[str, Place]:
"""One Place per curated locality clearing the threshold."""
from backend.localities import LOCALITY_OUTCODES
if "postcode" not in df.columns:
return {}
working = df.assign(_oc=df["postcode"].map(_outcode))
out: dict[str, Place] = {}
for slug, (name, outcodes) in LOCALITY_OUTCODES.items():
if slug in town_slugs:
# Skip, do not raise. The guard exists so a locality never
# silently shadows a town — skipping achieves that, and the error
# log makes it loud.
#
# Raising here took down sitemap generation for all 25,000 school
# pages when "richmond" met the GIAS town Richmond in North
# Yorkshire. Worse, GIAS town names change without any code change,
# so a raise means curated data can break the site spontaneously.
# A curation mistake must cost one page, not the sitemap.
logger.error(
"locality %r collides with the published town of the same "
"slug and has been skipped; rename it or remove it", slug)
continue
group = working[working["_oc"].isin(outcodes)]
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
# Not an error — a locality can legitimately be too small. Logged
# because one you meant to publish quietly vanishing is the
# failure worth hearing about.
logger.warning(
"locality %s (%s) has %d publishable schools, below the "
"threshold of %d - not published",
slug, ", ".join(outcodes), len(urns), MIN_SCHOOLS)
continue
authorities = _authorities(group)
place = Place(kind="locality", slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable))
out[place.key] = place
return out
def build_place_registry(df) -> dict[str, Place]:
"""Every place the site publishes, keyed by "<kind>:<slug>"."""
if df.empty or "urn" not in df.columns:
return {}
publishable = _publishable_urns(df)
registry: dict[str, Place] = {}
registry.update(_group(df, "local_authority", "authority", publishable))
towns = _group(df, "town", "town", publishable)
registry.update(towns)
town_slugs = {p.slug for p in towns.values()}
registry.update(_locality_places(df, publishable, town_slugs))
registry.update(_outcode_places(df, publishable))
return registry
+163
View File
@@ -0,0 +1,163 @@
"""Tests for the feature flag layer (spec 2026-08-23).
None of these need a running Unleash. That is the point: an unset UNLEASH_URL
means every flag is False, which is what local development and CI get.
"""
from datetime import date, timedelta
from backend import flags
def test_every_declared_flag_is_keyed_by_its_own_name():
# One string is the registry key, the Unleash flag name and the JSON key.
# A mismatch here would mean the UI toggles a flag the code never reads.
for key, flag in flags.REGISTRY.items():
assert key == flag.name
def test_flag_names_are_snake_case():
# Matches the API's existing convention (admission_distance,
# rwm_expected_pct) so no case transformation exists to get wrong.
for name in flags.REGISTRY:
assert name == name.lower()
assert "-" not in name and " " not in name
def test_an_unconfigured_client_evaluates_every_flag_false(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
for name in flags.REGISTRY:
assert flags.is_enabled(name) is False
def test_an_undeclared_flag_is_false_rather_than_an_error(monkeypatch):
# A typo'd flag name must not raise in a request path. It is logged as an
# error, because an undeclared flag is always a bug.
monkeypatch.setattr(flags, "_client", None)
assert flags.is_enabled("no_such_flag") is False
def test_an_exploding_client_is_false_rather_than_a_500(monkeypatch):
class Boom:
def is_enabled(self, *a, **kw):
raise RuntimeError("unleash is on fire")
monkeypatch.setattr(flags, "_client", Boom())
name = next(iter(flags.REGISTRY))
assert flags.is_enabled(name) is False
def test_all_flags_reports_every_declared_flag(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
assert set(flags.all_flags()) == set(flags.REGISTRY)
assert all(v is False for v in flags.all_flags().values())
def test_a_flag_older_than_the_limit_fails_this_test():
"""A tripwire, not an assertion about correctness.
Flags are temporary scaffolding and the failure mode of every flag system
is accumulation. This fails on the day a flag turns 90, on whatever PR
happens to be open — which is the point: someone has to decide.
To fix: delete the flag and the branches that read it, or, if it genuinely
still needs to exist, move its `added` date and say why in the commit.
"""
stale = [
f.name for f in flags.REGISTRY.values()
if date.today() - f.added > timedelta(days=flags.MAX_FLAG_AGE_DAYS)
]
assert not stale, (
f"Flags older than {flags.MAX_FLAG_AGE_DAYS} days: {stale}. "
"Remove the flag and the code branches it guards, or move its `added` "
"date deliberately."
)
def _client():
from fastapi.testclient import TestClient
from backend import app as app_module
return TestClient(app_module.app, raise_server_exceptions=False)
def test_the_flags_endpoint_lists_every_declared_flag(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
body = _client().get("/api/flags").json()
assert set(body) == set(flags.REGISTRY)
def test_the_flags_endpoint_answers_false_when_unleash_is_unreachable(monkeypatch):
# The endpoint must still answer. A frontend that cannot read flags renders
# everything dark, which is right; one that gets a 500 renders nothing.
monkeypatch.setattr(flags, "_client", None)
res = _client().get("/api/flags")
assert res.status_code == 200
assert all(v is False for v in res.json().values())
def _school_payload(monkeypatch, *, flag_on: bool):
"""Fetch one school's payload with the distance flag forced on or off.
The DataFrame shape is copied from test_school_details.py rather than
minimised: the endpoint reads a wide set of GIAS columns, and a trimmed
frame fails for reasons that have nothing to do with flags.
"""
import numpy as np
import pandas as pd
from fastapi.testclient import TestClient
from backend import app as app_module
df = pd.DataFrame([{
"urn": 150275,
"school_name": "West London Performing Arts Academy",
"phase": "Secondary",
"school_type": "Special post 16 institution",
"trust_name": None,
"religious_denomination": "Does not apply",
"gender": None,
"age_range": "16-25",
"admissions_policy": None,
"capacity": np.nan,
"gias_total_pupils": np.nan,
"headteacher_name": None,
"website": None,
"ofsted_grade": np.nan,
"local_authority": "Ealing",
"address": "268 Northfield Avenue, London, W5 4UB",
"postcode": "W5 4UB",
"latitude": 51.4986,
"longitude": -0.3148,
"year": np.nan,
"total_pupils": np.nan,
"eligible_pupils": np.nan,
"rwm_expected_pct": np.nan,
}])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
# Two arguments: get_supplementary_data(db, urn). See backend/app.py.
monkeypatch.setattr(
app_module, "get_supplementary_data",
lambda db, urn: {"admission_distance": {"distance_m": 772.49,
"year": 2024}})
monkeypatch.setattr(flags, "is_enabled", lambda name: flag_on)
client = TestClient(app_module.app, raise_server_exceptions=False)
res = client.get("/api/schools/150275")
assert res.status_code == 200, res.text
return res.json()
def test_the_distance_field_is_absent_when_the_flag_is_off(monkeypatch):
"""Absent, not null, and withheld at the source.
/api/schools/ is public and unauthenticated. Leaving a withheld field in
the payload while declining to render it hands the record to anyone who
opens the network tab — the reasoning already recorded in c9a1892.
"""
body = _school_payload(monkeypatch, flag_on=False)
assert "admission_distance" not in body
def test_the_distance_field_is_present_when_the_flag_is_on(monkeypatch):
body = _school_payload(monkeypatch, flag_on=True)
assert body["admission_distance"]["distance_m"] == 772.49
+420
View File
@@ -0,0 +1,420 @@
"""Tests for the place registry (spec 2026-08-21).
The registry is built from the in-memory school DataFrame, so these build a
small frame directly rather than touching a database.
"""
import numpy as np
import pandas as pd
import pytest
from backend.places import MIN_SCHOOLS, build_place_registry
def _df(rows: list[dict]) -> pd.DataFrame:
base = {
"year": 202425, "ofsted_grade": 2.0, "ofsted_date": None,
"rwm_expected_pct": 60.0, "attainment_8_score": np.nan,
"phase": "Primary", "postcode": "AA1 1AA",
}
return pd.DataFrame([{**base, **r} for r in rows])
def _town(n: int, town: str, la: str, start: int = 100000, **kw) -> list[dict]:
"""`start` offsets the URNs so two calls can describe different schools —
the Bedford case needs two authorities' worth of distinct URNs in one
town."""
return [
{"urn": start + i, "school_name": f"{town} School {i}",
"town": town, "local_authority": la, **kw}
for i in range(n)
]
def test_town_clearing_the_threshold_is_published():
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
assert "town:brentwood" in reg
assert reg["town:brentwood"].name == "Brentwood"
assert len(reg["town:brentwood"].urns) == MIN_SCHOOLS
def test_town_below_the_threshold_is_not_published():
reg = build_place_registry(_df(_town(MIN_SCHOOLS - 1, "Crosby", "Sefton")))
assert "town:crosby" not in reg
def test_a_town_below_threshold_still_names_its_authority():
# The route layer needs somewhere to 301 to.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS - 1, "Crosby", "Sefton") + _town(MIN_SCHOOLS, "Bootle", "Sefton")))
assert "authority:sefton" in reg
def test_town_and_authority_of_the_same_name_are_separate_places():
# 67 real collisions. Neither set contains the other: Bedford the town has
# 104 schools, Bedford the authority 86, because postal towns cross
# authority boundaries.
rows = (_town(MIN_SCHOOLS, "Bedford", "Bedford")
+ _town(MIN_SCHOOLS, "Bedford", "Central Bedfordshire", start=200000))
reg = build_place_registry(_df(rows))
town, authority = reg["town:bedford"], reg["authority:bedford"]
assert set(town.urns) != set(authority.urns)
assert len(town.urns) == MIN_SCHOOLS * 2 # both authorities' schools
assert len(authority.urns) == MIN_SCHOOLS # only this authority's
def test_schools_without_publishable_data_do_not_count_toward_the_threshold():
rows = _town(MIN_SCHOOLS, "Ghosttown", "Nowhere")
for r in rows:
r["rwm_expected_pct"] = np.nan
r["ofsted_grade"] = np.nan
reg = build_place_registry(_df(rows))
assert "town:ghosttown" not in reg
def test_blank_town_is_ignored():
rows = _town(MIN_SCHOOLS, "", "Essex")
reg = build_place_registry(_df(rows))
assert not any(k.startswith("town:") for k in reg)
def test_a_school_is_counted_once_even_with_several_years_of_rows():
rows = []
for year in (202324, 202425):
rows += [{**r, "year": year} for r in _town(MIN_SCHOOLS, "Beccles", "Suffolk")]
reg = build_place_registry(_df(rows))
assert len(reg["town:beccles"].urns) == MIN_SCHOOLS
def test_locality_groups_schools_by_outcode(monkeypatch):
# The GIAS town field collapses 1,819 London schools into "London", so a
# locality is defined by its postcode districts instead.
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"battersea": ("Battersea", ("SW11",))})
rows = _town(MIN_SCHOOLS, "London", "Wandsworth")
for r in rows:
r["postcode"] = "SW11 2AA"
reg = build_place_registry(_df(rows))
assert reg["locality:battersea"].name == "Battersea"
assert len(reg["locality:battersea"].urns) == MIN_SCHOOLS
def test_locality_below_the_threshold_is_not_published(monkeypatch):
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"nowhere": ("Nowhere", ("ZZ99",))})
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "London", "Wandsworth")))
assert "locality:nowhere" not in reg
def test_a_locality_may_not_shadow_a_viable_town(monkeypatch, caplog):
"""A colliding locality is skipped loudly, and the town survives.
This used to raise, which took down sitemap generation for all 25,000
school pages the first time a curated slug met a real GIAS town. Curated
data must not be able to break the site — and GIAS town names change with
no code change at all, so the raise could fire spontaneously.
"""
import logging
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"brentwood": ("Brentwood", ("CM13",))})
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
with caplog.at_level(logging.ERROR):
reg = build_place_registry(_df(rows))
assert "locality:brentwood" not in reg # skipped
assert "town:brentwood" in reg # the town is untouched
assert "brentwood" in caplog.text # and it was loud about it
def test_a_locality_collision_does_not_break_the_rest_of_the_registry(monkeypatch):
# The whole point of skipping rather than raising.
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"brentwood": ("Brentwood", ("CM13",))})
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
reg = build_place_registry(_df(rows))
assert "authority:essex" in reg
assert "outcode:cm13" in reg
def test_outcode_places_are_built_from_postcodes():
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:cm13"].name == "CM13"
assert len(reg["outcode:cm13"].urns) == MIN_SCHOOLS
def test_malformed_postcodes_do_not_create_places():
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "not a postcode"
reg = build_place_registry(_df(rows))
assert not any(k.startswith("outcode:") for k in reg)
def test_every_curated_locality_is_structurally_valid():
# Guards the hand-maintained file: real slug, real name, real outcodes.
import re
from backend.localities import LOCALITY_OUTCODES
assert LOCALITY_OUTCODES, "the curated locality list must not be empty"
for slug, (name, outcodes) in LOCALITY_OUTCODES.items():
assert re.fullmatch(r"[a-z0-9-]+", slug), slug
assert name.strip() == name and name, slug
assert outcodes, f"{slug} has no outcodes"
for oc in outcodes:
assert re.fullmatch(r"[A-Z]{1,2}\d{1,2}[A-Z]?", oc), (slug, oc)
def test_the_pipeline_seed_mirrors_the_canonical_module():
"""Two copies with no drift guard is worse than one copy.
backend/localities.py is canonical because the backend image does not
contain pipeline/. The seed exists so the warehouse can join on the same
definitions, and this is what stops the two diverging — the same
arrangement assert_gias_code_names_match_seed.sql gives gias_codes.
"""
import csv
from pathlib import Path
from backend.localities import LOCALITY_OUTCODES
seed_path = (Path(__file__).resolve().parents[2]
/ "pipeline/transform/seeds/locality_outcodes.csv")
assert seed_path.exists(), f"missing seed mirror at {seed_path}"
seed = {
row["locality_slug"]: (row["locality_name"],
tuple(row["outcodes"].split("|")))
for row in csv.DictReader(seed_path.open())
}
assert seed == LOCALITY_OUTCODES
def test_no_curated_locality_names_a_london_borough():
"""Boroughs are authorities and already have a page.
A locality defined by two or three outcodes inside a borough would be a
partial, near-duplicate subset of that authority page — the exact
thin-content failure the two-namespace design exists to avoid. Hackney,
Islington, Greenwich and Ealing were all in the first draft.
Hardcoded rather than read from the corpus because this must fail in CI,
where there is no database.
"""
from backend.localities import LOCALITY_OUTCODES
boroughs = {
"barking-and-dagenham", "barnet", "bexley", "brent", "bromley",
"camden", "croydon", "ealing", "enfield", "greenwich", "hackney",
"hammersmith-and-fulham", "haringey", "harrow", "havering",
"hillingdon", "hounslow", "islington", "kensington-and-chelsea",
"kingston-upon-thames", "lambeth", "lewisham", "merton", "newham",
"redbridge", "richmond-upon-thames", "southwark", "sutton",
"tower-hamlets", "waltham-forest", "wandsworth", "westminster",
}
named = boroughs & set(LOCALITY_OUTCODES)
assert not named, (
f"these are boroughs, not districts: {sorted(named)} - they already "
"have an authority page covering every school"
)
def test_a_place_names_every_authority_it_straddles():
"""SW19 is mostly Merton but partly Wandsworth.
A quarter of viable outcodes and a third of viable towns cross an
authority boundary, so naming only the largest asserts something false.
"""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
names = [n for n, _ in reg["outcode:sw19"].authorities]
assert names == ["Merton", "Wandsworth"] # largest first
assert dict(reg["outcode:sw19"].authorities)["Wandsworth"] == 7
def test_the_redirect_target_stays_a_single_authority():
# parent_authority and authorities do different jobs: a 301 needs one
# target, the page needs the truth.
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:sw19"].parent_authority == "Merton"
def test_a_stray_authority_below_the_share_threshold_is_not_named():
# GIAS carries postcode errors — EN6 lists two Shropshire schools among
# fourteen in Hertfordshire. Printing those as though real would be worse
# than omitting them.
rows = (_town(30, "Barnet", "Hertfordshire", start=300000)
+ _town(1, "Barnet", "Shropshire", start=400000))
for r in rows:
r["postcode"] = "EN6 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:en6"].authorities] == ["Hertfordshire"]
def test_a_sentinel_authority_is_never_named():
rows = (_town(20, "London", "Merton", start=300000)
+ _town(6, "London", "Does not apply", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:sw19"].authorities] == ["Merton"]
def test_a_place_always_names_at_least_one_authority():
# Even when every authority is below the share threshold, the page has to
# say where the place is.
rows = []
for i, la in enumerate(["A", "B", "C", "D", "E", "F", "G"]):
rows += _town(1, "Fragmented", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
place = reg.get("town:fragmented")
assert place is not None
assert len(place.authorities) == 1
def test_every_qualifying_authority_is_named_with_no_cap():
"""An earlier cut stopped at three, dropping the fourth silently.
That truncation bit exactly where the information matters most — a
genuinely fragmented place — and nothing recorded it.
"""
rows = []
for i, la in enumerate(["Hackney", "Lambeth", "Westminster", "Lewisham"]):
rows += _town(3, "Fourway", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
assert len(reg["town:fourway"].authorities) == 4
def test_the_redirect_target_is_the_authority_named_first():
"""They were computed separately — mode() against value_counts() — and on
an exact tie pandas does not guarantee the two agree."""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
place = build_place_registry(_df(rows))["outcode:sw19"]
assert place.parent_authority == place.authorities[0][0]
def test_a_place_never_redirects_to_a_sentinel_authority():
# Deriving the parent from `authorities` inherits its sentinel filter.
rows = (_town(6, "Someplace", "Does not apply", start=300000)
+ _town(5, "Someplace", "Essex", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:someplace"].parent_authority == "Essex"
def test_spellings_of_one_place_are_merged_not_overwritten():
"""GIAS spells the same place several ways, and they share a URL.
"London" (1,819 schools) and "LONDON" (12) both slugify to `london`.
Grouping by raw value let the later group overwrite the earlier one, so
the page could have shown twelve schools instead of 1,819 — silently, and
depending on row order.
"""
rows = (_town(6, "Weston-super-Mare", "North Somerset", start=300000)
+ _town(5, "Weston-Super-Mare", "North Somerset", start=400000))
reg = build_place_registry(_df(rows))
assert len(reg["town:weston-super-mare"].urns) == 11
def test_the_merged_place_takes_its_most_common_spelling():
rows = (_town(9, "Newcastle-under-Lyme", "Staffordshire", start=300000)
+ _town(5, "NEWCASTLE-UNDER-LYME", "Staffordshire", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:newcastle-under-lyme"].name == "Newcastle-under-Lyme"
def test_a_phase_page_needs_results_not_merely_publishable_schools():
"""/schools/kent/primary published with none of its five rows scored.
The threshold counted schools that were publishable — a result OR an
Ofsted grade — while the page exists for its results column. Forty-four
phase pages were majority-blank; one had no results at all.
"""
rows = _town(MIN_SCHOOLS, "Kent", "Kent")
for r in rows:
r["rwm_expected_pct"] = np.nan # Ofsted only, no results
reg = build_place_registry(_df(rows))
assert "town:kent" in reg # the place still publishes
assert not reg["town:kent"].publishes_phase("primary")
def test_a_phase_page_publishes_once_enough_schools_carry_a_result():
rows = _town(MIN_SCHOOLS, "Beccles", "Suffolk")
reg = build_place_registry(_df(rows))
assert reg["town:beccles"].publishes_phase("primary")
def test_a_publishing_phase_page_still_lists_its_unscored_schools():
"""The threshold gates whether the page exists; it does not filter rows.
A parent looking up a school by name has to find it whether or not it
published results.
"""
scored = _town(MIN_SCHOOLS, "Beccles", "Suffolk", start=300000)
unscored = _town(2, "Beccles", "Suffolk", start=400000)
for r in unscored:
r["rwm_expected_pct"] = np.nan
reg = build_place_registry(_df(scored + unscored))
place = reg["town:beccles"]
assert place.publishes_phase("primary")
assert len(place.phase_urns["primary"]) == MIN_SCHOOLS + 2
def test_the_secondary_threshold_counts_its_own_metric():
# A town full of scored primaries must not thereby publish a secondary page.
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
reg = build_place_registry(_df(rows))
assert not reg["town:brentwood"].publishes_phase("secondary")
def test_an_outcode_publishes_no_phase_variants():
"""There is no /schools/near/[outcode]/[phase] route, by design.
Nobody searches "primary schools in SW11", so the spec gives outcodes no
phase variants. The registry computed them anyway, and the place page —
which links whatever phases the registry reports — put two 404s on every
outcode page in the site.
This is the single rule now: a kind with no phase route reports no phases,
so neither the page nor the sitemap can offer one.
"""
rows = [{"urn": 500000 + i, "school_name": f"SW11 School {i}",
"town": "London", "local_authority": "Wandsworth",
"postcode": "SW11 1AA"} for i in range(MIN_SCHOOLS + 3)]
reg = build_place_registry(_df(rows))
place = reg["outcode:sw11"]
assert place.phase_urns == {}
assert not place.publishes_phase("primary")
assert not place.publishes_phase("secondary")
def test_an_authority_still_publishes_phase_variants():
"""Authorities keep theirs — "primary schools in Kent" is a real query,
and /schools/authority/[la]/[phase] is the route that serves it."""
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Maidstone", "Kent")))
assert reg["authority:kent"].publishes_phase("primary")
+139
View File
@@ -0,0 +1,139 @@
"""Tests for the places API (spec 2026-08-21)."""
import numpy as np
import pandas as pd
import pytest
from fastapi.testclient import TestClient
def _schools_df() -> pd.DataFrame:
base = {
"local_authority": "Essex", "school_type": "Academy",
"phase": "Primary", "year": 202425, "ofsted_grade": 2.0,
"ofsted_date": None, "attainment_8_score": np.nan,
"town": "Brentwood", "postcode": "CM13 1AA", "status": "Open",
"address": "1 Test Street", "latitude": 51.6, "longitude": 0.3,
}
return pd.DataFrame([
{**base, "urn": 100000 + i, "school_name": f"Brentwood School {i}",
"rwm_expected_pct": 50.0 + i}
for i in range(6)
])
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _schools_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_registry_lists_each_published_place(client):
body = client.get("/api/places").json()
slugs = {(p["kind"], p["slug"]) for p in body["places"]}
assert ("town", "brentwood") in slugs
assert ("authority", "essex") in slugs
assert ("outcode", "cm13") in slugs
def test_registry_carries_a_count_per_place(client):
body = client.get("/api/places").json()
town = next(p for p in body["places"] if p["slug"] == "brentwood")
assert town["count"] == 6
def test_place_detail_returns_its_schools_alphabetically(client):
"""A place page is read by someone looking for a school they can name.
Scanning for it is what the order should serve, so the list is A-Z.
/api/rankings is where the league-table ordering lives.
"""
body = client.get("/api/places/town/brentwood").json()
assert body["place"]["name"] == "Brentwood"
names = [s["school_name"] for s in body["schools"]]
assert names == sorted(names, key=str.lower)
def test_place_ordering_ignores_case(client):
body = client.get("/api/places/town/brentwood").json()
names = [s["school_name"] for s in body["schools"]]
# A capitalised name must not sort ahead of every lowercase one.
assert names == sorted(names, key=str.lower)
def test_the_rankings_endpoint_still_ranks_by_metric(client):
# Alphabetical is a place-page decision, not a site-wide one.
body = client.get("/api/rankings?metric=rwm_expected_pct&phase=primary").json()
scores = [r["rwm_expected_pct"] for r in body.get("rankings", [])
if r.get("rwm_expected_pct") is not None]
assert scores == sorted(scores, reverse=True)
def test_place_detail_carries_the_local_average(client):
body = client.get("/api/places/town/brentwood").json()
# 50..55 inclusive
assert body["averages"]["rwm_expected_pct"] == pytest.approx(52.5)
def test_phase_filter_narrows_the_school_list(client):
body = client.get("/api/places/town/brentwood?phase=secondary").json()
assert body["schools"] == []
def test_unknown_place_404s(client):
assert client.get("/api/places/town/atlantis").status_code == 404
def test_unknown_kind_404s(client):
assert client.get("/api/places/planet/mars").status_code == 404
def _straddling_df() -> pd.DataFrame:
"""Eight schools in CM13: six in Essex, which has a page, and two in an
authority too small to have one.
Two, not one: the registry ignores an authority holding a single school in
a place, because GIAS carries occasional postcode errors."""
df = _schools_df()
extra = df.iloc[:2].copy()
extra["urn"] = [200000, 200001]
extra["school_name"] = ["Scilly School 0", "Scilly School 1"]
extra["local_authority"] = "Isles Of Scilly"
return pd.concat([df, extra], ignore_index=True)
@pytest.fixture()
def straddling_client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _straddling_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _straddling_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_an_outcode_reports_no_phases_because_it_has_no_phase_route(client):
body = client.get("/api/places/outcode/cm13").json()
assert body["place"]["phases"] == []
def test_an_authority_reports_the_phases_it_publishes(client):
body = client.get("/api/places/authority/essex").json()
assert body["place"]["phases"] == ["primary"]
def test_an_authority_without_a_page_is_named_but_carries_no_slug(straddling_client):
"""Two English authorities — City of London and the Isles of Scilly — hold
fewer than the five schools a page needs, so they have no page.
Naming them is still right: the page says where the place is. Linking them
would not be. A null slug is what tells the page to print the name plainly
rather than invent a URL that 404s.
"""
body = straddling_client.get("/api/places/outcode/cm13").json()
by_name = {a["name"]: a for a in body["place"]["authorities"]}
assert by_name["Essex"]["slug"] == "essex"
assert by_name["Isles Of Scilly"]["slug"] is None
+147
View File
@@ -0,0 +1,147 @@
"""The rate-limit bucket must be the caller, not the proxy in front of them.
`get_remote_address` reads request.client.host. In staging and production the
backend has no published ports and its only caller is the Next proxy, so that
host is the Next container — one bucket for every browser user on the site.
Measured before this fix: 70 concurrent requests, 60 served and 10 refused.
"""
from starlette.datastructures import Headers
from backend.app import client_key
class _Req:
"""Enough of a Request for the key function: headers and a client host."""
def __init__(self, headers: dict, host: str = "10.0.0.9"):
self.headers = Headers(headers)
self.client = type("C", (), {"host": host})()
self.scope = {"type": "http", "client": (host, 0),
"headers": [(k.lower().encode(), v.encode())
for k, v in headers.items()]}
def test_cloudflare_header_wins():
# Cloudflare sets CF-Connecting-IP and overwrites any client-supplied
# value, so it is trustworthy in a way a parsed XFF chain is not.
assert client_key(_Req({"cf-connecting-ip": "203.0.113.7"})) == "203.0.113.7"
def test_forwarded_for_is_the_fallback_and_takes_the_first_entry():
# Left-most is the original client; everything after it is proxies.
assert client_key(
_Req({"x-forwarded-for": "203.0.113.7, 10.0.0.2"})) == "203.0.113.7"
def test_remote_address_is_the_last_resort():
assert client_key(_Req({}, host="10.0.0.9")) == "10.0.0.9"
def test_cloudflare_header_beats_forwarded_for():
key = client_key(_Req({"cf-connecting-ip": "203.0.113.7",
"x-forwarded-for": "198.51.100.1"}))
assert key == "203.0.113.7"
def test_two_callers_get_two_buckets():
# The whole point: one user exhausting their limit must not refuse another.
a = client_key(_Req({"cf-connecting-ip": "203.0.113.7"}))
b = client_key(_Req({"cf-connecting-ip": "203.0.113.8"}))
assert a != b
def test_whitespace_is_stripped():
# "a, b" split on comma leaves a leading space on every entry but the
# first; an unstripped key silently creates a second bucket per client.
assert client_key(_Req({"x-forwarded-for": " 203.0.113.7 ,10.0.0.2"})) \
== "203.0.113.7"
# ---------------------------------------------------------------------------
# The ceiling that header rotation cannot raise.
# ---------------------------------------------------------------------------
import pytest
from fastapi.testclient import TestClient
@pytest.fixture()
def api(monkeypatch):
from backend import app as app_module
from backend.config import settings
monkeypatch.setattr(settings, "global_rate_limit_per_minute", 5)
monkeypatch.setattr(app_module, "_global_window", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def _ceiling_req(path: str, host: str):
"""Enough of a Request for exempt_from_ceiling: a path and a peer host."""
return type("R", (), {
"url": type("U", (), {"path": path})(),
"client": type("C", (), {"host": host})(),
})()
def _get(client, path="/api/flags", cf=None):
headers = {"cf-connecting-ip": cf} if cf else {}
return client.get(path, headers=headers)
def test_rotating_the_cloudflare_header_cannot_buy_unlimited_requests(api):
"""The attack the per-client keying opened up.
client_key trusts CF-Connecting-IP, and nothing in this process can tell an
edge-set header from an attacker-set one — that distinction can only be
made at Cloudflare, with Authenticated Origin Pulls or an origin firewall.
A caller reaching the origin directly can therefore mint a fresh
rate-limit bucket per request and evade per-client limits entirely.
Per-client fairness is still the right default; this is the backstop that
bounds what evading it can achieve. Without it, correct keying would be a
net regression against abuse compared with the shared bucket it replaced.
"""
codes = [_get(api, cf=f"203.0.113.{i}").status_code for i in range(8)]
assert codes.count(200) == 5
assert codes.count(429) == 3
def test_the_ceiling_says_which_limit_was_hit(api):
# Distinguishable from slowapi's per-client 429, or an operator reading
# logs cannot tell "one noisy client" from "the origin is saturated".
for i in range(5):
_get(api, cf=f"203.0.113.{i}")
refused = _get(api, cf="203.0.113.99")
assert refused.status_code == 429
assert "capacity" in refused.json()["detail"].lower()
assert refused.headers.get("retry-after")
def test_traffic_below_the_ceiling_is_untouched(api):
codes = [_get(api, cf=f"203.0.113.{i}").status_code for i in range(5)]
assert codes == [200] * 5
def test_the_container_healthcheck_is_exempt(api):
"""The healthcheck runs `curl http://localhost:80/api/data-info` inside the
container. If the ceiling could starve it, saturation would fail the
healthcheck, restart the container, and turn a load spike into an outage
loop — the ceiling has to protect the origin, not kill it.
"""
from backend.app import exempt_from_ceiling
assert exempt_from_ceiling(_ceiling_req("/api/data-info", "127.0.0.1"))
assert exempt_from_ceiling(_ceiling_req("/api/data-info", "::1"))
# Everyone else is counted.
assert not exempt_from_ceiling(_ceiling_req("/api/data-info", "10.0.0.9"))
def test_the_ceiling_ignores_non_api_paths():
# Sitemaps and robots.txt are served by this app too, and a crawler
# fetching them must not be refused because the API is busy.
from backend.app import exempt_from_ceiling
assert exempt_from_ceiling(_ceiling_req("/sitemap.xml", "10.0.0.9"))
assert exempt_from_ceiling(_ceiling_req("/robots.txt", "10.0.0.9"))
+129 -1
View File
@@ -59,9 +59,14 @@ def static_child(monkeypatch) -> str:
def test_every_loc_uses_the_www_host(sitemaps):
# The apex 301s to www. A <loc> that redirects burns a crawl per URL.
# Checked across every file, index included, not just one.
#
# A child can legitimately be empty — this fixture holds two schools and no
# town clearing the threshold — so the presence check applies only to files
# that carry URLs. The absence check applies to all of them.
for name, xml in sitemaps.items():
assert "https://www.schoolcompare.co.uk" in xml, name
assert "https://schoolcompare.co.uk" not in xml, name
if "<loc>" in xml:
assert "https://www.schoolcompare.co.uk" in xml, name
def test_school_with_results_is_listed(schools_child):
@@ -176,3 +181,126 @@ def test_children_are_chunked_under_the_limit(monkeypatch):
def test_build_sitemap_still_returns_the_index(sitemap):
# lifespan and the admin endpoint call build_sitemap(); keep it working.
assert "<sitemapindex" in sitemap
def test_school_with_results_in_an_earlier_year_is_still_listed(monkeypatch):
"""Regression: The Mallard Academy (150367).
Real KS2 results 2015-16 to 2018-19, then null rows from 2022-23 onward
because the school stopped reporting. The first cut tested the latest
year's row alone and dropped it, along with ~220 others, even though its
detail page shows all four years of results.
"""
from backend import app as app_module
import pandas as _pd
base = {"local_authority": "Testshire", "school_type": "Academy",
"phase": "Primary", "ofsted_date": None, "ofsted_grade": np.nan,
"attainment_8_score": np.nan, "urn": 150367,
"school_name": "Mallard Academy"}
df = _pd.DataFrame([
{**base, "year": 201819, "rwm_expected_pct": 67.0},
{**base, "year": 202324, "rwm_expected_pct": np.nan},
{**base, "year": 202425, "rwm_expected_pct": np.nan},
])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
xml = app_module.build_sitemaps()["schools-1.xml"]
assert "/school/150367-mallard-academy" in xml
def test_school_with_no_results_in_any_year_is_still_omitted(monkeypatch):
"""The fix must not turn into "list everything"."""
from backend import app as app_module
import pandas as _pd
base = {"local_authority": "Testshire", "school_type": "Academy",
"phase": "Primary", "ofsted_date": None, "ofsted_grade": np.nan,
"attainment_8_score": np.nan, "rwm_expected_pct": np.nan,
"urn": 100002, "school_name": "Ghost Primary"}
df = _pd.DataFrame([{**base, "year": y} for y in (202324, 202425)])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
assert "/school/100002" not in app_module.build_sitemaps()["schools-1.xml"]
def _places_df() -> pd.DataFrame:
base = {
"local_authority": "Essex", "school_type": "Academy",
"phase": "Primary", "year": 202425, "ofsted_grade": 2.0,
"ofsted_date": None, "attainment_8_score": np.nan,
"town": "Brentwood", "postcode": "CM13 1AA",
}
return pd.DataFrame([
{**base, "urn": 100000 + i, "school_name": f"Brentwood School {i}",
"rwm_expected_pct": 60.0}
for i in range(6)
])
@pytest.fixture()
def place_sitemaps(monkeypatch) -> dict:
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _places_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return app_module.build_sitemaps()
def test_place_children_are_listed_in_the_index(place_sitemaps):
index = place_sitemaps["sitemap.xml"]
assert "/sitemaps/places-1.xml" in index
assert "/sitemaps/outcodes-1.xml" in index
def test_town_and_authority_urls_use_their_own_namespaces(place_sitemaps):
xml = place_sitemaps["places-1.xml"]
assert "<loc>https://www.schoolcompare.co.uk/schools/brentwood</loc>" in xml
assert "<loc>https://www.schoolcompare.co.uk/schools/authority/essex</loc>" in xml
def test_outcode_urls_live_in_their_own_child(place_sitemaps):
assert "/schools/near/cm13" in place_sitemaps["outcodes-1.xml"]
assert "/schools/near/cm13" not in place_sitemaps["places-1.xml"]
def test_place_urls_carry_no_priority_or_changefreq(place_sitemaps):
for name in ("places-1.xml", "outcodes-1.xml"):
assert "<priority>" not in place_sitemaps[name]
assert "<changefreq>" not in place_sitemaps[name]
def test_phase_variants_are_submitted_where_the_phase_clears_the_threshold(place_sitemaps):
# "primary schools in beccles" is the query shape the baseline showed, so
# each variant is its own page and has to be submitted. Emitting only the
# bare place URL left ~950 of them reachable by nothing.
xml = place_sitemaps["places-1.xml"]
assert "<loc>https://www.schoolcompare.co.uk/schools/brentwood/primary</loc>" in xml
def test_a_phase_below_its_own_threshold_is_not_submitted(place_sitemaps):
# The fixture is six primaries and no secondaries.
xml = place_sitemaps["places-1.xml"]
assert "/schools/brentwood/secondary" not in xml
def test_outcodes_get_no_phase_variants(place_sitemaps):
# Nobody searches "primary schools in CM13"; the routes do not exist.
xml = place_sitemaps["outcodes-1.xml"]
assert "/primary" not in xml and "/secondary" not in xml
def test_authority_phase_variants_are_submitted_in_their_own_namespace(place_sitemaps):
"""302 of these were already in the sitemap, and every one 404'd.
The spec gives authorities a phase route; the plan built the bare
authority route and dropped it. Nothing noticed because the sitemap was
written from the registry, which was right, while the routes were written
by hand. This test fails if the URL ever leaves the sitemap; the e2e
journey fails if the route ever leaves the app.
"""
xml = place_sitemaps["places-1.xml"]
assert ("<loc>https://www.schoolcompare.co.uk"
"/schools/authority/essex/primary</loc>") in xml
# And never in the town namespace, which is a different set of schools.
assert "/schools/essex/primary" not in xml
+157
View File
@@ -0,0 +1,157 @@
"""Tests for school autosuggest (spec 2026-08-26)."""
from backend import data_loader
class _FakeDocs:
def __init__(self, hits, explode=False):
self._hits = hits
self._explode = explode
self.last_params = None
def search(self, params):
self.last_params = params
if self._explode:
raise RuntimeError("typesense is down")
return {"hits": [{"document": d} for d in self._hits]}
class _FakeClient:
def __init__(self, hits, explode=False):
self.docs = _FakeDocs(hits, explode)
self.collections = {"schools": type("C", (), {"documents": self.docs})()}
_HIT = {
"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school",
}
def _use(monkeypatch, client):
monkeypatch.setattr(data_loader, "_get_typesense_client", lambda: client)
def test_returns_the_fields_a_suggestion_needs(monkeypatch):
# Local authority is not decoration: there are many schools called
# "St Mary's", and a list without it cannot be chosen between.
_use(monkeypatch, _FakeClient([_HIT]))
out = data_loader.suggest_schools_typesense("breck")
assert out == [{
"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school",
}]
def test_a_missing_optional_field_becomes_an_empty_string(monkeypatch):
# phase and school_type are optional in the Typesense schema. A missing
# key must not KeyError in the keystroke path.
_use(monkeypatch, _FakeClient([{"urn": 1, "school_name": "X",
"local_authority": "Y", "postcode": "Z"}]))
out = data_loader.suggest_schools_typesense("x")
assert out[0]["phase"] == "" and out[0]["school_type"] == ""
def test_typesense_unavailable_gives_no_suggestions_rather_than_raising(monkeypatch):
_use(monkeypatch, None)
assert data_loader.suggest_schools_typesense("anything") == []
def test_a_typesense_error_gives_no_suggestions_rather_than_raising(monkeypatch):
_use(monkeypatch, _FakeClient([], explode=True))
assert data_loader.suggest_schools_typesense("anything") == []
def test_the_limit_is_passed_through_and_clamped(monkeypatch):
client = _FakeClient([])
_use(monkeypatch, client)
data_loader.suggest_schools_typesense("x", limit=500)
assert client.docs.last_params["per_page"] == 20
def _client(monkeypatch, rows, *, blow_up_dataframe=False):
from fastapi.testclient import TestClient
from backend import app as app_module
monkeypatch.setattr(app_module, "suggest_schools_typesense",
lambda q, limit=8: rows)
if blow_up_dataframe:
def _boom():
raise AssertionError("the suggest path must not load the DataFrame")
monkeypatch.setattr(app_module, "load_school_data", _boom)
monkeypatch.setattr(app_module, "load_latest_school_data", _boom)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_the_endpoint_returns_suggestions(monkeypatch):
body = _client(monkeypatch, [_HIT]).get("/api/suggest?q=breck").json()
assert body["suggestions"][0]["school_name"] == "Brecknock Primary School"
def test_the_endpoint_never_touches_the_dataframe(monkeypatch):
"""The whole reason this is not a mode of /api/schools.
That endpoint filters and sorts 25,000 rows of pandas per query, holding
the GIL. Per keystroke, that is the cost this endpoint exists to avoid.
"""
res = _client(monkeypatch, [_HIT], blow_up_dataframe=True).get("/api/suggest?q=breck")
assert res.status_code == 200
assert res.json()["suggestions"]
def test_a_one_character_query_returns_nothing_and_does_not_error(monkeypatch):
# The keystroke path never errors on ordinary input.
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=b")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_a_blank_query_returns_nothing_and_does_not_error(monkeypatch):
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_typesense_down_is_an_empty_list_not_a_500(monkeypatch):
res = _client(monkeypatch, []).get("/api/suggest?q=breck")
assert res.status_code == 200
assert res.json() == {"suggestions": []}
def test_the_response_is_cacheable(monkeypatch):
# Prefix queries repeat enormously across users, and school names change
# once a year. Without this the endpoint pays full price every keystroke.
res = _client(monkeypatch, [_HIT]).get("/api/suggest?q=breck")
assert "s-maxage" in res.headers.get("cache-control", "")
assert res.headers.get("etag")
def test_a_malformed_urn_does_not_raise(monkeypatch):
"""The docstring promises "never raises"; the parsing loop sat outside the
try, so int(None) or int("abc") would have turned a keystroke into a 500.
Typesense declares urn as int32, so this should be unreachable — but the
contract is what the caller relies on, and a search index is a separate
system that can be reindexed by something other than this code.
"""
_use(monkeypatch, _FakeClient([{"urn": None, "school_name": "X",
"local_authority": "Y", "postcode": "Z"}]))
assert data_loader.suggest_schools_typesense("x") == []
def test_a_malformed_row_does_not_discard_the_good_ones(monkeypatch):
# One bad document must not blank the whole dropdown.
_use(monkeypatch, _FakeClient([
{"urn": "not-a-number", "school_name": "Bad", "local_authority": "Y",
"postcode": "Z"},
_HIT,
]))
out = data_loader.suggest_schools_typesense("x")
assert [r["urn"] for r in out] == [100010]
def test_a_hit_with_no_document_does_not_raise(monkeypatch):
_use(monkeypatch, _FakeClient([{}]))
assert data_loader.suggest_schools_typesense("x") == []
+9
View File
@@ -16,6 +16,8 @@
# ADMIN_API_KEY — Backend admin API key
# TYPESENSE_API_KEY — Typesense admin API key
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: development
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs)
# STAGING_DB_IP — macvlan IP for staging Postgres (default 10.0.1.190)
# STAGING_FRONTEND_IP — macvlan IP for staging frontend (default 10.0.1.151)
@@ -55,6 +57,12 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- unleash_cache:/app/.unleash
depends_on:
sc_database:
condition: service_healthy
@@ -212,3 +220,4 @@ volumes:
postgres_data:
typesense_data:
airflow_logs:
unleash_cache:
+73
View File
@@ -0,0 +1,73 @@
# Portainer Stack Definition for School Compare — UNLEASH (feature flags)
#
# Deploy as a *separate* Portainer stack ("schoolcompare-unleash"), alongside
# the production and staging stacks. It deliberately belongs to neither: a
# staging redeploy must not be able to disturb production's flag state, and a
# production redeploy must not disturb staging's.
#
# One instance serves both environments. Open-source Unleash ships with
# `development` and `production` environments and environment-scoped client
# tokens, so the same flag holds independent state in each — which is what
# lets a feature be on in staging, where the E2E journeys exercise it, while
# production stays dark.
#
# Portainer environment variables (set in Portainer UI -> Stack -> Environment):
# UNLEASH_DB_PASSWORD — PostgreSQL password for the Unleash database
# UNLEASH_ADMIN_PASSWORD — initial admin password for the Unleash UI
# UNLEASH_IP — macvlan IP for the Unleash server (default 10.0.1.152)
services:
# ── PostgreSQL (Unleash's own; nothing else uses it) ──────────────────
unleash_db:
container_name: sc_unleash_postgres
image: postgres:16-alpine
environment:
POSTGRES_USER: unleash
POSTGRES_PASSWORD: ${UNLEASH_DB_PASSWORD}
POSTGRES_DB: unleash
volumes:
- unleash_postgres_data:/var/lib/postgresql/data
networks:
- unleash
healthcheck:
test: ["CMD-SHELL", "pg_isready -U unleash"]
interval: 10s
timeout: 5s
retries: 5
start_period: 10s
restart: unless-stopped
# ── Unleash server (UI + client API on 4242) ──────────────────────────
unleash:
container_name: sc_unleash
image: unleashorg/unleash-server:6
environment:
DATABASE_URL: postgres://unleash:${UNLEASH_DB_PASSWORD}@unleash_db:5432/unleash
DATABASE_SSL: "false"
INIT_ADMIN_API_TOKENS: ""
UNLEASH_DEFAULT_ADMIN_PASSWORD: ${UNLEASH_ADMIN_PASSWORD}
depends_on:
unleash_db:
condition: service_healthy
networks:
unleash: {}
macvlan:
ipv4_address: ${UNLEASH_IP:-10.0.1.152}
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:4242/health || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 30s
restart: unless-stopped
networks:
unleash:
driver: bridge
macvlan:
external:
name: macvlan
volumes:
unleash_postgres_data:
+9
View File
@@ -7,6 +7,8 @@
# ADMIN_API_KEY — Backend admin API key
# TYPESENSE_API_KEY — Typesense admin API key
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: production
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs)
services:
@@ -44,6 +46,12 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- unleash_cache:/app/.unleash
depends_on:
sc_database:
condition: service_healthy
@@ -201,3 +209,4 @@ volumes:
postgres_data:
typesense_data:
airflow_logs:
unleash_cache:
+4
View File
@@ -36,6 +36,10 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- ./data:/app/data:ro
depends_on:
+100
View File
@@ -150,3 +150,103 @@ token Gitea Actions provides automatically (`secrets.GITEA_TOKEN` — no setup
needed), and fails the check only when a finding is rated
**severe** (would break prod, leak data, or corrupt data). Minor findings are
informational and never block a merge.
## Rate limiting, and the Cloudflare gap
Two independent limits protect the API:
- **Per client**, via slowapi, keyed on `CF-Connecting-IP` (falling back to
`X-Forwarded-For`, then the peer address). 60/minute by default;
`/api/suggest` gets 120/minute because typing is bursty.
- **Globally**, via `GlobalRateLimitMiddleware`: a fixed 60-second window over
all `/api/` traffic, `GLOBAL_RATE_LIMIT_PER_MINUTE` (default 3000),
independent of any client identity. Requests from `127.0.0.1` are exempt so
the container healthcheck cannot be starved into a restart loop.
### Open: the origin must only accept Cloudflare
`CF-Connecting-IP` is only meaningful for requests that actually reached the
origin through Cloudflare, and **the application cannot verify that they did**.
Anything able to reach the origin directly can set that header freely and, by
rotating it, mint a fresh rate-limit bucket per request — defeating per-client
limits on every endpoint.
The global ceiling bounds the damage to total origin capacity. It does not fix
the underlying gap, and nothing in the code can. Closing it needs one of:
- **Authenticated Origin Pulls** — Cloudflare presents a client certificate the
origin requires, so non-Cloudflare traffic is refused at TLS.
- **An origin firewall** restricted to Cloudflare's published IP ranges.
Until one is in place, treat per-client limits as protection against accidents
and ordinary load, not against a determined caller.
## Feature flags (Unleash)
Flag state lives in a self-hosted Unleash instance, deployed as its own
Portainer stack from `docker-compose.portainer.unleash.yml`. It is separate
from the application stacks on purpose — redeploying staging must not be able
to disturb production's flags.
The flags themselves are declared in `backend/flags.py`. Unleash holds the
state; the registry holds the list. A flag in the UI that is not in the
registry is orphaned and nothing reads it.
### First-time setup
1. Deploy the stack in Portainer. Set `UNLEASH_DB_PASSWORD`,
`UNLEASH_ADMIN_PASSWORD` and (optionally) `UNLEASH_IP`.
2. Log in to the UI at `http://<UNLEASH_IP>:4242` as `admin`.
3. Create one **client** API token per environment:
- `schoolcompare-staging`, environment **development**
- `schoolcompare-prod`, environment **production**
Client tokens, not admin tokens — the backend only reads.
4. Put each token in the matching Portainer stack's `UNLEASH_API_TOKEN`
variable, and set `UNLEASH_URL` to `http://<UNLEASH_IP>:4242/api`.
5. Redeploy the application stacks.
### Adding a flag to Unleash
**Unleash does not create flags by itself.** The SDK reads definitions from the
server and never registers anything, and metrics for a flag the server has
never heard of are discarded. So a flag declared in `backend/flags.py` will be
evaluated on every request, stay `False` forever, and never appear in the UI
until someone creates it there by hand.
For each flag in the registry, create one in Unleash with:
- **Name** — character for character what `backend/flags.py` declares.
snake_case, no hyphens or spaces. A typo produces a flag that looks correct
in the UI and is read by nothing.
- **Type** — Release. No strategies, constraints or variants: these are plain
on/off switches, by design.
### Turning a feature on
Toggle the flag in the environment matching the stack you mean: **development**
for staging, **production** for prod. The token in each stack is scoped to one
environment, so toggling the other one has no visible effect.
The SDK refreshes every 15 seconds, so the API reflects the change almost at
once; the pages follow on their own schedule, below.
A flip reaches school pages within about five minutes and place pages within
the hour. Next's ISR does the propagating — it revalidates a route at the
*lowest* `revalidate` among that route's fetches, which is 300s for
`/school/[slug]` and 3600s for the place pages. There is no webhook, and
adding one would only be worth it if flips ever needed to be instant.
### When Unleash is unreachable
Every flag evaluates to `False` and the site serves as though nothing were
switched on. That is deliberate — an unfinished feature staying hidden is the
safe direction — but it means a *released* feature disappears if a backend
container cold-starts with an empty cache while Unleash is down. The SDK's
disk cache is on a named volume so restarts keep last-known state, and flags
are removed from the code within 90 days (enforced by a test), which bounds
how long any feature is exposed to this.
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
attempted. That is the correct behaviour for local development and CI, and it
means the test suites need no flag server.
+85
View File
@@ -0,0 +1,85 @@
# Remote branch cleanup, 2026-08-20
# Restore any branch with: git push origin <sha>:refs/heads/<name>
## Deleted: fully merged into main (content is in main)
75677f4252b759ef895e7d5f7c19f8f1745bdb59 add-contact-form-footer
fa1abff642683dfd26ba88a295a0a6710d77147a chore/byline-removal-and-audit-figure
95081d38bdf87764ef5d298676c25fae4cd3b792 chore/remove-parent-view
6877abedebfc1c7d95f1f6ebe68945c62be528ab chore/staged-prod-promotion
090d5f7bec824e083d3252e2c6e636686016304d ci/frontend-checks-speedup
8c3a5cc4e9f551f0190d85357ce7741ad87f3a4c design/cohort-identity
955659580067ca8b81bd77e31e4e2554103f094d feat/allow-analytics-iframe-embed
6828f6cd4417284ea3eb6f088fa20945b8b40ed3 feat/compare-chips-two-per-row
d5cd0abfee226885119665da5d2aa8288b59217f feat/compare-data-foundation
6dd9b04b50bee146682da87efad8fc8b526251c5 feat/compare-frontend-rebuild
96d5fcf5b07b6f175b48e9b20fcb765a320a907f feat/detail-header-details-reveal
f1388ff5bd0af1409823a1e047b7ba84246a0f70 feat/gias-sixth-form-flag
eddf74745f86c9c6d9eb07d867246ff9cc90dc20 feat/hero-artwork-v2
3015c37bac6dc28db58c80fdb9235942c25b83d7 feat/hero-byline
8e763e39d17a4964cf558e51c03f044186371f6d feat/info-popover-tooltip
88c653215d520ab6e902c9de55bf27d86eeb90c3 feat/last-distance-offered
a72323874f7aebdb2e64b6d64a5febd61152d09d feat/last-distance-offered-full
c9a1892bfb0370e0672e5849cba294ddfabe3c65 feat/latest-cutoff-only
4e8df006d75d8be2a1d8529ddf855c445854cba1 feat/near-me-by-search
45ab479062c6a1639facad636fc0cc0cf0fd9155 feat/proposed-to-close-schools
3bf2e8f262cbe058fda6de6f8ea3e050224a51b9 feat/school-detail-visualisations
1f80571b1ff217dc92b660a936c02b5f49d07f0b feat/umami-heatmap-recorder
609bb923d96aa5730131463ea35d5efdd404bf96 feature/ingest-independent-schools
94151c58ea38a9256d66d15161293486a500c7a2 fix/admissions-section-height
79246edc22961c2beb3520437e9d064d3b809d10 fix/annual-dag-ks4-national-selector
59ac9c10b97e0ef1143f57fea06b324e72ac3d4a fix/chart-marker-contrast
9f8dba227c95706ca3527bd48d381e7622cc0a5e fix/compare-chart-refetch-resilience
e74d3882ce78a141fa1a57daa3102d7a58852dc3 fix/compare-expert-fixes
80176cac4db4820e76ea2a156c7a2974bee2f203 fix/compare-final-review-mustfix
f579630fab6c456e26a6984a3e8eebdbe3184202 fix/compare-mockup-drift
dc85254ad2ddf134b4434065d4762cef374d2220 fix/compare-null-year-blanks-chart
43a2c4a6bc539b621f31655aec05ef319a25f343 fix/compare-refresh-and-fetch
d677b5453365b72c81d6df2de62b1fa0d05d684d fix/daily-dag-cache-invalidation
f6bb037c471553e8195b5a8b147467ce0d07a688 fix/detail-all-through
17bd4d5a5eb0b12ca79b97db587f14f7d671e85b fix/detail-chart-truthfulness
4e6be0ce65647410b4ff74f8b763207c28920c26 fix/detail-inclusion-admissions
fdda52ff0af3fae03a4b059a655973cdbb269f91 fix/detail-ofsted-correctness
e36125b24aba254a8d15c5c33b4d2a296e691995 fix/detail-provenance-anchoring
b31e71ac884df6507f567adc046c9fc52d9d310a fix/detail-report-card-render-date
32f8a02862be6a1d49f4c3928b17fcf15d4c94cc fix/detail-trend-chart-taller
e65688d600a86818fe21ae4c61ba27e5b6ec8d7c fix/e2e-brand-assertions
3adea73ee04cdedfab54b0351878b297f72756ad fix/e2e-compare-chips-phase
06e4898c30feaedc471f97aba28ddb0d61379f4d fix/e2e-compare-samephase
9abd020967670a855e80fe5a908c8048a3aa9f14 fix/e2e-distance-locator
acec8135e1ec7c9c3c255e5b23733a0ef862b590 fix/e2e-rankings-year-pick
5944d88f0b1517ef1ef56af1b62270d2cf28e717 fix/expert-signoff-mustfixes
2433101fa08be5df6f170d41790512ffe823d33e fix/font-cascade-and-map-palette
74ca76d150deec6725259d9637ea86d7bb90c683 fix/gias-legacy-fallback
bdaa05cd542f563ef74c45307cd8f7fc465193c9 fix/hero-fallback-and-sharp
d52d384cf23d282b44e9251176f8f3402d808600 fix/hero-map-ios-fullscreen
4043270a77bbe4edb18207fa5f1d94d4747fe12f fix/hero-mobile-and-wording
22e9eb2d48b0d6623e88fd67cb6ca8e5e4074583 fix/homepage-education-accuracy
8d50afef1e8a2b621b7344eadf475b0d609ad7a2 fix/leaflet-specificity-and-font-assertion
b2b2cad5acf534ae7a667d3fb2be15efff37c4a7 fix/list-map-report-card-signal
dc21e80a5e9eebd13aaab84735642eb77cef35e6 fix/mobile-cell-name-size
e5f7f4c959f024c073472122d858333a0f24866c fix/mobile-compare-polish
a00cbe916182d1e04661c750a1ce2e7ad27a68ae fix/mobile-sort-select-overflow
2fd997bfe640c419a6713e85df008463de7f56e8 fix/modal-keyboard-viewport
3e7705756776a0c44d27966dfc972023f1f38b69 fix/ofsted-link-text
ce422e64363e2b03c186ef316832b6de15502d67 fix/promote-status-token
4522cbf64560db1e1cd519e119aa466b42cd16a1 fix/proposed-to-close-copy
15da060e4af37fbae919e0edf25f108266de2585 fix/rankings-admissions-accuracy
6c872ce726f210433354ca38dc5314bc6a467534 fix/rankings-year-validation
1c1df7796194af3d47f8e5ac0a0fbe6f323700e7 fix/report-card-chip-alignment
b2dc4d0779ced02709429afaa5985acf4abda794 fix/results-map-ios-fullscreen
95a5783da1fc994df76cb97238b55596dce4cd8f fix/runtime-api-proxy
8a9ba30cc24e29653a72c03c0e817684b7db7c07 fix/sats-per-level-national
536832a524fc4f9ed858f047e00b9a4d9d429c46 fix/school-detail-nan-500
fef83b3bf244a9bf3cb4afa75dbc433d8725014f fix/schoolbar-sticky-offset
b0c5b6bb57c879477da12c37b15cff950c29ebd3 fix/secondary-anchors-button-affordance
261403bcd2b01aa4f26ee212e26302fc0f769bf9 fix/special-note-full-width
ea5249a2ea6faf6bfa5a1522644387ba59770ec8 fix/standardise-distance-units
e4565e9f158721d4df6b918f2b065de82851f8d9 fix/trends-chart-height
3aad5101a842539105022f9059e85c224fc5973a fix/welsh-establishment-leak
315f1feede70bdf3101d2fdd305b3d06e037fdac perf/batch-supplementary
d2dc78aeb599e16b7ef5019be2b360b08df463bc perf/compare-loading
e098ad4bd1130152705788d4773837b7d1e7112e perf/server-client-split
## Deleted: superseded by PR #110 (content preserved on feat/seo-crawl-hygiene-main)
786ec80dd4de4a3cb674a89e35b3b5e461639289 feat/seo-crawl-hygiene
a5ac0bcd1b37bc10dcbce88f8601d01bf7b3eaaf feat/england-only-corpus
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,278 @@
# W2: The Location Layer — Design
Date: 2026-08-21
Status: awaiting review
Supersedes: workstream W2 in `2026-08-20-seo-programme-design.md`
Scope note: this covers four page families in one spec. Splitting them — towns
and authorities first, outcodes and localities after — was proposed and
declined in favour of building the layer in one pass. The decomposition
argument was that the curated locality seed needs human review and would hold
up 783 pages of measured demand behind it; that risk is accepted here, and the
implementation plan should sequence the seed early enough that review time
does not become the critical path.
## Problem
Location intent is the largest unserved demand the site has. In the 16-month
Search Console baseline it draws **874 impressions, one click, average
position 49.5**. The site does not compete.
Unlike named-school queries — which the same baseline showed to be
navigational and unwinnable, since a parent typing "audley junior school"
wants that school's own website — location queries have no incumbent owner.
Nobody owns "primary schools in Brentwood" the way a school owns its name.
The cause is structural: the site has no page about a place. Every competitor
ranking above it does.
## What the demand actually looks like
Every location query in the baseline is **town or district level**. Not one is
an administrative area:
| Query | Impressions | Position |
|-------|-------------|----------|
| colleges in solihull | 112 | 51.2 |
| schools in ramsey | 64 | 42.5 |
| schools in crosby | 57 | 47.7 |
| primary schools in beccles | 44 | 40.9 |
| private schools in battersea | 41 | 71.9 |
| secondary schools in brentwood | 37 | 56.1 |
| secondary schools in canary wharf | 30 | 35.9 |
Three patterns follow directly, and they drive the whole design.
**Towns, not authorities.** The superseded W2 put `/schools/[la]` first and
towns second. The data inverts that. Brentwood appears four times in different
phrasings; Beccles twice. Both are towns, not authorities.
**Phase is part of the query**, not a filter applied afterwards: "primary
schools in beccles", "secondary schools in brentwood", "colleges in solihull".
**London is searched by district** — Battersea, Canary Wharf — and the GIAS
`town` field cannot serve it at all.
## Measured sizing
Counted against the live corpus of 25,185 schools, not estimated.
| Family | Viable (≥5 schools) | Below threshold |
|--------|--------------------|-----------------|
| Towns | **783** | 907 → redirect to authority |
| Outcodes | **1,760** | 305 |
| Local authorities | 154 | — |
| London localities | ~100–150 (curated) | — |
With phase variants — 783 town pages plus roughly 700 primary and 250
secondary variants, 154 authorities across three variants, 1,760 outcodes and
the curated localities — the total lands near **4,000 pages**. Phase variants
need their own threshold: there are 17,426 primaries but only 4,456 secondaries nationally,
so most towns will support a primary page and not a secondary one.
## Two design problems this spec exists to solve
### 1. Town and authority names collide, and neither contains the other
67 viable towns share a name with a local authority. The obvious fix — let the
authority absorb the town, since it sounds like a superset — **does not work**:
| Place | Schools in the town | Schools in the authority |
|-------|--------------------|-----------------------|
| Bedford | 104 | 86 |
| Birmingham | 520 | 518 |
| Derby | 157 | 119 |
| Doncaster | 152 | 145 |
The authority is the larger set in only 43 of the 67. Postal towns cross
authority boundaries, so these are overlapping sets that happen to share a
name. Publishing both into one namespace produces near-duplicate pages, which
is the specific failure that sinks programmatic SEO.
**Resolution: two namespaces.**
```
/schools/[place] towns and London localities
/schools/[place]/primary
/schools/[place]/secondary
/schools/authority/[la] local authorities
/schools/authority/[la]/primary
/schools/authority/[la]/secondary
/schools/near/[outcode]
```
Outcodes carry no phase variants: nobody searches "primary schools in SW11",
so the variants would be pages without demand.
Every collision disappears by construction. `/schools/[place]` keeps the clean
URL for the pattern that carries the demand; authorities get a namespace whose
purpose is genuinely different — admissions are authority-run, and the
authority page is the one that can speak to catchment policy and LA averages.
A place page and an authority page of the same name must each say plainly
which set of schools they cover, or they read as duplicates to a reader even
when they differ in fact.
### 2. London has no locality field
`town` collapses **1,819 London schools into the single value "London"**. A
page listing all of them is useless, and borough pages do not help because
people search "Battersea", not "Wandsworth".
No single field solves it:
| Search term | `parliamentary_constituency` | postcodes.io `admin_ward` |
|-------------|------------------------------|---------------------------|
| Battersea | **Battersea** ✓ | Northcote / Wandsworth Town ✗ |
| Canary Wharf | Poplar and Limehouse ✗ | **Canary Wharf** ✓ |
| Vauxhall | Vauxhall and Camberwell Green ✗ | **Vauxhall** ✓ |
And neither covers Clapham, Shoreditch or Peckham, which are postal and
colloquial rather than administrative.
**Resolution: a curated seed mapping locality to outcodes.**
```
pipeline/transform/seeds/locality_outcodes.csv
locality_slug,locality_name,outcodes,region
battersea,Battersea,"SW11|SW8",London
canary-wharf,Canary Wharf,"E14",London
clapham,Clapham,"SW4|SW9",London
```
This needs **no new ingestion** — the corpus already has postcodes. It puts
the fuzzy, contested part of the problem in a reviewable file rather than in
derived logic, which suits it: locality boundaries are a judgement, not a
fact. The repo already uses dbt seeds for curated reference data
(`la_code_names.csv`, `gias_code_names.csv`), so this follows an established
pattern.
The seed generalises past London. Any colloquial place — Jesmond, Chorlton,
Clifton — can be defined by its outcodes without a schema change.
**Constraint:** a locality slug may not collide with a viable town slug. The
place registry enforces this and fails the build rather than silently
shadowing a town.
## Architecture
### The place registry
One module owns the question "what places do we publish, and what is in each".
Everything else reads from it: the pages, the sitemap, the internal links.
```
backend/places.py
Place = { kind: "town"|"locality"|"authority"|"outcode",
slug, name, urn_list, parent_authority | None }
build_place_registry(df) -> dict[str, Place]
place_schools(slug, phase=None) -> list[School]
```
Built once at startup from the same DataFrame the sitemap uses, and rebuilt by
the existing `/api/admin/regenerate-sitemap` path after a pipeline run.
Registry construction is where the threshold, the collision rules and the
seed's uniqueness constraint are enforced — in one place, testable without a
browser or a database.
### API
```
GET /api/places the registry: slug, kind, name, count
GET /api/places/{slug}?phase= aggregate + ranked schools for one place
```
`/api/places` is what the sitemap and the internal-link modules enumerate.
### Routes
Next App Router, ISR with the same 7-day revalidate the school pages use.
`generateStaticParams` gated behind an env flag, matching
`PRERENDER_SCHOOLS`, because 3,900 more routes cannot be statically built in
CI on every deploy.
## What each page must contain
A place page that is a name substituted into a template is the thing Google's
helpful-content stance exists to demote. Each page carries computed local
facts that exist nowhere else on the site:
- **H1** matching the query: "Primary schools in Brentwood"
- **Counts framed usefully**: "29 schools, 4 rated Outstanding"
- **A ranked table** of the top 20 on the phase's headline metric —
`rwm_expected_pct` for primary, `attainment_8_score` for secondary, and for
an unphased place page the metric matching whichever phase it holds more of
- **The local average against the England average** — the one number a parent
cannot get from a list
- **Ofsted grade distribution** for the place
- **A map**
- **Links to neighbouring places** and to the parent authority
- **An FAQ block**, feeding `FAQPage` structured data
- **A link to every school page in scope** — this is what finally de-orphans
the 23,000 school pages the original spec identified as near-orphans
## Thin-page controls
Three, and they are the difference between a location layer and index bloat:
1. **Five schools with current data minimum.** Below it, 301 to the parent
authority. This drops 907 towns and 305 outcodes.
2. **Per-phase thresholds.** A town with 30 primaries and 2 secondaries
publishes a primary page and no secondary page.
3. **No page without a local average.** If a place has too few schools with
results to compute one, it has nothing to say that a list does not, and it
falls back to the authority.
## Sitemap
Two new children in the existing index: `/sitemaps/places-{n}.xml` and
`/sitemaps/outcodes-{n}.xml`. Per-family children are why the index was built
in W1 — Search Console reports coverage per submitted sitemap, so indexation
of the location layer is measurable separately from the school pages.
## Testing
Per `CLAUDE.md`, user-facing behaviour extends `e2e/tests/journeys.spec.ts` in
the same PR.
**Unit (registry, no DB):** threshold enforcement; a sub-threshold town
resolves to its authority; a locality slug colliding with a town fails the
build; Bedford's town and authority pages hold different URN sets; per-phase
thresholds.
**Backend:** `/api/places` shape; `/api/places/{slug}` aggregate correctness
against a fixture; unknown slug 404s.
**e2e:** a known town, authority, locality and outcode page each render with
the expected count; a below-threshold town 301s; every place page declares a
canonical and appears in the sitemap; `/schools/bedford` and
`/schools/authority/bedford` both resolve and state which set they cover.
## Risks
**Index bloat** is the failure mode of every programmatic SEO programme. The
three controls above are the answer, and the per-family sitemap is how we find
out early if they were not enough.
**Helpful-content exposure.** Templated location pages are exactly what
Google's stance targets. The mitigation is that every page carries real
computed local data — counts, distributions, local-versus-national comparison
— rather than a name dropped into boilerplate. If indexation of the places
sitemap stalls below roughly half, that is the signal to stop and rethink
rather than to add more pages.
**Build cost.** ~4,000 additional ISR routes on top of 23,000 school pages.
The env-flag gate on `generateStaticParams` keeps CI viable.
**Curation drift.** The locality seed is hand-maintained and will go stale as
places change. It is small and reviewable, and a dbt test asserts every seed
outcode matches at least one school so a typo fails the pipeline rather than
publishing an empty page.
## Out of scope
Catchment-area estimation. It is a strong driver for this cluster and
`fact_admissions` carries the distances, but it is a modelling problem with
real accuracy risk and deserves its own design.
@@ -0,0 +1,294 @@
# Feature Flags — Design
**Date:** 2026-08-23
**Status:** approved for planning
**First consumer:** the last-distance-offered feature (`admission_distance`)
## Goal
Let work merge to `main` and deploy to production without becoming visible,
so that releasing a feature stops being the same event as deploying it.
The site has no way to do this today. A feature is either on `main` and live,
or it is on a branch. That forces long-lived branches for anything not ready,
and it makes every promotion to production an all-or-nothing decision about
everything queued behind it.
This is a **ship-dark** capability, not a kill switch. Flags are expected to
flip on the order of once a month, by a person, deliberately. Nothing here is
designed for flipping something off in seconds under pressure, and nothing
here does percentage rollouts, user targeting or A/B tests — the site has no
user identity to target.
## Decision: Unleash
Flag state is held in a self-hosted [Unleash](https://www.getunleash.io/)
instance (Apache-2.0), not in the repository.
A lighter option was considered and rejected by the project owner: a typed
registry in each runtime with environment-variable overrides set in the
Portainer stack files, which would have needed no new container and kept flag
state in git. The argument for Unleash is that it provides a UI and an audit
log without a deploy, and that flags are expected to become an ongoing
operational tool rather than an occasional one.
Two consequences follow from choosing a service, and this design exists mostly
to handle them:
1. **Flag state lives outside the repository.** `main` is no longer the whole
truth about what is switched on. The registry in §2 exists to bound that.
2. **A flag can change without a deploy**, so nothing else clears the caches
that a deploy would have cleared. §4 establishes how long a flip takes to
become visible, and why that is short enough to need no extra mechanism.
Also considered: Flagsmith (heavier — Django, Postgres and Redis), GrowthBook
(requires MongoDB), and Flipt v2 (the closest conceptual fit, git-native, but
now under the Fair Core Licence — source-available, not OSI open source).
## 1. Topology
A third Portainer stack, `docker-compose.portainer.unleash.yml`, holding
`unleashorg/unleash-server` and its own PostgreSQL 16. It is on the macvlan so
both application stacks can reach it, and it belongs to neither of them — a
staging redeploy must not be able to disturb production's flag state, and vice
versa.
One instance serves both environments. Open-source Unleash ships with
`development` and `production` environments and environment-scoped client
tokens, so the same flag holds independent state in each: staging's FastAPI
carries a `development` token, production's carries a `production` one.
That property is what makes ship-dark testable. A feature can be **on in
staging and off in production** for as long as it takes, which means the `e2e/`
journeys exercise it against staging while production stays unchanged.
## 2. The registry
Unleash supplies flag *state* and the toggle UI. It does not supply the list of
flags. `backend/flags.py` declares every flag the code knows about:
```python
@dataclass(frozen=True)
class Flag:
name: str # identical in the registry, in Unleash, and in JSON
description: str # one line: what turning this on reveals
added: date # for the staleness test in §8
```
**Every flag defaults to `False`.** There is no per-flag default field, because
a flag that defaults on is not a ship-dark flag — it is a kill switch, and this
design does not offer one. A single unconditional default also means the
fallback path has no branching to get wrong.
Three reasons the registry is not optional:
- The Unleash SDK evaluates an unknown flag to `False`. Without a registry that
is an *undeclared* false — indistinguishable from a typo in a flag name.
- `/api/flags` needs a key set to return when Unleash is unreachable. It cannot
enumerate flags it has never heard of.
- A flag present in the Unleash UI but absent from the registry is orphaned,
and should be visibly so rather than quietly authoritative.
**Naming.** One string, used unchanged as the registry key, the Unleash flag
name, and the JSON key in `/api/flags`. It is snake_case, matching the API's
existing convention (`admission_distance`, `rwm_expected_pct`) and the mirrored
types in `nextjs-app/lib/types.ts`. No case transformation anywhere, so there
is no mapping layer to get wrong.
## 3. Read paths
### Backend
`backend/flags.py` wraps `UnleashClient` behind `is_enabled(name: str) -> bool`.
Fail-closed is the default rather than something added: the Python SDK
evaluates every flag to `False` until it has synchronised with the server. An
unfinished feature therefore stays hidden when Unleash is unreachable, which is
the correct direction for ship-dark.
The SDK's fcache directory is mounted on a named volume so a container restart
during an Unleash outage keeps last-known state rather than reverting a
released feature to dark. The registry default remains `False`, so the worst
case is a feature disappearing, never one appearing.
### Frontend
`nextjs-app/lib/flags.ts` exposes `getFlags(): Promise<Flags>`, a single
server-side fetch of `/api/flags` returning a typed record. Server components
only — no flag value reaches the browser bundle, and `package.json` gains no
Unleash dependency. The Unleash client library stays entirely inside the
service that already owns every other piece of data the frontend renders.
The cost, named plainly: a purely front-end flag must still be declared in a
Python file. It is a flat data edit rather than programming, and the return is
one list, so nobody has to ask which service knows about a given flag.
### `/api/flags` must not be publicly reachable
`nextjs-app/app/api/[...path]/route.ts` proxies **everything** under `/api/` to
FastAPI. Left alone, `https://www.schoolcompare.co.uk/api/flags` would return
`{"admission_distance": false, ...}` — publishing the name and state of every
unreleased feature, which defeats the purpose of shipping dark.
The proxy therefore gains a denylist, and `flags` is on it: a request for a
denied path returns 404 rather than being forwarded. Next's own `getFlags()` is
unaffected because it calls `FASTAPI_URL` directly across the Docker network
and never transits the public proxy.
This is a general hole rather than a flags-specific one — the proxy will
forward any future internal endpoint too — so the denylist is written as a
named constant with a comment saying what belongs on it.
## 4. Propagation
**Time-based revalidation is sufficient. There is no webhook.**
An earlier draft of this section specified two Unleash webhooks and a
`revalidateTag('flags')` purge, on the premise that pages cache for seven days.
That premise was wrong, and checking it removed the most complex part of the
design.
Next uses the **lowest** `revalidate` among a route's fetches to set the
revalidation frequency of the whole route — the segment-level
`export const revalidate` does not override a lower value inside it. Measured
against this codebase:
| Page family | Segment | Lowest fetch | Effective |
|---|---|---|---|
| `/school/[slug]` | 604800 | `fetchSchoolDetails` at 300 | **5 minutes** |
| `/schools/*` | 604800 | `fetchNationalAverages` at 3600 | **1 hour** |
The Unleash SDK polls every 15 seconds, so a flip reaches school pages within
about five minutes and place pages within the hour, unaided. Flags flip
monthly, by hand, deliberately. That is fast enough.
What this removes: two webhook integrations, a `/api/revalidate-flags` route, a
shared-secret-in-a-query-string scheme, an idempotency requirement against
duplicate and out-of-order delivery, and a rule that every fetch in
`nextjs-app/lib/` carry a cache tag. None of it has to be built, maintained, or
kept correct as new fetches are added.
**If instant flips are ever wanted**, the webhook is the way to add them, and it
is purely additive — nothing in this design has to change first.
### Two constraints this leaves behind
**Never flag content on a `force-static` page.** `app/admissions/page.tsx`
declares `export const dynamic = 'force-static'`, so it is baked at build time
and never revalidates. A flag gating anything on such a page would not take
effect until the next deploy, silently. If a flag ever needs to reach one, that
page must first move to ISR.
**A route-family flag still needs the sitemap rebuilt.** The sitemap is held in
memory and rebuilt only at startup or via `POST /api/admin/regenerate-sitemap`.
No flag in scope touches the sitemap (§6), so this is deferred with the route
case rather than solved now — but a route flag must not ship without it, or the
sitemap will advertise URLs that `notFound()`.
## 5. What "off" means, per surface
| Surface | Off |
|---|---|
| Route | `notFound()`, **and** absent from the sitemap, **and** absent from nav |
| UI element | Not rendered; surrounding page byte-identical to today |
| API field | Key **absent**, not `null` |
| API endpoint | 404, not 403 |
The three parts of the route rule move together or not at all. Submitting URLs
to Google that return 404 is the bug fixed in PR #124, and a flag is a new way
to reintroduce it.
An API field is withheld **at the source**, never rendered-but-hidden. The
precedent is already set in this codebase by commit `c9a1892`: `/api/schools/`
is public and unauthenticated, so leaving a withheld field in the payload hands
the record to anyone who opens the network tab.
## 6. First consumer: `admission_distance`
The last-distance-offered feature is merged to `main` and live on staging.
Production has never received it: `/api/schools/100010` on production carries
no `admission_distance` key, and no Distance section renders.
It needs **exactly one gate** — `backend/app.py:809`, where the field is
attached to the school payload:
```python
"admission_distance": (
supplementary.get("admission_distance")
if flags.is_enabled("admission_distance") else None
),
```
The frontend follows with no change. `DistanceSection` already returns `null`
when `admission_distance?.distance_m == null`, and `PrimarySchoolSections`
already conditions the admissions block on `(admissions || admissionDistance)`.
The off-state is the commonest state on the site — only 57 local authorities
publish cut-off distances at all — so it is well covered by construction.
The flag does not touch the sitemap: school pages exist either way.
Intended lifecycle: default off, so production receives the code dark on the
next promotion; on in the `development` environment so staging keeps testing
it; flipped on in `production` when the owner chooses.
**This flag exercises two of the three surfaces** in §5 — API field and UI
element. No route case ships with it. The route rule is specified but unproven
until a route-shaped flag exists, and should be treated as such.
## 7. Testing
**Backend unit.** The registry is well-formed; an unknown flag evaluates
`False`; `/api/flags` returns every declared flag with its default when the
SDK is unreachable; `admission_distance` is absent from the school payload when
the flag is off and present when on.
**Frontend unit.** `getFlags()` returns declared defaults when `/api/flags`
fails, rather than throwing and taking the page with it.
**E2E.** Journeys read `/api/flags` and gate flag-dependent assertions on it,
matching the `test.skip` shape the suite already uses.
One trap to avoid, worth stating because the existing distance journeys walk
straight into it: they already skip when no school has a published figure, so
with the flag off they would skip silently and the suite would go green. The
gate must be explicit — **if `/api/flags` reports `admission_distance` on, then
a school with a cut-off must be found**, converting a silent skip into a real
assertion.
## 8. Lifecycle
A flag is temporary scaffolding, and the failure mode of every flag system is
accumulation.
The registry records the date each flag was added, and a backend test fails any
flag older than **90 days**. Removing a flag means deleting the registry entry,
the branches that read it, and the flag in the Unleash UI.
Unleash SDK usage metrics stay enabled, so the UI shows which flags are still
being evaluated — the evidence needed to retire one safely.
## 9. Risks
**Production gains a homelab dependency.** If Unleash is unreachable when a
production container cold-starts with an empty cache, every flag evaluates
`False` and any feature currently switched on disappears. The fcache volume
covers restarts; the 90-day lifecycle rule bounds how long any feature is
exposed to this. It is a real regression risk and the reason flags must be
retired rather than left on indefinitely.
**Flag state is not in git.** `main` no longer tells you what production is
showing. The registry lists what *can* be flagged; only the Unleash UI says
what *is*. This is inherent to the choice of a service.
**A large promotion backlog exists.** Production is running the
pre-SEO-programme build — no place pages, and a sitemap still declaring the
apex host. The first promotion after this work ships that entire backlog. The
flag isolates the distance feature from it and nothing else.
## Out of scope
- Percentage rollouts, user targeting, A/B testing, and Unleash strategies
beyond simple on/off. Flags are booleans.
- Pipeline and dbt flags. Airflow and dbt are not flag consumers.
- Client-side flag evaluation. Flags are server-side only.
- Automatic flag removal. The staleness test reports; a person deletes.
@@ -0,0 +1,294 @@
# School Autosuggest — Design
**Date:** 2026-08-26
**Status:** approved for planning
**Depends on:** the feature-flag layer (PR #125, merged)
## Goal
Suggest schools by name as someone types in the site's main search box, so a
parent who knows the school they want reaches it in one step instead of
searching, scanning a result list, and clicking.
Scope is **schools only**. Places and postcodes were considered and excluded —
see *Out of scope*.
## The finding that shapes everything
The site's rate limiter does not do what it looks like it does.
`limiter = Limiter(key_func=get_remote_address)` with `60/minute` reads
`request.client.host`. In staging and production the backend has no published
ports and sits on the internal `backend` network, so its only caller is the
Next proxy — and `request.client.host` is therefore **the Next container**, for
every browser user on the site.
Measured against staging: 70 concurrent requests to `/api/schools` returned
**60 × 200 and 10 × 429**. One machine consumed the whole site's budget for
that minute.
Autosuggest is the worst possible feature to build on that. One person typing
"st marys primary" produces six to eight debounced requests; **eight concurrent
searchers would 429 the site.** The compare modal's search-as-you-type already
shares this bucket, so the exposure exists today — autosuggest makes it
certain.
Fixing the keying is therefore part of this work, not a follow-up.
## 1. Rate-limit keying
Both environments sit behind Cloudflare (`server: cloudflare`, `cf-ray` present
on staging and production). Cloudflare sets `CF-Connecting-IP` on every request
to the origin and **overwrites any client-supplied value**, which makes it
trustworthy in a way a parsed `X-Forwarded-For` chain is not.
```python
def client_key(request: Request) -> str:
"""Rate-limit bucket: the real caller, not the proxy in front of them."""
cf = request.headers.get("cf-connecting-ip")
if cf:
return cf.strip()
xff = request.headers.get("x-forwarded-for")
if xff:
return xff.split(",")[0].strip()
return get_remote_address(request)
```
`nextjs-app/app/api/[...path]/route.ts` already forwards every inbound header
except `host` and `connection`, so `CF-Connecting-IP` reaches the backend with
no proxy change.
**This header is trustworthy only for traffic that actually passed through
Cloudflare, and nothing in the application can verify that it did.** An earlier
draft of this section claimed Cloudflare "replaces the header, so a browser
cannot forge it", and that only the `X-Forwarded-For` fallback was forgeable.
That was wrong. Cloudflare does overwrite the header *on requests it handles* —
but a caller reaching the origin directly sets whatever it likes, and this
process cannot distinguish an edge-set header from an attacker-set one. Both
headers are equally forgeable in that scenario.
The consequence is sharper than a weakened defence. An attacker rotating
`CF-Connecting-IP` per request mints a fresh rate-limit bucket every time and
evades per-client limits entirely — including on the DataFrame-heavy
`/api/schools`. Against abuse that is *worse* than the shared bucket it
replaced, which at least capped everyone at 60/minute together.
Two mitigations, and they are not interchangeable:
1. **The real fix is at Cloudflare** — Authenticated Origin Pulls, or an origin
firewall that refuses connections not from Cloudflare's ranges. Only the
edge can vouch for its own header. This is infrastructure work and is not
part of this change; it is the thing that makes the header mean anything.
2. **The ceiling in §1.1 bounds what evading the keying can achieve** while
that remains open. It does not make the header trustworthy — it makes
trusting it survivable.
The backend being unreachable from outside the Docker network is a real second
layer, but it depends on the ingress path in front of the frontend, which this
design does not control and should not assume.
### 1.1 The ceiling, which is back
The shared bucket was acting as an accidental global throttle on a
single-process uvicorn backend that filters a 25,000-row DataFrame in-process.
Correct per-user keying removes it: the origin becomes reachable at 60/min *per
user* rather than 60/min in total, and — per above — at an unbounded rate by
anyone willing to rotate a header.
An earlier draft dropped the in-app ceiling, arguing it belonged at Cloudflare.
That argument assumed the keying was sound. It is not, so the ceiling is
load-bearing rather than redundant, and it ships here:
`GlobalRateLimitMiddleware` counts all `/api/` requests in a fixed 60-second
window against `global_rate_limit_per_minute` (3000), independent of any client
identity, and refuses with a 429 that names capacity rather than the client —
an operator has to be able to tell "one noisy client" from "the origin is
saturated". It is registered last so it is outermost: a ceiling that applies
after the expensive work has run is not a ceiling.
slowapi cannot express this. `default_limits` and `application_limits` are both
evaluated with the same `key_func`, making them per-client rather than global,
and `application_limits` only apply with `SlowAPIMiddleware` installed, which
this app does not use. Hence the explicit middleware — about thirty lines, and
obviously correct, which is what a backstop needs to be.
Requests from `127.0.0.1` are exempt. The container healthcheck runs
`curl http://localhost:80/api/data-info` from inside the container, and
starving it would fail the check, restart the container, and turn a load spike
into an outage loop. The exemption keys on the peer address, never the `Host`
header, which the caller sets.
3000/minute is an estimate, not a measurement, and worth revisiting against
real traffic.
### Per-user limits
Per-user fairness and origin protection are different jobs, and this design now
does both separately: the ceiling above for the origin, and per-route limits
for fairness. Conflating them is what produced the original behaviour, where
one bucket served the whole internet.
The existing 60/minute default is unchanged, and `/api/suggest` gets
120/minute. Both are estimates rather than measurements, and are a starting
point to revisit once the keying is correct enough for real per-user traffic to
be visible — which it was not before, because everyone shared one bucket.
## 2. `GET /api/suggest`
A dedicated endpoint, not a mode of `/api/schools`.
The existing search path calls Typesense for URNs and then filters, ranks and
sorts the full in-memory DataFrame — a pandas pass per keystroke, holding the
GIL and blocking other requests in the same worker. Suggestions need none of
it: `urn`, `school_name`, `phase`, `school_type`, `local_authority`,
`postcode` and `ofsted_rating` are all already in the Typesense document
(`pipeline/scripts/sync_typesense.py`).
```
GET /api/suggest?q=<query>&limit=8
→ 200 {"suggestions": [
{"urn": 100010, "school_name": "Brecknock Primary School",
"local_authority": "Camden", "postcode": "NW1 1AA",
"phase": "Primary", "school_type": "Community school"}
]}
```
- **Under two characters** returns `{"suggestions": []}` with 200. The
keystroke path never returns an error for ordinary input.
- **Typesense unavailable** returns `{"suggestions": []}` with 200. There is
deliberately **no DataFrame fallback**: the substring scan `/api/schools`
falls back to is precisely the cost this endpoint exists to avoid, and a
silent 25,000-row scan per keystroke is worse than no suggestions.
- **`limit` is clamped** to 20. It is a public endpoint.
- **Rate limit `120/minute`** per client, not the default 60. A 200 ms
debounce tops out near 5 requests/second while someone is actively typing,
but averages far below that across a real search; 120 leaves headroom for
bursts while still bounding one client.
- **Local authority is part of the payload, not decoration.** There are many
schools called "St Mary's"; a suggestion list without the authority is
unusable for exactly the queries autosuggest is meant to serve.
### Caching
`CACHE_RULES` gains `("/api/suggest", (60, 3600, 86400))`. Prefix queries
repeat enormously across users and school names change once a year.
The client fetch must **not** use `cache: "no-store"`. The compare modal does,
and copying that pattern would throw away both the browser cache and the ETag
304s the existing `CacheAndETagMiddleware` already provides.
Both environments currently report `cf-cache-status: DYNAMIC` — Cloudflare
ignores the `Cache-Control` the API already sends, because it does not cache
dynamic paths by default. **A Cloudflare Cache Rule for `/api/suggest*` would
let the edge absorb most of this traffic and never reach the origin.** That is
a dashboard change, it is optional, and nothing here depends on it.
## 3. The combobox
This is an ARIA combobox, not a text input with a list underneath.
**Files.** `FilterBar.tsx` is already long. The work splits three ways:
`hooks/useSchoolSuggest.ts` owns fetching, debouncing and cancellation;
`components/SuggestList.tsx` owns rendering and ARIA; `FilterBar.tsx` wires
them to the existing input and form.
**Fetching.** 200 ms debounce; minimum two characters; an `AbortController`
cancels the superseded request on every keystroke. Cancellation is not an
optimisation — without it, a slow response for `"st"` can land after the fast
one for `"st marys"` and replace a correct list with a stale one.
**Suppressed during postcode entry.** The box takes a school name *or* a
postcode, and `isValidPostcode` already distinguishes them. Suggestions do not
appear once the value parses as a postcode.
**Keyboard.** `ArrowDown`/`ArrowUp` move the active option, `Escape` closes and
keeps the typed text, `Tab` closes. `Enter` **with an option active** navigates
to that school's page. `Enter` **with none active** submits the free-text
search exactly as it does today — the existing behaviour is preserved, not
replaced.
**ARIA.** `role="combobox"` with `aria-expanded` and `aria-controls` on the
input, `aria-activedescendant` pointing at the active option, `role="listbox"`
on the list and `role="option"` on each row.
**Both instances get it.** `HomeView` renders `FilterBar` twice — hero and
sticky — from one component, so there is one implementation.
## 4. Behind a flag
Flag `school_autosuggest`, declared in `backend/flags.py`, default off.
This is the most-used control on the site and the first change to it in a
while. `app/page.tsx` is an async server component, so it reads the flag and
threads it to `FilterBar` through `HomeView` — two prop hops, explicit, no
client-side flag read.
Off means the input behaves exactly as it does today: no listener, no fetch, no
markup. Not a rendered-then-hidden dropdown.
The rate-limit keying is **not** flagged. It is a correctness fix that should
apply whether or not autosuggest is on, and flagging it would mean shipping a
known-wrong limiter into production deliberately.
## 5. Analytics
`search_submitted` already carries `via: 'input'`. Selecting a suggestion fires
it with `via: 'suggestion'` plus the chosen `urn`, so the obvious question —
does this actually help, or do people ignore it — has an answer in the data
rather than an opinion.
## 6. Testing
**Backend.** `client_key` prefers `CF-Connecting-IP`, falls back through
`X-Forwarded-For` to the remote address, and two different values get two
different buckets. `/api/suggest` returns matches, returns empty below two
characters, returns empty and 200 when Typesense is unavailable, and clamps
`limit`. That it never touches the DataFrame is asserted by making
`load_school_data` raise and requiring the endpoint to answer anyway.
**Frontend.** The hook debounces, aborts superseded requests, and drops a
late-arriving response for a stale query. The list renders the ARIA
attributes. Keyboard navigation moves the active option; `Enter` on an option
navigates; `Enter` on none submits the search.
**E2E.** With the flag on, typing a known school name shows it and selecting it
lands on that school's page. With the flag off, no combobox markup exists.
Gated on the flag the same way the distance journeys are — read the observable
effect, since `/api/flags` is denied to the public.
## 7. Risks
**Removing the accidental throttle.** Covered in §1. Correct per-user keying
means the origin is reachable at 60/minute *per user* where it was 60/minute
in total, and no in-app global cap replaces it — that job goes to Cloudflare,
which is not done as part of this change. Until it is, a determined caller
with many source addresses can put more load on a single-process origin than
they can today. Against this site's traffic that is a theoretical risk rather
than a live one, but it is a real one and it is the price of the fix.
**Cloudflare bypass — the open one.** If the origin is reachable without
passing through Cloudflare, `CF-Connecting-IP` is attacker-controlled, and
rotating it per request defeats per-client limits on every endpoint. The
ceiling in §1.1 bounds the damage to the origin's total capacity; it does not
restore per-client fairness under attack, and it cannot. Closing this properly
means Authenticated Origin Pulls or an origin firewall restricted to
Cloudflare's published ranges — infrastructure work, outside this change, and
the single most valuable follow-up here.
**Typesense becomes user-visible.** Today a Typesense outage degrades search to
a slow substring match. With autosuggest it also means the dropdown silently
stops appearing. That is the correct failure — quiet, not broken — but it makes
Typesense health worth monitoring in a way it was not before.
## Out of scope
- **Place suggestions.** The 2,646 town, authority and outcode pages are a
strong candidate and would route people onto the pages W2 built, but they
live in the place registry rather than Typesense, so it is a second index and
a ranking rule for comparing two kinds of result. Worth its own change.
- **Postcode completion.** Would put postcodes.io in the keystroke path, with
its own latency and rate limits.
- **The compare modal.** It already has search-as-you-type. Converting it to
this component is a reasonable follow-up, not part of this.
- **Recent or popular searches.** No storage for either, and no evidence yet
that they are wanted.
+481 -2
View File
@@ -1226,6 +1226,27 @@ const CUTOFF_CANDIDATE_URNS = [
101099, 100553, 102574, 100769, // mixed
];
/**
* Whether the last-distance-offered feature is switched on here.
*
* Read from the data rather than from /api/flags, which the public proxy
* denies on purpose — the endpoint names unreleased features. The observable
* effect is the field's presence: the flag is off iff no candidate school
* carries an `admission_distance` key at all.
*
* The distinction that matters: `admission_distance: null` means this school
* has no published cut-off, and the key being ABSENT means cut-offs are not
* being published at all.
*/
async function distanceFeatureIsOn(page: Page): Promise<boolean> {
for (const urn of CUTOFF_CANDIDATE_URNS) {
const res = await page.request.get(`/api/schools/${urn}`);
if (!res.ok()) continue;
if ('admission_distance' in (await res.json())) return true;
}
return false;
}
async function schoolWithCutoff(page: Page) {
for (const urn of CUTOFF_CANDIDATE_URNS) {
const res = await page.request.get(`/api/schools/${urn}`);
@@ -1237,6 +1258,59 @@ async function schoolWithCutoff(page: Page) {
return null;
}
test('when the distance feature is on, a school with a cut-off is findable', async ({ page }) => {
/*
* The gate that stops the other distance journeys passing vacuously.
*
* They all skip when schoolWithCutoff() finds nothing, which is right when
* the feature is off — but it means a feature that is *supposed* to be on
* and is silently broken shows up as a green run full of skips. This test
* fails in that case.
*/
test.skip(!(await distanceFeatureIsOn(page)),
'the admission_distance flag is off in this environment');
expect(await schoolWithCutoff(page),
'the distance feature is on, but no candidate school has a cut-off — '
+ 'the flag is on and the data or the query behind it is broken')
.not.toBeNull();
});
test('with the distance feature off, the section is absent rather than empty', async ({ page }) => {
// Shipping dark means the page renders as it did before the feature existed,
// not as a feature with its content removed.
test.skip(await distanceFeatureIsOn(page),
'the admission_distance flag is on in this environment');
// A school that exists, found rather than hardcoded — a 404 page would
// satisfy the absent-heading assertion without proving anything.
//
// A plain loop, not Array.find: find's predicate is synchronous, so an async
// one returns a Promise, every Promise is truthy, and it would always hand
// back the first URN whether or not that school exists.
let urn: number | null = null;
for (const candidate of CUTOFF_CANDIDATE_URNS) {
if ((await page.request.get(`/api/schools/${candidate}`)).ok()) {
urn = candidate;
break;
}
}
expect(urn, 'no candidate school resolves in this environment').not.toBeNull();
await page.goto(`/school/${urn}`);
await expect(page.locator('h1')).toBeVisible();
await expect(page.getByRole('heading', { name: /How far away are you\?/ }))
.toHaveCount(0);
});
test('/api/flags is not reachable from the public internet', async ({ page }) => {
// It names every unreleased feature and whether it is on. Next reads it
// server-side over the Docker network; the public proxy must deny it.
const res = await page.request.get('/api/flags');
expect(res.status()).toBe(404);
});
test('a published cut-off distance is shown with the year it belongs to', async ({ page }) => {
const found = await schoolWithCutoff(page);
test.skip(found === null, 'no school in the sample has a published cut-off distance yet');
@@ -1649,13 +1723,29 @@ const CANONICAL_ROUTES: Array<[string, string]> = [
['/admissions', 'https://www.schoolcompare.co.uk/admissions'],
];
/**
* Next normalises canonical URLs against `trailingSlash: false`, so the root
* ships as `https://www.schoolcompare.co.uk` with no slash while every other
* route keeps its path. Both forms address the same document, and which one
* Next emits is its business, not something worth pinning a test to.
*
* The first cut hardcoded the slash and failed only on the homepage — the
* same gap as the doubled brand: it asserted the metadata object rather than
* what the page actually renders.
*/
function sameUrl(a: string | null, b: string): boolean {
const strip = (u: string) => u.replace(/\/+$/, '');
return strip(a ?? '') === strip(b);
}
for (const [path, expected] of CANONICAL_ROUTES) {
test(`${path} declares exactly one canonical, on the www host`, async ({ page }) => {
await page.goto(path);
const hrefs = await page.locator('link[rel="canonical"]').evaluateAll(
(els) => els.map((e) => e.getAttribute('href')));
expect(hrefs, `${path} should declare one canonical`).toHaveLength(1);
expect(hrefs[0]).toBe(expected);
expect(sameUrl(hrefs[0], expected),
`${path} canonical was ${hrefs[0]}, expected ${expected}`).toBe(true);
});
}
@@ -1663,7 +1753,8 @@ test('a filtered homepage still canonicalises to the bare root', async ({ page }
await page.goto('/?search=primary&phase=primary&sort=name&page=2');
const href = await page.locator('link[rel="canonical"]').first()
.getAttribute('href');
expect(href).toBe('https://www.schoolcompare.co.uk/');
expect(sameUrl(href, 'https://www.schoolcompare.co.uk/'),
`filtered homepage canonical was ${href}`).toBe(true);
});
test('a school page canonicalises to its own slug on the www host', async ({ page }) => {
@@ -1695,3 +1786,391 @@ test('a bare /compare is indexable, a parameterised one is not', async ({ page }
.getAttribute('href');
expect(canonical).toBe('https://www.schoolcompare.co.uk/compare');
});
/*
* Staging must not be indexable (spec 2026-08-20, W1 hygiene).
*
* These journeys only ever run against staging — deploy.yml passes
* STAGING_BASE_URL, and promote.yml only smoke-polls production without
* Playwright — so asserting the noindex header here is safe.
*/
test('staging answers noindex, and stays crawlable so the noindex is seen', async ({ page }) => {
const res = await page.request.get('/');
expect(res.ok()).toBeTruthy();
const tag = res.headers()['x-robots-tag'];
expect(tag, 'staging must send X-Robots-Tag').toBeTruthy();
expect(tag).toContain('noindex');
// The other half, and the reason this is one test rather than two: a
// Disallow would stop Google fetching the page at all, so it would never
// see the noindex above. The two only work together.
//
// Scoped to the `*` group. The first cut matched `Disallow: /` anywhere in
// the file and tripped over the AI-crawler groups Cloudflare injects —
// ClaudeBot, GPTBot, Amazonbot and friends all carry a blanket disallow,
// deliberately, and none of them is Googlebot.
const robots = await (await page.request.get('/robots.txt')).text();
expect(blocksEverything(robots, '*'),
'the * group must not disallow the whole site, or the noindex is never seen')
.toBe(false);
});
/** True when `agent`'s group in a robots.txt disallows the entire site. */
function blocksEverything(robots: string, agent: string): boolean {
let current: string | null = null;
let blocked = false;
for (const raw of robots.split('\n')) {
const line = raw.split('#')[0].trim();
if (!line) continue;
const [key, ...rest] = line.split(':');
const value = rest.join(':').trim();
const k = key.trim().toLowerCase();
if (k === 'user-agent') current = value;
else if (current === agent && k === 'disallow' && value === '/') blocked = true;
}
return blocked;
}
test('a school page on staging is noindexed too, not just the homepage', async ({ page }) => {
const list = await page.request.get('/api/schools?search=primary&per_page=1');
const [first] = (await list.json()).schools ?? [];
expect(first, 'no school available').toBeTruthy();
const res = await page.request.get(`/school/${first.urn}-x`);
expect(res.headers()['x-robots-tag']).toContain('noindex');
});
/*
* W8 — the C1 pages must ship a description, and it must differentiate.
*
* Baseline was 0.43% CTR at position 6.1 on "compare school performance",
* against 9.16% for the brand query from the same neighbourhood. The SERP is
* owned by the DfE's own service, so a description that paraphrases it earns
* nothing. Google may rewrite a snippet, but it cannot use one we never sent.
*/
test('every C1 page ships a description, and none opens its title with the brand', async ({ page }) => {
for (const path of ['/', '/compare', '/rankings', '/admissions']) {
await page.goto(path);
const desc = await page.locator('meta[name="description"]').first()
.getAttribute('content');
expect(desc, `${path} must ship a description`).toBeTruthy();
expect(desc!.length, `${path} description too short to be worth reading`)
.toBeGreaterThan(100);
const title = await page.title();
expect(title.toLowerCase().startsWith('schoolcompare'),
`${path} spends its most valuable pixels on the brand`).toBe(false);
}
});
test('the homepage snippet names what gov.uk does not publish', async ({ page }) => {
await page.goto('/');
const desc = await page.locator('meta[name="description"]').first()
.getAttribute('content');
// Admissions distance is the one fact the DfE service has no equivalent for.
expect(desc).toMatch(/close you had to live|distance/i);
});
/*
* The location layer (spec 2026-08-21, W2).
*
* Location intent sat at position 49.5 with one click across the whole 16-month
* baseline — the site published no page about a place. These assert the four
* families render, stay in their own namespaces, and reach the sitemap.
*/
async function firstPlaceOfKind(page: Page, kind: string) {
const res = await page.request.get('/api/places');
expect(res.ok()).toBeTruthy();
const { places } = await res.json();
const hit = places.find((p: { kind: string }) => p.kind === kind);
expect(hit, `no ${kind} in the registry`).toBeTruthy();
return hit as { kind: string; slug: string; name: string; count: number };
}
for (const [kind, prefix, article] of [
['town', '/schools/', 'a'],
['authority', '/schools/authority/', 'an'],
['outcode', '/schools/near/', 'an'],
] as const) {
test(`${article} ${kind} page renders with its school count`, async ({ page }) => {
const place = await firstPlaceOfKind(page, kind);
await page.goto(`${prefix}${place.slug}`);
await expect(page.locator('h1')).toContainText(place.name, { ignoreCase: true });
await expect(page.locator('a[href^="/school/"]').first()).toBeVisible();
});
}
test('a town and an authority sharing a name are different pages', async ({ page }) => {
// 67 real collisions, and the authority is the larger set in only 43 — so
// one namespace would have published near-duplicates.
const { places } = await (await page.request.get('/api/places')).json();
const townSlugs = new Set(
places.filter((p: { kind: string }) => p.kind === 'town')
.map((p: { slug: string }) => p.slug));
const clash = places.find((p: { kind: string; slug: string }) =>
p.kind === 'authority' && townSlugs.has(p.slug));
test.skip(!clash, 'no town/authority name collision in this environment');
const townRes = await page.request.get(`/api/places/town/${clash.slug}`);
const laRes = await page.request.get(`/api/places/authority/${clash.slug}`);
expect(townRes.ok() && laRes.ok()).toBeTruthy();
const townUrns = (await townRes.json()).schools.map((s: { urn: number }) => s.urn).sort();
const laUrns = (await laRes.json()).schools.map((s: { urn: number }) => s.urn).sort();
expect(townUrns).not.toEqual(laUrns);
});
test('a place below the threshold has no page', async ({ page }) => {
// Crosby holds one school; publishing it would be a page with nothing to say.
const res = await page.request.get('/api/places/town/crosby');
expect(res.status()).toBe(404);
});
test('place pages declare a canonical and reach the sitemap', async ({ page }) => {
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
const canonical = await page.locator('link[rel="canonical"]').first()
.getAttribute('href');
expect(canonical).toBe(`https://www.schoolcompare.co.uk/schools/${place.slug}`);
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toContain(`/schools/${place.slug}`);
});
test('the place page ships ItemList structured data that parses', async ({ page }) => {
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
const raw = await page.locator('script[type="application/ld+json"]').first()
.textContent();
const parsed = JSON.parse(raw!);
const types = (parsed['@graph'] ?? []).map((n: { '@type': string }) => n['@type']);
expect(types).toContain('ItemList');
expect(types).toContain('BreadcrumbList');
});
test('a place page states the local average against England', async ({ page }) => {
// The one number a list cannot give, and the reason these pages are not
// a name dropped into a template.
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
await expect(page.getByTestId('local-vs-england')).toContainText(/across England/i);
});
test('a place page links its phase variants, and they resolve', async ({ page }) => {
// "primary schools in beccles" is the query shape the baseline showed. The
// first cut submitted only the bare place URL and linked nothing, leaving
// ~950 variant pages reachable by nothing at all.
const res = await page.request.get('/api/places');
const { places } = await res.json();
const town = places.find((p: { kind: string }) => p.kind === 'town');
expect(town).toBeTruthy();
const detail = await (await page.request.get(`/api/places/town/${town.slug}`)).json();
test.skip(!(detail.place.phases ?? []).length, 'no phase clears the threshold here');
await page.goto(`/schools/${town.slug}`);
const phase = detail.place.phases[0];
const link = page.locator(`a[href="/schools/${town.slug}/${phase}"]`).first();
await expect(link).toBeVisible();
await link.click();
await expect(page.locator('h1')).toContainText(new RegExp(`${phase} schools in`, 'i'));
});
test('phase variants are submitted in the places sitemap', async ({ page }) => {
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toMatch(/\/schools\/[a-z0-9-]+\/primary</);
});
test('authority phase variants are submitted, and in their own namespace', async ({ page }) => {
// 302 of these were in the sitemap for weeks and every one 404'd: the spec
// called for the route, the plan built the bare authority page and dropped
// it, and the sitemap — written from the registry — kept submitting them.
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toMatch(/\/schools\/authority\/[a-z0-9-]+\/primary</);
});
test('every place link a place page emits resolves', async ({ page }) => {
/*
* The guard that was missing. Each family built its own links, so a URL
* shape belonging to one namespace was used by all four: an authority page
* offered "Primary schools in Barnet" pointing at /schools/barnet/primary,
* the *town*. For 87 of 151 authorities that 404'd; for the other 64 it
* quietly served a different set of schools under the same name.
*
* Only /schools links are followed. The per-school links are the same
* component the school-page journeys already cover, and there are hundreds
* of them on a page.
*/
for (const kind of ['town', 'authority', 'outcode'] as const) {
const place = await firstPlaceOfKind(page, kind);
const prefix = kind === 'authority' ? '/schools/authority/'
: kind === 'outcode' ? '/schools/near/' : '/schools/';
await page.goto(`${prefix}${place.slug}`);
const hrefs = [...new Set(
await page.locator('a[href^="/schools"]').evaluateAll(
(els) => els.map((e) => e.getAttribute('href')!)))];
expect(hrefs.length, `${kind} page links no other place`).toBeGreaterThan(0);
for (const href of hrefs) {
const res = await page.request.get(href);
expect(res.status(), `${kind} page links ${href}`).toBe(200);
}
}
});
test('an outcode page offers no phase link, because no such page exists', async ({ page }) => {
// Nobody searches "primary schools in SW11", so the spec gives outcodes no
// phase route. The registry computed the variants anyway and the page
// linked them, putting two 404s on each of 1,720 outcode pages.
const place = await firstPlaceOfKind(page, 'outcode');
const detail = await (await page.request.get(
`/api/places/outcode/${place.slug}`)).json();
expect(detail.place.phases).toEqual([]);
await page.goto(`/schools/near/${place.slug}`);
await expect(page.getByRole('navigation', { name: 'By phase' })).toHaveCount(0);
});
test('no page title repeats the brand', async ({ page }) => {
// The root layout appends '| schoolcompare' to a plain-string title. Any
// route whose title already carries the brand must opt out with
// `absolute`, or it ships '... | schoolcompare | schoolcompare' — which is
// how ~2,600 place pages first went out.
const res = await page.request.get('/api/places');
const { places } = await res.json();
const town = places.find((p: { kind: string }) => p.kind === 'town');
for (const path of ['/', '/rankings', '/admissions', `/schools/${town.slug}`]) {
await page.goto(path);
const title = await page.title();
const brands = (title.match(/schoolcompare/gi) ?? []).length;
expect(brands, `${path} repeats the brand: ${title}`).toBeLessThanOrEqual(1);
}
});
test('a place straddling a boundary names every authority it sits in', async ({ page }) => {
// A quarter of outcodes and a third of towns cross an authority boundary —
// SW19 is mostly Merton but partly Wandsworth. Naming only the largest
// asserts something false about the place.
const { places } = await (await page.request.get('/api/places')).json();
const outcode = places.find((p: { kind: string }) => p.kind === 'outcode');
expect(outcode).toBeTruthy();
// Find any place the registry reports as straddling.
let straddling: { kind: string; slug: string } | null = null;
for (const p of places.filter((p: { kind: string }) => p.kind === 'outcode').slice(0, 40)) {
const d = await (await page.request.get(`/api/places/outcode/${p.slug}`)).json();
if ((d.place.authorities ?? []).length > 1) { straddling = p; break; }
}
test.skip(!straddling, 'no straddling outcode found in the sample');
const detail = await (await page.request.get(
`/api/places/outcode/${straddling!.slug}`)).json();
await page.goto(`/schools/near/${straddling!.slug}`);
for (const a of detail.place.authorities) {
if (a.slug) {
await expect(page.locator(`a[href="/schools/authority/${a.slug}"]`).first())
.toBeVisible();
} else {
// No page of its own — City of London and the Isles of Scilly are
// under the threshold. Named, deliberately not linked.
await expect(page.locator('header p')).toContainText(a.name);
await expect(page.getByRole('link', { name: a.name })).toHaveCount(0);
}
}
});
test('a place page lists its schools alphabetically', async ({ page }) => {
// Someone on a place page is usually looking for a school they can name,
// so the order should serve scanning for it. /rankings is where the
// league-table ordering lives.
const { places } = await (await page.request.get('/api/places')).json();
const town = places.find((p: { kind: string; count: number }) =>
p.kind === 'town' && p.count >= 5);
expect(town).toBeTruthy();
await page.goto(`/schools/${town.slug}`);
const names = await page.locator('a[href^="/school/"]').allTextContents();
expect(names.length).toBeGreaterThan(1);
const sorted = [...names].sort((a, b) =>
a.toLowerCase().localeCompare(b.toLowerCase()));
expect(names).toEqual(sorted);
});
test('the rankings page still orders by score, not name', async ({ page }) => {
// Alphabetical is a place-page decision, not a site-wide one.
const res = await page.request.get('/api/rankings?metric=rwm_expected_pct&phase=primary');
expect(res.ok()).toBeTruthy();
const scores = ((await res.json()).rankings ?? [])
.map((r: { rwm_expected_pct: number | null }) => r.rwm_expected_pct)
.filter((v: number | null) => v != null);
expect(scores).toEqual([...scores].sort((a: number, b: number) => b - a));
});
/*
* School autosuggest (spec 2026-08-26).
*/
async function autosuggestIsOn(page: Page): Promise<boolean> {
await page.goto('/');
return (await page.getByRole('combobox').count()) > 0;
}
test('the suggest endpoint answers from Typesense', async ({ page }) => {
// Not flagged — the endpoint is live even while the UI is dark, so it can
// be smoke-tested before the feature is switched on.
const res = await page.request.get('/api/suggest?q=brecknock');
expect(res.ok()).toBeTruthy();
const { suggestions } = await res.json();
expect(Array.isArray(suggestions)).toBeTruthy();
if (suggestions.length) {
// Local authority is what tells two "St Mary's" apart.
expect(suggestions[0]).toHaveProperty('school_name');
expect(suggestions[0]).toHaveProperty('local_authority');
}
});
test('a one-character query is answered, not rejected', async ({ page }) => {
// The keystroke path never errors on ordinary input.
const res = await page.request.get('/api/suggest?q=b');
expect(res.status()).toBe(200);
expect((await res.json()).suggestions).toEqual([]);
});
test('the suggest response is cacheable', async ({ page }) => {
const res = await page.request.get('/api/suggest?q=brecknock');
expect(res.headers()['cache-control'] ?? '').toContain('s-maxage');
});
test('typing a school name suggests it, and choosing it opens that school', async ({ page }) => {
test.skip(!(await autosuggestIsOn(page)),
'the school_autosuggest flag is off in this environment');
// A school certain to exist in any environment with data.
const { schools } = await (await page.request.get('/api/schools?page_size=1')).json();
test.skip(!schools?.length, 'no schools in this environment');
const name = schools[0].school_name as string;
await page.goto('/');
await page.getByRole('combobox').first().fill(name.slice(0, 12));
const option = page.getByRole('option').first();
await expect(option).toBeVisible();
await option.click();
await expect(page).toHaveURL(/\/school\/\d+/);
});
test('with autosuggest off, the search box is a plain input', async ({ page }) => {
test.skip(await autosuggestIsOn(page),
'the school_autosuggest flag is on in this environment');
await page.goto('/');
await expect(page.getByRole('combobox')).toHaveCount(0);
// And the box still works: the existing search must be untouched.
await page.getByPlaceholder(/School name or postcode/i).first().fill('abbey');
await page.getByRole('button', { name: /Search/i }).first().click();
await expect(page).toHaveURL(/search=abbey/);
});
@@ -0,0 +1,31 @@
/**
* The /api/* proxy is public. Anything it forwards is on the internet.
*
* @jest-environment node
*/
// The docblock above is load-bearing. jest.config.js sets jsdom globally, and
// NextRequest/NextResponse need the Web Fetch API globals that only the node
// environment provides — under jsdom this suite fails on import, not on an
// assertion.
import { NextRequest } from 'next/server';
import { GET } from '@/app/api/[...path]/route';
function request(path: string) {
return new NextRequest(`http://localhost:3000/api/${path}`);
}
describe('public API proxy', () => {
it('refuses to forward internal-only paths', async () => {
// /api/flags names every unreleased feature and its state. Forwarding it
// publishes the thing shipping dark exists to keep quiet.
const res = await GET(request('flags'), { params: Promise.resolve({ path: ['flags'] }) });
expect(res.status).toBe(404);
});
it('does not deny a path that merely starts with the same letters', async () => {
// A prefix match would take /api/flagship down with /api/flags.
const res = await GET(
request('flagship'), { params: Promise.resolve({ path: ['flagship'] }) });
expect(res.status).not.toBe(404);
});
});
+77
View File
@@ -51,3 +51,80 @@ describe('/compare indexability', () => {
.toBe('https://www.schoolcompare.co.uk/compare');
});
});
/*
* W8 — snippet copy for the C1 cluster.
*
* The baseline (GSC, 16 months to 2026-08-20) showed these pages ranking on
* page one and converting at a tenth of the normal rate: "compare school
* performance" at position 6.1 with 0.43% CTR, against 9.16% for the brand
* query from the same neighbourhood. The SERP is dominated by the DfE's own
* "Compare school performance" service, so the job of this copy is to say
* what that service does not offer, without losing intent match on the title.
*
* These tests guard the mechanics that make a snippet work — length, intent
* keyword, differentiator, no brand-first — not the exact wording, which
* should stay free to iterate.
*/
// Google truncates titles near 60 characters and descriptions near 155.
const TITLE_MAX = 60;
const DESC_MIN = 110;
const DESC_MAX = 155;
type Meta = { title?: unknown; description?: unknown };
const titleOf = (m: Meta): string => {
const t = m.title as string | { absolute?: string } | undefined;
return typeof t === 'string' ? t : (t?.absolute ?? '');
};
describe('C1 snippet copy', () => {
const pages: Array<[string, Meta, RegExp]> = [
['home', homeMetadata as Meta, /compare schools/i],
['rankings', rankingsMetadata as Meta, /league table/i],
['admissions', admissionsMetadata as Meta, /admission/i],
];
for (const [name, meta, intent] of pages) {
it(`${name}: title carries the search intent and fits the SERP`, () => {
const t = titleOf(meta);
expect(t).toMatch(intent);
expect(t.length).toBeLessThanOrEqual(TITLE_MAX);
});
it(`${name}: title does not open with the brand`, () => {
// The measured 0.43% CTR came from a brand-first title. The most
// valuable pixels go to the thing the searcher typed.
expect(titleOf(meta).toLowerCase().startsWith('schoolcompare')).toBe(false);
});
it(`${name}: description is long enough to be worth reading, short enough to survive`, () => {
const d = meta.description as string;
expect(d.length).toBeGreaterThanOrEqual(DESC_MIN);
expect(d.length).toBeLessThanOrEqual(DESC_MAX);
});
}
it('the homepage description names what gov.uk does not publish', () => {
// Admissions distance is the one fact the DfE service has no equivalent
// for. If it ever leaves this description, the snippet is competing with
// gov.uk on gov.uk's own ground.
expect(homeMetadata.description).toMatch(/close you had to live|distance/i);
});
it('/compare targets the tool phrasing rather than repeating the homepage', () => {
// Two pages chasing one phrase is how a site competes with itself.
return compareMetadata({ searchParams: Promise.resolve({}) }).then((m) => {
expect(m.title).toMatch(/comparison tool/i);
expect(m.title).not.toBe(titleOf(homeMetadata as Meta));
});
});
it('no C1 page claims a school count that will drift', () => {
// The corpus moves with every data refresh; this repo has already shipped
// one copy bug of that kind ("three schools" against MAX_SCHOOLS = 5).
for (const [, meta] of pages) {
expect(meta.description as string).not.toMatch(/\b\d{2},\d{3}\b|\b\d{2},000\b/);
}
});
});
@@ -0,0 +1,42 @@
import { generateMetadata as placeMeta } from '@/app/schools/[place]/page';
jest.mock('@/lib/places', () => ({
...jest.requireActual('@/lib/places'),
fetchPlace: jest.fn(async (kind: string, slug: string) =>
slug === 'atlantis' ? null : ({
place: { kind, slug, name: 'Brentwood', count: 29,
parent_authority: 'Essex' },
schools: [], averages: { rwm_expected_pct: 63, attainment_8_score: null },
})),
fetchPlaces: jest.fn(async () => []),
}));
describe('place page metadata', () => {
it('titles the page the way the place is searched', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'brentwood' }) });
expect((m.title as { absolute: string }).absolute).toMatch(/schools in brentwood/i);
});
it('canonicalises to its own path on the www host', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'brentwood' }) });
expect(m.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/schools/brentwood');
});
it('opts out of the layout template, which would double the brand', () => {
// The root layout appends '| schoolcompare' to a plain string title, and
// these titles already carry it — every place page shipped reading
// '... | schoolcompare | schoolcompare' until this was made absolute.
return placeMeta({ params: Promise.resolve({ place: 'brentwood' }) })
.then((m) => {
expect(typeof m.title).toBe('object');
expect((m.title as { absolute: string }).absolute)
.not.toMatch(/schoolcompare.*schoolcompare/);
});
});
it('an unknown place gets a not-found title rather than inventing one', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'atlantis' }) });
expect(m.title).toMatch(/not found/i);
});
});
@@ -0,0 +1,76 @@
import { render, screen, fireEvent, waitFor } from '@testing-library/react';
import userEvent from '@testing-library/user-event';
import { FilterBar } from '@/components/FilterBar';
const push = jest.fn();
jest.mock('next/navigation', () => ({
useRouter: () => ({ push, replace: jest.fn(), prefetch: jest.fn() }),
usePathname: () => '/',
useSearchParams: () => new URLSearchParams(),
}));
const FILTERS = {
local_authorities: [], school_types: [], years: [], phases: [],
genders: [], admissions_policies: [],
};
const realFetch = global.fetch;
beforeEach(() => {
global.fetch = jest.fn(async () => ({
ok: true,
json: async () => ({ suggestions: [{
urn: 100010, school_name: 'Brecknock Primary School',
local_authority: 'Camden', postcode: 'NW1 1AA',
phase: 'Primary', school_type: 'Community school' }] }),
})) as unknown as typeof fetch;
push.mockClear();
});
afterEach(() => { global.fetch = realFetch; });
describe('FilterBar autosuggest', () => {
it('is a combobox only when the flag is on', () => {
const { rerender } = render(<FilterBar filters={FILTERS} autosuggest={false} />);
expect(screen.queryByRole('combobox')).not.toBeInTheDocument();
rerender(<FilterBar filters={FILTERS} autosuggest />);
expect(screen.getByRole('combobox')).toBeInTheDocument();
});
it('makes no request while the flag is off', async () => {
// Off means off: no listener, no fetch, no markup.
render(<FilterBar filters={FILTERS} autosuggest={false} />);
await userEvent.type(screen.getByPlaceholderText(/School name or postcode/i),
'brecknock');
expect(global.fetch).not.toHaveBeenCalled();
});
it('shows suggestions and navigates when one is chosen', async () => {
render(<FilterBar filters={FILTERS} autosuggest />);
await userEvent.type(screen.getByRole('combobox'), 'brecknock');
const option = await screen.findByRole('option', { name: /Brecknock/ });
await userEvent.click(option);
expect(push).toHaveBeenCalledWith(
expect.stringContaining('/school/100010'));
});
it('suppresses suggestions once the value is a postcode', async () => {
// The box takes a name OR a postcode; suggestions must get out of the way.
//
// fireEvent.change, not userEvent.type: typing sets "N", "NW", "NW1"... and
// "NW1" is not a postcode, so a request for it is correct behaviour. Only
// the settled value is the assertion, so set it in one go.
render(<FilterBar filters={FILTERS} autosuggest />);
fireEvent.change(screen.getByRole('combobox'), { target: { value: 'NW1 1AA' } });
await new Promise((r) => setTimeout(r, 300)); // past the 200ms debounce
expect(global.fetch).not.toHaveBeenCalled();
});
it('Enter with no active option still submits the free-text search', async () => {
// The existing behaviour is preserved, not replaced.
render(<FilterBar filters={FILTERS} autosuggest />);
const input = screen.getByRole('combobox');
await userEvent.type(input, 'brecknock{Enter}');
// updateURL pushes inside startTransition, so the call is not synchronous.
await waitFor(() => expect(push).toHaveBeenCalledWith(
expect.stringContaining('search=brecknock')));
});
});
@@ -0,0 +1,348 @@
import { render, screen } from '@testing-library/react';
import { PlaceView } from '@/components/places/PlaceView';
import type { PlaceDetail } from '@/lib/places';
const detail: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 29,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', rwm_expected_pct: 82,
ofsted_grade: 1, phase: 'Primary' } as never,
{ urn: 2, school_name: 'Beta Primary', rwm_expected_pct: 44,
ofsted_grade: 3, phase: 'Primary' } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
describe('PlaceView', () => {
it('leads with an H1 that matches how the place is searched', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByRole('heading', { level: 1 }))
.toHaveTextContent(/primary schools in brentwood/i);
});
it('states the count so the page says something before the table', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByText(/29 schools/i)).toBeInTheDocument();
});
it('compares the local average against England, which a list cannot', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByTestId('local-vs-england')).toHaveTextContent('63');
expect(screen.getByTestId('local-vs-england')).toHaveTextContent('61');
});
it('links every school in scope, which is what de-orphans them', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getAllByRole('link', { name: /Primary$/ })).toHaveLength(2);
});
it('links to the parent authority so the place sits in a hierarchy', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByRole('link', { name: /Essex/i }))
.toHaveAttribute('href', '/schools/authority/essex');
});
it('shows the Ofsted distribution, not just a count of Outstanding', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByTestId('ofsted-distribution')).toBeInTheDocument();
});
it('links to neighbouring places so the page is not a dead end', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[{ kind: 'town', slug: 'romford', name: 'Romford', count: 40 }]} />);
expect(screen.getByRole('link', { name: /Romford/ }))
.toHaveAttribute('href', '/schools/romford');
});
it('says nothing about an average it does not have', () => {
render(<PlaceView detail={{ ...detail, averages:
{ rwm_expected_pct: null, attainment_8_score: null } }}
phase="primary" englandAverage={61} neighbours={[]} />);
expect(screen.queryByTestId('local-vs-england')).not.toBeInTheDocument();
});
});
describe('PlaceView structured data', () => {
function jsonLd() {
const { container } = render(<PlaceView detail={detail} phase="primary"
englandAverage={61} neighbours={[]} />);
const el = container.querySelector('script[type="application/ld+json"]');
return JSON.parse(el!.textContent!);
}
it('declares the page as a ranked list, not prose', () => {
const types = jsonLd()['@graph'].map((n: { '@type': string }) => n['@type']);
expect(types).toContain('ItemList');
expect(types).toContain('BreadcrumbList');
});
it('gives every listed school an absolute URL on the canonical host', () => {
const list = jsonLd()['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListElement).toHaveLength(2);
for (const item of list.itemListElement) {
expect(item.url).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\/school\//);
}
});
});
describe('PlaceView phase variants', () => {
it('links the phase variants that exist', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /Primary schools in Brentwood/i }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('links no variant for a phase below its own threshold', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.queryByRole('link', { name: /Secondary schools in Brentwood/i }))
.not.toBeInTheDocument();
});
it('does not link sideways from a variant page to itself', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.queryByRole('link', { name: /Primary schools in Brentwood/i }))
.not.toBeInTheDocument();
});
});
describe('PlaceView presentation', () => {
// /schools/brentwood shipped with 8 of 27 rows blank: an unphased page shows
// one primary-only measure for a list that also holds secondaries.
const mixed: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 4,
parent_authority: 'Essex', phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta High', phase: 'Secondary',
rwm_expected_pct: null, attainment_8_score: 47 } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: 45 },
};
it('gives each phase its own table rather than one column of blanks', () => {
render(<PlaceView detail={mixed} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('heading', { name: /^Primary schools/ })).toBeInTheDocument();
expect(screen.getByRole('heading', { name: /^Secondary schools/ })).toBeInTheDocument();
expect(screen.getByText('82%')).toBeInTheDocument();
expect(screen.getByText('47')).toBeInTheDocument();
});
it('names the measure in plain words, not jargon', () => {
// The first cut said "RWM expected", which appears nowhere else on the site.
render(<PlaceView detail={mixed} englandAverage={61} neighbours={[]} />);
expect(screen.getByText('Reading, writing & maths')).toBeInTheDocument();
expect(screen.getByText('Attainment 8')).toBeInTheDocument();
expect(screen.queryByText(/RWM expected/i)).not.toBeInTheDocument();
});
it('says a missing result is unpublished rather than showing a bare dash', () => {
const noResult: PlaceDetail = {
...mixed,
schools: [{ urn: 3, school_name: 'New Primary', phase: 'Primary',
rwm_expected_pct: null, attainment_8_score: null } as never],
};
render(<PlaceView detail={noResult} englandAverage={61} neighbours={[]} />);
expect(screen.getByText('Not published')).toBeInTheDocument();
});
it('styles school links to the site convention rather than browser default', () => {
const { container } = render(<PlaceView detail={mixed} englandAverage={61}
neighbours={[]} />);
const link = container.querySelector('a[href^="/school/"]');
expect(link?.className).toBeTruthy();
});
it('a phased page shows one table and no phase headings', () => {
render(<PlaceView detail={mixed} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.queryByRole('heading', { name: /^Secondary schools/ }))
.not.toBeInTheDocument();
});
});
describe('PlaceView table alignment', () => {
const aligned: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('aligns the measure heading and its values with the same class', () => {
// They were aligned by two different selectors whose specificity did not
// match: `.table th:last-child` (0,2,1) won and went right, while `.num`
// (0,1,0) lost to `.table td` (0,1,1) and stayed left. Sharing one class
// is what makes them impossible to drift apart.
const { container } = render(<PlaceView detail={aligned} englandAverage={61}
neighbours={[]} />);
const th = container.querySelectorAll('th')[1];
const td = container.querySelectorAll('tbody td')[1];
expect(th.className).toBeTruthy();
expect(td.className).toBe(th.className);
});
it('leaves the school-name column unclassed so it takes the spare width', () => {
const { container } = render(<PlaceView detail={aligned} englandAverage={61}
neighbours={[]} />);
expect(container.querySelectorAll('th')[0].className).toBe('');
});
});
describe('PlaceView authorities', () => {
const straddling: PlaceDetail = {
place: { kind: 'outcode', slug: 'sw19', name: 'SW19', count: 33,
parent_authority: 'Merton', phases: ['primary'],
authorities: [
{ name: 'Merton', slug: 'merton', count: 26 },
{ name: 'Wandsworth', slug: 'wandsworth', count: 7 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names every authority the place straddles, not just the largest', () => {
// SW19 is mostly Merton but partly Wandsworth. Naming one asserts
// something false about a quarter of outcodes.
render(<PlaceView detail={straddling} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' }))
.toHaveAttribute('href', '/schools/authority/merton');
expect(screen.getByRole('link', { name: 'Wandsworth' }))
.toHaveAttribute('href', '/schools/authority/wandsworth');
});
it('joins them readably rather than as a bare list', () => {
// Asserted on the summary line's whole text: a loose /and/ matcher also
// hits "Wandsworth".
const { container } = render(<PlaceView detail={straddling}
englandAverage={61} neighbours={[]} />);
const summary = container.querySelector('header p');
expect(summary?.textContent).toContain('Merton and Wandsworth');
});
it('falls back to the single parent when the field is absent', () => {
// A cached API response predating the authorities field must not blank
// the line entirely.
const legacy = { ...straddling,
place: { ...straddling.place, authorities: undefined } };
render(<PlaceView detail={legacy} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' })).toBeInTheDocument();
});
});
describe('PlaceView list ordering', () => {
const detail3: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 40, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta Primary', phase: 'Primary',
rwm_expected_pct: 90, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 65, attainment_8_score: null },
};
it('renders schools in the order the API sent them, not by score', () => {
// The API sorts alphabetically now; the component must not re-sort.
render(<PlaceView detail={detail3} englandAverage={61} neighbours={[]} />);
const links = screen.getAllByRole('link', { name: /Primary$/ });
expect(links.map((l) => l.textContent))
.toEqual(['Alpha Primary', 'Beta Primary']);
});
it('declares the list as ascending rather than implying a ranking', () => {
// An ItemList carrying `position` reads as a ranking unless it says
// otherwise, and the table is A-Z.
const { container } = render(<PlaceView detail={detail3} englandAverage={61}
neighbours={[]} />);
const ld = JSON.parse(
container.querySelector('script[type="application/ld+json"]')!.textContent!);
const list = ld['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListOrder).toBe('https://schema.org/ItemListOrderAscending');
});
});
describe('PlaceView phase links', () => {
const authority: PlaceDetail = {
place: { kind: 'authority', slug: 'barnet', name: 'Barnet', count: 156,
parent_authority: null, phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('keeps an authority phase link in the authority namespace', () => {
// The link was built as `/schools/${slug}/${phase}` for every kind, so an
// authority page pointed into the town namespace. For 87 of 151
// authorities that 404'd; for the other 64 it silently landed on the town
// page of the same name — a different set of schools, and exactly the
// duplicate the two namespaces exist to prevent. Barnet is one of the 64.
render(<PlaceView detail={authority} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /^Primary schools in Barnet$/ }))
.toHaveAttribute('href', '/schools/authority/barnet/primary');
expect(screen.getByRole('link', { name: /^Secondary schools in Barnet$/ }))
.toHaveAttribute('href', '/schools/authority/barnet/secondary');
});
it('still uses the bare namespace for a town', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /^Primary schools in Brentwood$/ }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('offers no phase link when the place publishes none', () => {
// Outcodes are the case: no phase route exists for them, so the registry
// reports no phases and the nav does not render.
const outcode = { ...detail,
place: { ...detail.place, kind: 'outcode', slug: 'cm13', name: 'CM13',
phases: [] } };
render(<PlaceView detail={outcode} englandAverage={61} neighbours={[]} />);
expect(screen.queryByRole('navigation', { name: 'By phase' }))
.not.toBeInTheDocument();
});
});
describe('PlaceView unlinkable authorities', () => {
const withUnpublished: PlaceDetail = {
place: { kind: 'outcode', slug: 'tr21', name: 'TR21', count: 8,
parent_authority: 'Cornwall', phases: [],
authorities: [
{ name: 'Cornwall', slug: 'cornwall', count: 6 },
{ name: 'Isles Of Scilly', slug: null, count: 2 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names an authority with no page without linking it', () => {
// City of London and the Isles of Scilly hold fewer schools than a page
// needs. Saying where the place is stays right; linking there would 404.
const { container } = render(<PlaceView detail={withUnpublished}
englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Cornwall' })).toBeInTheDocument();
expect(screen.queryByRole('link', { name: 'Isles Of Scilly' }))
.not.toBeInTheDocument();
expect(container.querySelector('header p')?.textContent)
.toContain('Isles Of Scilly');
});
});
@@ -0,0 +1,50 @@
import { render, screen } from '@testing-library/react';
import { SuggestList, suggestOptionId } from '@/components/SuggestList';
const ROWS = [
{ urn: 1, school_name: "St Mary's Primary", local_authority: 'Camden',
postcode: 'NW1 1AA', phase: 'Primary', school_type: 'Voluntary aided school' },
{ urn: 2, school_name: "St Mary's Primary", local_authority: 'Barnet',
postcode: 'EN5 2AA', phase: 'Primary', school_type: 'Community school' },
];
describe('SuggestList', () => {
it('is a listbox of options', () => {
render(<SuggestList id="s" suggestions={ROWS} activeIndex={-1}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getByRole('listbox')).toBeInTheDocument();
expect(screen.getAllByRole('option')).toHaveLength(2);
});
it('shows the local authority, which is what tells two schools apart', () => {
// Both rows are "St Mary's Primary". Without the authority the list is
// unusable for exactly the query autosuggest exists to serve.
render(<SuggestList id="s" suggestions={ROWS} activeIndex={-1}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getByText('Camden')).toBeInTheDocument();
expect(screen.getByText('Barnet')).toBeInTheDocument();
});
it('marks only the active option selected', () => {
render(<SuggestList id="s" suggestions={ROWS} activeIndex={1}
onPick={() => {}} onHover={() => {}} />);
const options = screen.getAllByRole('option');
expect(options[0]).toHaveAttribute('aria-selected', 'false');
expect(options[1]).toHaveAttribute('aria-selected', 'true');
});
it('gives each option the id the input will point at', () => {
// aria-activedescendant on the input has to name a real element id, or
// a screen reader announces nothing as the user arrows through.
render(<SuggestList id="s" suggestions={ROWS} activeIndex={0}
onPick={() => {}} onHover={() => {}} />);
expect(screen.getAllByRole('option')[0]).toHaveAttribute(
'id', suggestOptionId('s', 0));
});
it('renders nothing when there is nothing to suggest', () => {
const { container } = render(<SuggestList id="s" suggestions={[]}
activeIndex={-1} onPick={() => {}} onHover={() => {}} />);
expect(container).toBeEmptyDOMElement();
});
});
@@ -0,0 +1,78 @@
import fs from 'fs';
import path from 'path';
/**
* Guards against light-theme-only CSS.
*
* The site themes entirely through tokens redefined under
* `@media (prefers-color-scheme: dark)`. A hardcoded colour therefore does not
* fail loudly — it renders perfectly in the theme it was written for and
* quietly wrongly in the other, which nobody sees unless they happen to be in
* dark mode when they look.
*
* Both rules below are drawn from real defects in SchoolHeroMap.module.css,
* found by eye rather than by any test:
*
* - the map's fade to the header ramped through hardcoded white and landed on
* `var(--bg-card)`. Invisible in light; a bright band across the full width
* of a near-black card in dark.
* - the controls floating over the map paired a hardcoded white background
* with `color: var(--text-primary)`, which resolves to #E9EEF0 in dark —
* near-white text on a near-white button.
*/
const COMPONENTS = path.join(__dirname, '..', '..', 'components');
function stylesheets(dir: string): string[] {
return fs.readdirSync(dir, { withFileTypes: true }).flatMap((entry) => {
const full = path.join(dir, entry.name);
if (entry.isDirectory()) return stylesheets(full);
return entry.name.endsWith('.module.css') ? [full] : [];
});
}
/** Innermost `selector { body }` pairs. Nested at-rules never match as rules,
* because their body contains braces. */
function rules(css: string): Array<{ selector: string; body: string }> {
return Array.from(css.matchAll(/([^{}]+)\{([^{}]*)\}/g), (m) => ({
selector: m[1].trim().split('\n').pop()!.trim(),
body: m[2],
}));
}
const HARDCODED_WHITE_BG = /background[^;]*(?:255,\s*255,\s*255|#fff\b|#ffffff\b)/i;
const THEMED_COLOR = /(?:^|[^-])color:\s*var\(--/;
const files = stylesheets(COMPONENTS);
describe('dark-theme safety', () => {
it('finds stylesheets to check', () => {
expect(files.length).toBeGreaterThan(0);
});
it('never pairs a hardcoded white background with a themed text colour', () => {
const offenders = files.flatMap((file) =>
rules(fs.readFileSync(file, 'utf8'))
.filter((r) => HARDCODED_WHITE_BG.test(r.body) && THEMED_COLOR.test(r.body))
.map((r) => `${path.relative(COMPONENTS, file)} ${r.selector}`));
// Either the surface follows the theme and so should the text, or it does
// not and the text must be literal too. Mixing them is how near-white text
// ends up on a near-white button.
expect(offenders).toEqual([]);
});
it('never fades to a themed colour through a hardcoded one', () => {
const offenders = files.flatMap((file) =>
rules(fs.readFileSync(file, 'utf8'))
.filter((r) => /linear-gradient/.test(r.body)
&& /var\(--bg-(card|primary|secondary)\)/.test(r.body)
&& /255,\s*255,\s*255|#fff\b/i.test(r.body))
.map((r) => `${path.relative(COMPONENTS, file)} ${r.selector}`));
// A gradient that lands on a token has to be made of that token, or the
// ramp and its destination disagree in one theme. Use the matching
// `--*-rgb` token for the transparent stops.
expect(offenders).toEqual([]);
});
});
@@ -0,0 +1,73 @@
import { renderHook, act, waitFor } from '@testing-library/react';
import { useSchoolSuggest } from '@/hooks/useSchoolSuggest';
const realFetch = global.fetch;
function mockFetch(rows: unknown[], delayMs = 0) {
global.fetch = jest.fn(async (_url: unknown, init?: { signal?: AbortSignal }) => {
if (delayMs) {
await new Promise((resolve, reject) => {
const t = setTimeout(resolve, delayMs);
init?.signal?.addEventListener('abort', () => {
clearTimeout(t);
reject(Object.assign(new Error('aborted'), { name: 'AbortError' }));
});
});
}
return { ok: true, json: async () => ({ suggestions: rows }) };
}) as unknown as typeof fetch;
}
const ROW = {
urn: 1, school_name: 'Brecknock Primary School', local_authority: 'Camden',
postcode: 'NW1 1AA', phase: 'Primary', school_type: 'Community school',
};
describe('useSchoolSuggest', () => {
beforeEach(() => { jest.useFakeTimers(); });
afterEach(() => { jest.useRealTimers(); global.fetch = realFetch; });
it('does not fetch below the minimum query length', () => {
mockFetch([ROW]);
renderHook(() => useSchoolSuggest('b', true));
act(() => { jest.advanceTimersByTime(500); });
expect(global.fetch).not.toHaveBeenCalled();
});
it('does not fetch at all when disabled', () => {
// The flag being off must mean no request, not a hidden dropdown.
mockFetch([ROW]);
renderHook(() => useSchoolSuggest('brecknock', false));
act(() => { jest.advanceTimersByTime(500); });
expect(global.fetch).not.toHaveBeenCalled();
});
it('debounces rather than firing per keystroke', () => {
mockFetch([ROW]);
const { rerender } = renderHook(
({ q }) => useSchoolSuggest(q, true), { initialProps: { q: 'br' } });
rerender({ q: 'bre' });
rerender({ q: 'brec' });
act(() => { jest.advanceTimersByTime(199); });
expect(global.fetch).not.toHaveBeenCalled();
act(() => { jest.advanceTimersByTime(2); });
expect(global.fetch).toHaveBeenCalledTimes(1);
});
it('opens with results once they arrive', async () => {
mockFetch([ROW]);
const { result } = renderHook(() => useSchoolSuggest('brecknock', true));
act(() => { jest.advanceTimersByTime(200); });
await waitFor(() => expect(result.current.suggestions).toHaveLength(1));
expect(result.current.open).toBe(true);
});
it('close() hides the list without clearing the query', async () => {
mockFetch([ROW]);
const { result } = renderHook(() => useSchoolSuggest('brecknock', true));
act(() => { jest.advanceTimersByTime(200); });
await waitFor(() => expect(result.current.open).toBe(true));
act(() => { result.current.close(); });
expect(result.current.open).toBe(false);
});
});
+34
View File
@@ -0,0 +1,34 @@
import { getFlags } from '@/lib/flags';
// jsdom provides no global fetch, so there is nothing for jest.spyOn to attach
// to — assign it and restore the original afterwards. This is the first test
// here to mock fetch; later ones should follow this shape.
const realFetch = global.fetch;
function mockFetch(impl: () => Promise<unknown>) {
global.fetch = jest.fn(impl) as unknown as typeof fetch;
}
describe('getFlags', () => {
afterEach(() => { global.fetch = realFetch; });
it('returns the flags the API reports', async () => {
mockFetch(async () => ({
ok: true,
json: async () => ({ admission_distance: true }),
}));
await expect(getFlags()).resolves.toEqual({ admission_distance: true });
});
it('returns no flags rather than throwing when the API is down', async () => {
// A page that cannot read flags must render everything dark, not 500.
// Fail-closed is the same direction as the backend's default.
mockFetch(async () => { throw new Error('ECONNREFUSED'); });
await expect(getFlags()).resolves.toEqual({});
});
it('returns no flags rather than throwing on a non-200', async () => {
mockFetch(async () => ({ ok: false, status: 503 }));
await expect(getFlags()).resolves.toEqual({});
});
});
+4 -2
View File
@@ -5,9 +5,11 @@ import { AdmissionsView } from '@/components/AdmissionsView';
export const dynamic = 'force-static';
export const metadata: Metadata = {
title: 'School Admissions Guide',
// Deadlines and offer days are what gets searched, and what this page is
// genuinely best at — the countdowns are live.
title: { absolute: 'School Admissions Deadlines & Offer Days | schoolcompare' },
description:
'Understand the Primary and Secondary school admissions process in England, with live countdowns to every key deadline and National Offer Day.',
'Every key date for primary and secondary school admissions in England, with live countdowns to the application deadline and National Offer Day.',
alternates: { canonical: absoluteUrl('/admissions') },
};
+19
View File
@@ -26,8 +26,27 @@ function backendBase(): string {
const STRIPPED_RESPONSE_HEADERS = ['content-encoding', 'content-length', 'transfer-encoding', 'connection'];
const METHODS_WITH_BODY = new Set(['POST', 'PUT', 'PATCH', 'DELETE']);
/*
* API paths this public proxy must not forward.
*
* Matched on the first segment, exactly — a prefix match would take
* /api/flagship down with /api/flags.
*
* `flags` is here because GET /api/flags names every unreleased feature the
* codebase knows about, along with whether it is on. Publishing that defeats
* the point of shipping dark. Next reads it server-side via FASTAPI_URL, on
* the Docker network, which never transits this route.
*
* Anything else internal-only belongs here too.
*/
const INTERNAL_ONLY_SEGMENTS = new Set(['flags']);
async function handler(req: NextRequest, ctx: { params: Promise<{ path: string[] }> }) {
const { path } = await ctx.params;
if (INTERNAL_ONLY_SEGMENTS.has(path[0])) {
return NextResponse.json({ detail: 'Not Found' }, { status: 404 });
}
const target = `${backendBase()}/${path.join('/')}${req.nextUrl.search}`;
const headers = new Headers(req.headers);
+5 -2
View File
@@ -30,9 +30,12 @@ export async function generateMetadata(
const { urns } = await searchParams;
const base: Metadata = {
title: 'Compare Schools',
// Deliberately not the homepage's phrase. Two pages chasing "compare
// schools" is how a site competes with itself; this one takes the tool
// phrasing instead.
title: 'School Comparison Tool — Up to Five at Once | schoolcompare',
description:
'Compare schools in England side by side — Ofsted inspections, KS2 and GCSE results against the England average, admissions odds and school community.',
'Put up to five English schools in one table: SATs and GCSE results against the England average, Ofsted grades, and the distance places were offered.',
keywords:
'school comparison, compare schools, Ofsted comparison, school admissions, KS2 comparison, primary school performance',
alternates: { canonical: absoluteUrl('/compare') },
+4
View File
@@ -28,6 +28,9 @@
--bg-primary: #FAFAF8; /* Warm White */
--bg-secondary: #F5EFE6; /* Sand — hero panels, sunken rows */
--bg-card: #FFFFFF;
/* For gradients that have to fade to the card colour. A hardcoded white
ramp reads as a bright band against a dark card. */
--bg-card-rgb: 255, 255, 255;
--surface-inverse: #0F766E;
/* ── Ink ────────────────────────────────────────────────────────── */
@@ -234,6 +237,7 @@
--bg-primary: #111A20;
--bg-secondary: #16222A;
--bg-card: #18242C;
--bg-card-rgb: 24, 36, 44;
--surface-inverse: #E9EEF0;
--text-primary: #E9EEF0;
+9 -6
View File
@@ -48,10 +48,11 @@ export const metadata: Metadata = {
statusBarStyle: 'default',
},
title: {
default: 'schoolcompare | Compare School Performance',
default: 'Compare Schools Side by Side | schoolcompare',
template: '%s | schoolcompare',
},
description: 'Compare primary and secondary school SATs and GCSE performance across England',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place. Free, no sign-up.',
keywords: 'school comparison, KS2 results, KS4 results, primary school, secondary school, England schools, SATs results, GCSE results',
authors: [{ name: 'schoolcompare' }],
manifest: '/manifest.json',
@@ -61,16 +62,18 @@ export const metadata: Metadata = {
metadataBase: new URL(SITE_URL),
openGraph: {
type: 'website',
title: 'schoolcompare | Compare School Performance',
description: 'Compare primary and secondary school SATs and GCSE performance across England',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
url: SITE_URL,
siteName: 'schoolcompare',
},
twitter: {
// summary_large_image now that there is an image worth showing.
card: 'summary_large_image',
title: 'schoolcompare | Compare School Performance',
description: 'Compare primary and secondary school SATs and GCSE performance across England',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
},
};
+23 -2
View File
@@ -8,6 +8,7 @@ import type { Metadata } from 'next';
import { fetchSchools, fetchFilters, fetchDataInfo } from '@/lib/api';
import { formatAcademicYear } from '@/lib/utils';
import { HomeView } from '@/components/HomeView';
import { getFlags } from '@/lib/flags';
import { HowItWorksSection } from '@/components/HowItWorksSection';
import { EditorialSection } from '@/components/EditorialSection';
@@ -34,8 +35,21 @@ interface HomePageProps {
* saying the brand twice.
*/
export const metadata: Metadata = {
title: { absolute: 'schoolcompare | Compare every school in England' },
description: 'Search and compare school performance across England',
/*
* Intent in the title, differentiator in the description.
*
* These queries are owned by the DfE's own "Compare school performance"
* service, and the old title — brand first, then a near-paraphrase of that
* service's name — gave a searcher no reason to pick us over it. It drew
* 0.43% CTR at position 6.1 while the brand query drew 9.16% from the same
* neighbourhood, so the ranking was never the problem.
*
* The title now matches what people type. The description carries the one
* fact gov.uk does not publish: how close you had to live to get a place.
*/
title: { absolute: 'Compare Schools Side by Side | schoolcompare' },
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place. Free, no sign-up.',
// This page reads eleven search params. They filter a result set; they do
// not make a new document. Collapsing every combination onto "/" stops the
// homepage competing with itself for its own head terms.
@@ -50,6 +64,11 @@ export default async function HomePage({ searchParams }: HomePageProps) {
// Await search params (Next.js 15 requirement)
const params = await searchParams;
// Server-read: no flag value reaches the browser bundle. Threaded down to
// both FilterBar instances via HomeView.
const flags = await getFlags();
const autosuggest = flags.school_autosuggest === true;
// Parse search params
const page = parseInt(params.page || '1');
const radius = params.radius ? parseFloat(params.radius) : undefined;
@@ -98,6 +117,7 @@ export default async function HomePage({ searchParams }: HomePageProps) {
const years = dataInfo?.years_available ?? [];
return (
<HomeView
autosuggest={autosuggest}
initialSchools={schoolsData}
filters={resolvedFilters}
totalSchools={total}
@@ -118,6 +138,7 @@ export default async function HomePage({ searchParams }: HomePageProps) {
const emptyFilters = { local_authorities: [], school_types: [], years: [], phases: [], genders: [], admissions_policies: [] };
return (
<HomeView
autosuggest={autosuggest}
initialSchools={{ schools: [], page: 1, page_size: 50, total: 0, total_pages: 0 }}
filters={emptyFilters}
totalSchools={null}
+5 -2
View File
@@ -18,8 +18,11 @@ interface RankingsPageProps {
}
export const metadata: Metadata = {
title: 'School Rankings',
description: 'Top-ranked schools by SATs and GCSE performance across England',
// 'School Rankings' matched nothing anyone types. League tables is the
// phrase parents actually search, and it spikes each results day.
title: { absolute: 'Primary & Secondary School League Tables | schoolcompare' },
description:
'Rank English schools by SATs results, GCSEs, Progress 8 or Attainment 8, and filter by local authority or year. Built from the DfE’s own figures.',
keywords: 'school rankings, top schools, best schools, KS2 rankings, KS4 rankings, school league tables',
// Param forms (?metric=&local_authority=&year=&phase=) collapse here for
// now. W3 replaces them with real indexable paths.
@@ -0,0 +1,63 @@
/**
* Phase variants of a place page.
*
* Phase is part of the query — "primary schools in beccles", "secondary
* schools in brentwood" — not a filter applied afterwards, so each gets its
* own indexable path. A place with no schools of the phase has no page: the
* per-phase threshold, not an error.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ place: string; phase: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
const PHASES = ['primary', 'secondary'] as const;
type Phase = (typeof PHASES)[number];
const isPhase = (v: string): v is Phase => (PHASES as readonly string[]).includes(v);
async function resolve(slug: string, phase: Phase) {
return (await fetchPlace('town', slug, phase))
?? (await fetchPlace('locality', slug, phase));
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { place: slug, phase } = await params;
if (!isPhase(phase)) return { title: 'Place Not Found' };
const detail = await resolve(slug, phase);
if (!detail || detail.schools.length === 0) return { title: 'Place Not Found' };
const word = phase === 'secondary' ? 'Secondary' : 'Primary';
const { name } = detail.place;
return {
// Not "Ranked": the table is alphabetical, so the word would be a claim
// the page does not keep.
title: { absolute: `${word} Schools in ${name} | schoolcompare` },
description:
`Every ${phase} school in ${name}, with results, Ofsted grades and the local `
+ `average against England.`,
alternates: { canonical: absoluteUrl(`/schools/${slug}/${phase}`) },
};
}
export default async function PlacePhasePage({ params }: Props) {
const { place: slug, phase } = await params;
if (!isPhase(phase)) notFound();
const detail = await resolve(slug, phase);
if (!detail || detail.schools.length === 0) notFound();
const national = await fetchNationalAverages().catch(() => null);
const englandAverage = phase === 'secondary'
? national?.secondary?.attainment_8_score ?? null
: national?.primary?.rwm_expected_pct ?? null;
return <PlaceView detail={detail} phase={phase}
englandAverage={englandAverage} neighbours={[]} />;
}
+99
View File
@@ -0,0 +1,99 @@
/**
* Town and locality pages.
*
* A place below the five-school threshold is not in the registry, so
* fetchPlace returns null and the request 404s rather than rendering a page
* with nothing to say.
*/
import { notFound, redirect } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces, authoritySlug } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ place: string }> }
// ISR: place aggregates change only when the pipeline runs.
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ place: string }>> {
// Off by default: ~2,000 place routes cannot be built in CI on every deploy.
// Matches the PRERENDER_SCHOOLS gate on the school route.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'town' || p.kind === 'locality')
.map((p) => ({ place: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
async function resolve(slug: string) {
return (await fetchPlace('town', slug)) ?? (await fetchPlace('locality', slug));
}
/** Other towns in the same authority — the cheapest honest definition of
* "nearby", and enough to stop each place page being a dead end. */
async function neighboursOf(detail: { place: { slug: string; parent_authority: string | null } }) {
if (!detail.place.parent_authority) return [];
const all = await fetchPlaces();
return all
.filter((p) => p.kind === 'town' && p.slug !== detail.place.slug)
.slice(0, 12);
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { place: slug } = await params;
const detail = await resolve(slug);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
// absolute: the root layout's template appends '| schoolcompare' to a
// plain string, and this title already carries it. Without this every
// place title read '... | schoolcompare | schoolcompare'.
title: { absolute: `Schools in ${name} — Compare ${count} Schools | schoolcompare` },
description:
`Every school in ${name}, with SATs and GCSE results, Ofsted grades, the local `
+ `average against England, and how close you had to live to get a place.`,
alternates: { canonical: absoluteUrl(`/schools/${slug}`) },
};
}
export default async function PlacePage({ params }: Props) {
const { place: slug } = await params;
const detail = await resolve(slug);
if (!detail) notFound();
// Global constraint: no page without a local average. A place with too few
// schools carrying results has nothing to say that a list does not, so it
// defers to its authority rather than publishing a thin page.
if (detail.averages.rwm_expected_pct == null
&& detail.averages.attainment_8_score == null) {
// The API's own slug, which is null when that authority is itself under
// the threshold and has no page. Re-slugifying the name here would send
// the reader to a 404 instead of telling them the place has no page.
const target = detail.place.authorities?.[0]?.slug
?? (detail.place.parent_authority
? authoritySlug(detail.place.parent_authority)
: null);
if (target) redirect(`/schools/authority/${target}`);
notFound();
}
const national = await fetchNationalAverages().catch(() => null);
// NationalAverages is nested by phase — { primary: {...}, secondary: {...} }
// — not flat. Reading it flat silently yields undefined and the page renders
// with no comparison, which is the one thing that makes it not a list.
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={await neighboursOf(detail)}
/>
);
}
@@ -0,0 +1,65 @@
/**
* Phase variants of an authority page.
*
* The spec called for these; the plan built the bare authority route and
* dropped them. Nothing caught it, because the sitemap is written from the
* place registry — which was right about them all along — while the routes
* were written by hand. 302 authority phase URLs were submitted to Google and
* every one 404'd, and every authority page linked to a phase page in the
* *town* namespace, which is a different set of schools entirely.
*
* "Primary schools in Kent" is the query these serve, and it is a real one:
* admissions are authority-run, so the authority is the unit a parent thinks
* in when they have not settled on a town.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ la: string; phase: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
const PHASES = ['primary', 'secondary'] as const;
type Phase = (typeof PHASES)[number];
const isPhase = (v: string): v is Phase => (PHASES as readonly string[]).includes(v);
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { la, phase } = await params;
if (!isPhase(phase)) return { title: 'Place Not Found' };
const detail = await fetchPlace('authority', la, phase);
if (!detail || detail.schools.length === 0) return { title: 'Place Not Found' };
const word = phase === 'secondary' ? 'Secondary' : 'Primary';
const { name } = detail.place;
return {
// "Local Authority" stays in the title for the same reason it is on the
// bare authority page: 67 town names collide with an authority name, and
// a reader landing on both needs to know which set each covers.
title: { absolute: `${word} Schools in ${name} — Local Authority | schoolcompare` },
description:
`Every ${phase} school in the ${name} local authority, with results, Ofsted `
+ `grades and the authority average against England.`,
alternates: { canonical: absoluteUrl(`/schools/authority/${la}/${phase}`) },
};
}
export default async function AuthorityPhasePage({ params }: Props) {
const { la, phase } = await params;
if (!isPhase(phase)) notFound();
const detail = await fetchPlace('authority', la, phase);
if (!detail || detail.schools.length === 0) notFound();
const national = await fetchNationalAverages().catch(() => null);
const englandAverage = phase === 'secondary'
? national?.secondary?.attainment_8_score ?? null
: national?.primary?.rwm_expected_pct ?? null;
return <PlaceView detail={detail} phase={phase}
englandAverage={englandAverage} neighbours={[]} />;
}
@@ -0,0 +1,67 @@
/**
* Local authority pages.
*
* A separate namespace from /schools/[place] because 67 town names collide
* with an authority name and neither set contains the other — Bedford the
* town holds 104 schools, Bedford the authority 86, because postal towns
* cross authority boundaries. The title says "Local Authority" so a reader
* landing on both knows which set each covers.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ la: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ la: string }>> {
// Gated like every other prerender in this app. There are only ~154
// authorities, but "few enough to always build" still means the API must be
// reachable at build time, and in CI it is not — the build fails with
// ECONNREFUSED rather than degrading. The catch is the same fallback the
// school route uses.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'authority')
.map((p) => ({ la: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { la } = await params;
const detail = await fetchPlace('authority', la);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
title: { absolute: `Schools in ${name} — Local Authority | schoolcompare` },
description:
`All ${count} schools in the ${name} local authority, with SATs and GCSE results, `
+ `Ofsted grades and the authority average against England.`,
alternates: { canonical: absoluteUrl(`/schools/authority/${la}`) },
};
}
export default async function AuthorityPage({ params }: Props) {
const { la } = await params;
const detail = await fetchPlace('authority', la);
if (!detail) notFound();
const national = await fetchNationalAverages().catch(() => null);
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={[]}
/>
);
}
@@ -0,0 +1,62 @@
/**
* Postcode district pages.
*
* No phase variants: nobody searches "primary schools in SW11", so the
* variants would be pages without demand. These exist to catch
* "schools near <postcode>" and to give London districts a geographic page
* where the GIAS town field cannot.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ outcode: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ outcode: string }>> {
// 1,760 of these; same CI budget argument as the town routes.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'outcode')
.map((p) => ({ outcode: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { outcode } = await params;
const detail = await fetchPlace('outcode', outcode);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
title: { absolute: `Schools near ${name} | schoolcompare` },
description:
`${count} schools in the ${name} postcode district, with results, Ofsted grades `
+ `and how close you had to live to get a place.`,
alternates: { canonical: absoluteUrl(`/schools/near/${outcode}`) },
};
}
export default async function OutcodePage({ params }: Props) {
const { outcode } = await params;
const detail = await fetchPlace('outcode', outcode);
if (!detail) notFound();
const national = await fetchNationalAverages().catch(() => null);
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={[]}
/>
);
}
+1 -1
View File
@@ -9,7 +9,7 @@ export const runtime = 'nodejs';
* validated here rather than passed through, so this route cannot be used to
* reach arbitrary backend paths.
*/
const CHILD = /^(static|schools-\d+)\.xml$/;
const CHILD = /^(static|schools-\d+|places-\d+|outcodes-\d+)\.xml$/;
export async function GET(
_request: Request,
@@ -48,6 +48,8 @@
display: flex;
align-items: center;
gap: 0.5rem;
/* The suggestion dropdown is absolutely positioned against this box. */
position: relative;
}
/* The hero pill: hairline, soft corner, everything else sits inside it. */
+66 -1
View File
@@ -3,8 +3,11 @@
import { useState, useCallback, useTransition, useRef, useEffect } from "react";
import type { ReactNode } from "react";
import { useRouter, useSearchParams, usePathname } from "next/navigation";
import { isValidPostcode } from "@/lib/utils";
import { isValidPostcode, schoolUrl } from "@/lib/utils";
import { track } from "@/lib/analytics";
import { useSchoolSuggest } from "@/hooks/useSchoolSuggest";
import { SuggestList, suggestOptionId } from "./SuggestList";
import type { Suggestion } from "@/lib/suggest";
import type { Filters, ResultFilters } from "@/lib/types";
import styles from "./FilterBar.module.css";
@@ -17,6 +20,8 @@ interface FilterBarProps {
onNearMe?: () => void;
geoState?: "idle" | "requesting" | "error";
geoError?: string | null;
/** Server-read feature flag. Off means no listener, no fetch, no markup. */
autosuggest?: boolean;
}
/**
@@ -48,6 +53,7 @@ export function FilterBar({
onNearMe,
geoState = "idle",
geoError,
autosuggest = false,
}: FilterBarProps) {
const router = useRouter();
const pathname = usePathname();
@@ -62,6 +68,45 @@ export function FilterBar({
const [omniValue, setOmniValue] = useState(initialOmniValue);
const suggestId = `school-suggest-${isHero ? "hero" : "bar"}`;
// Suppressed once the value parses as a postcode: the box takes a school
// name OR a postcode, and suggesting schools during postcode entry fights
// the user rather than helping them.
const suggestEnabled = autosuggest && !isValidPostcode(omniValue);
const { suggestions, open, activeIndex, setActiveIndex, close } =
useSchoolSuggest(omniValue, suggestEnabled);
const pickSuggestion = (s: Suggestion) => {
close();
track('search_submitted', {
query: s.school_name.toLowerCase(),
via: 'suggestion',
urn: s.urn,
has_postcode: false,
filters_active: '',
filters_count: 0,
});
router.push(schoolUrl(s.urn, s.school_name));
};
const handleOmniKeyDown = (e: React.KeyboardEvent<HTMLInputElement>) => {
if (!open) return;
if (e.key === "ArrowDown") {
e.preventDefault();
setActiveIndex(activeIndex + 1 >= suggestions.length ? 0 : activeIndex + 1);
} else if (e.key === "ArrowUp") {
e.preventDefault();
setActiveIndex(activeIndex <= 0 ? suggestions.length - 1 : activeIndex - 1);
} else if (e.key === "Escape") {
close();
} else if (e.key === "Enter" && activeIndex >= 0) {
// Only when an option is active. With none, the event falls through to
// the form's submit handler and searches the typed text, as it does now.
e.preventDefault();
pickSuggestion(suggestions[activeIndex]);
}
};
const currentLA = searchParams.get("local_authority") || "";
const currentType = searchParams.get("school_type") || "";
const currentPhase = searchParams.get("phase") || "";
@@ -227,8 +272,19 @@ export function FilterBar({
type="search"
value={omniValue}
onChange={(e) => setOmniValue(e.target.value)}
onKeyDown={handleOmniKeyDown}
onBlur={close}
placeholder="School name or postcode"
className={styles.omniInput}
{...(autosuggest ? {
role: "combobox",
"aria-expanded": open,
"aria-controls": suggestId,
"aria-autocomplete": "list" as const,
"aria-activedescendant":
activeIndex >= 0 ? suggestOptionId(suggestId, activeIndex) : undefined,
autoComplete: "off",
} : {})}
/>
<button
type="submit"
@@ -237,6 +293,15 @@ export function FilterBar({
>
{isPending ? <div className={styles.spinner}></div> : isHero ? "Search schools" : "Search"}
</button>
{autosuggest && open && (
<SuggestList
id={suggestId}
suggestions={suggestions}
activeIndex={activeIndex}
onPick={pickSuggestion}
onHover={setActiveIndex}
/>
)}
</div>
{isHero && (
<>
+5 -1
View File
@@ -29,6 +29,8 @@ interface HomeViewProps {
// show (e.g. an active search).
howItWorks?: React.ReactNode;
editorial?: React.ReactNode;
/** Server-read feature flag, threaded to both FilterBar instances. */
autosuggest?: boolean;
}
function daysUntil(month: number, day: number): number {
@@ -193,7 +195,7 @@ const VALUE_PROPS: ValueProp[] = [
},
];
export function HomeView({ initialSchools, filters, totalSchools, howItWorks, editorial }: HomeViewProps) {
export function HomeView({ initialSchools, filters, totalSchools, howItWorks, editorial, autosuggest = false }: HomeViewProps) {
const searchParams = useSearchParams();
const router = useRouter();
const pathname = usePathname();
@@ -462,6 +464,7 @@ export function HomeView({ initialSchools, filters, totalSchools, howItWorks, ed
onNearMe={handleNearMe}
geoState={geoState}
geoError={geoError}
autosuggest={autosuggest}
/>
</div>
@@ -500,6 +503,7 @@ export function HomeView({ initialSchools, filters, totalSchools, howItWorks, ed
onNearMe={handleNearMe}
geoState={geoState}
geoError={geoError}
autosuggest={autosuggest}
/>
)}
+31 -7
View File
@@ -47,7 +47,13 @@
width: 100%;
height: 100%;
background:
linear-gradient(100deg, rgba(255, 255, 255, 0) 40%, rgba(255, 255, 255, .5) 50%, rgba(255, 255, 255, 0) 60%) var(--bg-secondary);
/* Sweeps toward the card colour, which is a shade lighter than this
ground in both themes. Hardcoded white was a bright flash across a
dark page every 1.4s while the tiles loaded. */
linear-gradient(100deg,
rgba(var(--bg-card-rgb), 0) 40%,
rgba(var(--bg-card-rgb), .5) 50%,
rgba(var(--bg-card-rgb), 0) 60%) var(--bg-secondary);
background-size: 200% 100%;
animation: shimmer 1.4s infinite;
}
@@ -76,6 +82,15 @@
justify-content: center;
}
/*
* Controls that float ON the map.
*
* The map tiles are light in both themes, so these deliberately do NOT follow
* the theme — they follow the map. The literal ink below is the point: paired
* with a hardcoded white background, `color: var(--text-primary)` resolved to
* #E9EEF0 in the dark theme and put near-white text on a near-white button.
* A themed token is the wrong tool for a surface that never changes.
*/
.openHint {
display: inline-flex;
align-items: center;
@@ -85,7 +100,8 @@
border-radius: 999px;
font-size: 13px;
font-weight: 600;
color: var(--text-primary);
/* See "Controls that float ON the map" above. */
color: #1C2731;
background: rgba(255, 255, 255, .85);
-webkit-backdrop-filter: blur(6px);
backdrop-filter: blur(6px);
@@ -113,11 +129,18 @@
on top of the blend. */
z-index: 450;
pointer-events: none;
/* The card colour, not white.
This ramp was hardcoded white and ended at var(--bg-card). In the light
theme that is white into white and invisible, as intended. In the dark
theme it climbed to 95% WHITE and then met a near-black card — a bright
band across the full width, right where the map is supposed to dissolve
into the header. Fading to the same colour the gradient lands on is the
whole trick, and it only works if that colour is a token. */
background: linear-gradient(to bottom,
rgba(255, 255, 255, 0) 0%,
rgba(255, 255, 255, .35) 35%,
rgba(255, 255, 255, .75) 62%,
rgba(255, 255, 255, .95) 82%,
rgba(var(--bg-card-rgb), 0) 0%,
rgba(var(--bg-card-rgb), .35) 35%,
rgba(var(--bg-card-rgb), .75) 62%,
rgba(var(--bg-card-rgb), .95) 82%,
var(--bg-card) 100%);
}
@@ -134,7 +157,8 @@
border: none;
border-radius: 8px;
background: rgba(255, 255, 255, .92);
color: var(--text-primary);
/* See "Controls that float ON the map" above. */
color: #1C2731;
cursor: pointer;
box-shadow: 0 2px 10px rgba(var(--shadow-rgb), .2);
}
@@ -0,0 +1,53 @@
/*
* Anchored to .omniBoxContainer, which is position: relative for this reason.
*
* Every colour is a token, so the dropdown follows the theme. The dark theme
* redefines --bg-card, --border, --text-muted and --shadow-soft, and this
* inherits all four without a second rule.
*/
.list {
position: absolute;
top: calc(100% + 4px);
left: 0;
right: 0;
/* Above the sticky filter bar (10) and the hero layers (0–2), below the
skip-link (10000) and the modal overlay (1000). */
z-index: 40;
margin: 0;
padding: 4px;
list-style: none;
max-height: 320px;
overflow-y: auto;
background: var(--bg-card);
border: 1px solid var(--border);
border-radius: var(--radius-md);
box-shadow: var(--shadow-soft);
}
.option {
display: flex;
align-items: baseline;
justify-content: space-between;
gap: 12px;
padding: 10px 12px;
border-radius: var(--radius-sm);
cursor: pointer;
color: var(--text-primary);
}
/* Hover and keyboard share one style: the active option is the active option
however it became active. Two rules would drift. */
.option:hover,
.active {
background: var(--bg-secondary);
}
.name {
font-weight: 500;
}
.meta {
font-size: 0.85em;
color: var(--text-muted);
white-space: nowrap;
}
+53
View File
@@ -0,0 +1,53 @@
'use client';
/**
* The autosuggest dropdown. Presentational only — it fetches nothing and owns
* no state, so the fetching rules and the ARIA rules can be read separately.
*/
import type { Suggestion } from '@/lib/suggest';
import styles from './SuggestList.module.css';
/** The id the input's aria-activedescendant points at. */
export function suggestOptionId(id: string, index: number): string {
return `${id}-option-${index}`;
}
interface Props {
/** Shared with the input's aria-controls. */
id: string;
suggestions: Suggestion[];
activeIndex: number;
onPick: (s: Suggestion) => void;
onHover: (index: number) => void;
}
export function SuggestList({ id, suggestions, activeIndex, onPick, onHover }: Props) {
if (suggestions.length === 0) return null;
return (
<ul className={styles.list} id={id} role="listbox">
{suggestions.map((s, i) => (
<li
key={s.urn}
id={suggestOptionId(id, i)}
role="option"
aria-selected={i === activeIndex}
className={`${styles.option} ${i === activeIndex ? styles.active : ''}`}
/*
* onMouseDown, not onClick: the input's blur handler closes the list,
* and blur fires before click — so a click handler never runs. This
* is the classic autosuggest bug where the dropdown is unclickable
* with a mouse while working perfectly with a keyboard.
*/
onMouseDown={(e) => { e.preventDefault(); onPick(s); }}
onMouseEnter={() => onHover(i)}
>
<span className={styles.name}>{s.school_name}</span>
{/* Not decoration: there are many "St Mary's". */}
<span className={styles.meta}>{s.local_authority}</span>
</li>
))}
</ul>
);
}
@@ -0,0 +1,210 @@
/* Tokens only — see globals.css. Follows RankingsView's conventions, and in
particular its link treatment: table links take --text-primary with no
underline and a brand-coloured hover, not the browser default. The first
cut used bare <Link> with no class at all, which rendered as default blue
underlined links and read as unstyled beside the rest of the site. */
.container {
width: 100%;
min-width: 0;
}
.header {
margin-bottom: 1.5rem;
}
.header h1 {
font-size: 2.25rem;
font-weight: 700;
color: var(--text-primary);
margin-bottom: 0.5rem;
font-family: var(--font-display);
text-wrap: balance;
}
.summary {
font-size: 1rem;
color: var(--text-secondary);
margin: 0;
line-height: 1.6;
}
/* Links in running copy: brand colour, underline on hover only. */
.inlineLink {
color: var(--brand);
text-decoration: none;
transition: color 0.2s ease;
}
.inlineLink:hover {
color: var(--brand-strong);
text-decoration: underline;
}
/* Phase variants are separate indexable pages, so the bare place page has to
link them — a sitemap entry alone leaves them with no internal path in. */
.phaseLinks {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 0.75rem;
margin: 0 0 1.25rem;
}
.phaseLink {
display: inline-block;
padding: 0.4rem 0.875rem;
border: 1px solid var(--border-strong);
border-radius: 999px;
font-size: 0.875rem;
font-weight: 500;
color: var(--text-primary);
text-decoration: none;
transition: border-color 0.2s ease, color 0.2s ease;
}
.phaseLink:hover {
border-color: var(--brand);
color: var(--brand-strong);
}
/* The one number a list cannot give you, so it gets its own band. */
.compare {
background: var(--bg-secondary);
border: 1px solid var(--border);
border-radius: 8px;
padding: 0.875rem 1.125rem;
margin: 0 0 1.5rem;
color: var(--text-primary);
font-size: 1rem;
}
.ofsted {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1.25rem;
list-style: none;
padding: 0;
margin: 0 0 1.5rem;
font-size: 0.9375rem;
color: var(--text-secondary);
}
.group {
margin-bottom: 2rem;
}
.groupHeading {
display: flex;
align-items: baseline;
gap: 0.625rem;
font-size: 1.25rem;
font-weight: 600;
color: var(--text-primary);
font-family: var(--font-display);
margin: 0 0 0.75rem;
}
.groupCount {
font-size: 0.8125rem;
font-weight: 500;
color: var(--text-secondary);
background: var(--bg-secondary);
border-radius: 999px;
padding: 0.125rem 0.5rem;
}
/* Wide content scrolls in its own container so the page body never does. */
.tableWrap {
overflow-x: auto;
border: 1px solid var(--border);
border-radius: 8px;
background: var(--bg-card);
}
.table {
width: 100%;
border-collapse: collapse;
font-size: 0.9375rem;
}
.table th,
.table td {
padding: 0.75rem 1rem;
text-align: left;
border-bottom: 1px solid var(--border);
}
.table th {
background: var(--bg-secondary);
color: var(--text-secondary);
font-weight: 600;
font-size: 0.8125rem;
}
.table tbody tr:last-child td {
border-bottom: none;
}
/*
* Header and value share one class and one rule, so they cannot drift apart.
*
* The first cut aligned them with two different selectors: `.table th:last-child`
* at (0,2,1) beat the element rule and went right, while `.num` at (0,1,0) lost
* to `.table td` at (0,1,1) and stayed left. The heading and its numbers sat on
* opposite edges of the column.
*
* width:1% with nowrap makes the measure column hug its content so the school
* name takes the remaining width — without it the two columns split evenly and
* the gap between heading and value reads as misalignment on a wide screen.
*/
.table th.num,
.table td.num {
text-align: right;
font-variant-numeric: tabular-nums;
width: 1%;
white-space: nowrap;
}
/* The measure is spelled out; the tooltip carries the definition. */
.metricHead {
text-decoration: none;
cursor: help;
border-bottom: 1px dotted var(--border-strong);
}
/* Table links: site convention is body colour, brand on hover. */
.schoolLink {
color: var(--text-primary);
text-decoration: none;
transition: color 0.2s ease;
}
.schoolLink:hover {
color: var(--brand-strong);
}
/* "Not published" is a fact about the school, not an error. */
.noData {
color: var(--text-muted);
font-size: 0.8125rem;
}
.neighbours {
margin-top: 2rem;
}
.neighbours h2 {
font-size: 1.125rem;
font-weight: 600;
color: var(--text-primary);
margin: 0 0 0.75rem;
font-family: var(--font-display);
}
.neighbours ul {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1rem;
list-style: none;
padding: 0;
margin: 0;
}
+258
View File
@@ -0,0 +1,258 @@
/**
* One place page, shared by all four families.
*
* They differ in what fills the registry, not in what the page shows, so a
* second component would be a second place to forget the same change.
*
* The local-versus-England comparison is the reason this page is not a list:
* it is the one number a parent cannot get by reading the schools one by one,
* and it is what keeps the page from reading as a name dropped into a
* template.
*/
import Link from 'next/link';
import type { PlaceDetail, PlaceSummary } from '@/lib/places';
import { placeUrl, authoritySlug } from '@/lib/places';
import type { School } from '@/lib/types';
import { schoolUrl } from '@/lib/utils';
import { absoluteUrl } from '@/lib/site';
import styles from './PlaceView.module.css';
interface Props {
detail: PlaceDetail;
phase?: 'primary' | 'secondary';
englandAverage: number | null;
/** Nearby places, so the page links onward instead of dead-ending. */
neighbours: PlaceSummary[];
}
// Ofsted grades in the order they are reported.
const OFSTED_LABELS: Array<[number, string]> = [
[1, 'Outstanding'], [2, 'Good'],
[3, 'Requires improvement'], [4, 'Inadequate'],
];
/**
* Column headings, taken from the site's own metric dictionary rather than
* invented here — see METRIC_DEFINITIONS in backend/schemas.py, surfaced at
* /api/metrics. The first cut said "RWM expected", which is jargon that
* appears nowhere else on the site.
*/
const METRICS = {
primary: {
key: 'rwm_expected_pct' as const,
heading: 'Reading, writing & maths',
hint: '% meeting the expected standard in reading, writing and maths',
unit: '%',
},
secondary: {
key: 'attainment_8_score' as const,
heading: 'Attainment 8',
hint: "Average grade across a pupil's best 8 GCSEs, including English and maths",
unit: '',
},
};
type PhaseKey = keyof typeof METRICS;
/** All-through schools sit in both phases, matching the search filters. */
function isPhase(school: School, phase: PhaseKey): boolean {
const p = (school.phase ?? '').toLowerCase();
if (p === 'all-through') return true;
return phase === 'secondary'
? p.includes('secondary') || p === '16 plus'
: p.includes('primary') || p.includes('middle');
}
function SchoolTable({ schools, phase }: { schools: School[]; phase: PhaseKey }) {
const metric = METRICS[phase];
return (
<div className={styles.tableWrap}>
<table className={styles.table}>
<thead>
<tr>
<th scope="col">School</th>
{/* Same class as the value cell below: one rule aligns both, so
they cannot drift apart. */}
<th scope="col" className={styles.num}>
<abbr className={styles.metricHead} title={metric.hint}>
{metric.heading}
</abbr>
</th>
</tr>
</thead>
<tbody>
{schools.map((s) => {
const value = s[metric.key];
return (
<tr key={s.urn}>
<td>
<Link href={schoolUrl(s.urn, s.school_name)} className={styles.schoolLink}>
{s.school_name}
</Link>
</td>
<td className={styles.num}>
{value == null
? <span className={styles.noData}>Not published</span>
: `${Math.round(Number(value))}${metric.unit}`}
</td>
</tr>
);
})}
</tbody>
</table>
</div>
);
}
export function PlaceView({ detail, phase, englandAverage, neighbours }: Props) {
const { place, schools, averages } = detail;
// Fall back to the single parent when the API predates the authorities
// field, so a stale cache never blanks the line entirely.
const authorities = place.authorities?.length
? place.authorities
: place.parent_authority
? [{ name: place.parent_authority, slug: authoritySlug(place.parent_authority), count: 0 }]
: [];
const local = averages[METRICS[phase ?? 'primary'].key];
const phaseWord = phase === 'secondary' ? 'Secondary schools'
: phase === 'primary' ? 'Primary schools' : 'Schools';
const graded = OFSTED_LABELS
.map(([grade, label]) => [label, schools.filter((s) => s.ofsted_grade === grade).length] as const)
.filter(([, n]) => n > 0);
/*
* An unphased page holds both primaries and secondaries, and they are
* scored on different measures — a percentage and a 0-90 score. Showing one
* column for both left 30% of rows blank on /schools/brentwood and put two
* incomparable scales in one column when it did not.
*
* So the phases get a table each. A blank cell inside one now means the
* school genuinely has no published result, which is worth saying.
*/
const groups: Array<[PhaseKey, School[]]> = phase
? [[phase, schools]]
: (['primary', 'secondary'] as PhaseKey[])
.map((p) => [p, schools.filter((s) => isPhase(s, p))] as [PhaseKey, School[]])
.filter(([, list]) => list.length > 0);
const jsonLd = {
'@context': 'https://schema.org',
'@graph': [
{
'@type': 'ItemList',
name: `${phaseWord} in ${place.name}`,
numberOfItems: schools.length,
// Alphabetical, and said so. Without this an ItemList carrying
// `position` reads as a ranking, which would be a claim the page
// stopped making when the table became A-Z.
itemListOrder: 'https://schema.org/ItemListOrderAscending',
itemListElement: schools.slice(0, 20).map((s, i) => ({
'@type': 'ListItem',
position: i + 1,
url: absoluteUrl(schoolUrl(s.urn, s.school_name)),
name: s.school_name,
})),
},
{
'@type': 'BreadcrumbList',
itemListElement: [
{ '@type': 'ListItem', position: 1, name: 'Schools', item: absoluteUrl('/') },
{ '@type': 'ListItem', position: 2, name: place.name },
],
},
],
};
return (
<div className={styles.container}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<header className={styles.header}>
<h1>{phaseWord} in {place.name}</h1>
<p className={styles.summary}>
{place.count} schools
{authorities.length > 0 && (
<>
{' · '}
{/* Every authority, not just the largest. A quarter of outcodes
and a third of towns cross a boundary: SW19 is mostly Merton
but partly Wandsworth, and naming one asserts otherwise. */}
{authorities.map((a, i) => (
<span key={a.name}>
{i > 0 && (i === authorities.length - 1 ? ' and ' : ', ')}
{/* No slug means no page: City of London and the Isles of
Scilly hold too few schools for one. Saying where the
place is stays right; linking there would 404. */}
{a.slug
? (
<Link href={`/schools/authority/${a.slug}`} className={styles.inlineLink}>
{a.name}
</Link>
)
: a.name}
</span>
))}
</>
)}
</p>
</header>
{!phase && (place.phases ?? []).length > 0 && (
<nav className={styles.phaseLinks} aria-label="By phase">
{(place.phases ?? []).map((ph) => (
/* placeUrl, not a template: the bare `/schools/[slug]/[phase]`
shape belongs to towns alone, and using it everywhere sent
every authority page into the town namespace. */
<Link key={ph} href={placeUrl(place.kind, place.slug, ph)}
className={styles.phaseLink}>
{ph === 'secondary' ? 'Secondary schools' : 'Primary schools'} in {place.name}
</Link>
))}
</nav>
)}
{local != null && englandAverage != null && (
<p className={styles.compare} data-testid="local-vs-england">
{place.name} averages <strong>{Math.round(local)}</strong> against{' '}
<strong>{Math.round(englandAverage)}</strong> across England.
</p>
)}
{graded.length > 0 && (
<ul className={styles.ofsted} data-testid="ofsted-distribution">
{graded.map(([label, n]) => (
<li key={label}>{label}: <strong>{n}</strong></li>
))}
</ul>
)}
{groups.map(([p, list]) => (
<section key={p} className={styles.group}>
{groups.length > 1 && (
<h2 className={styles.groupHeading}>
{p === 'secondary' ? 'Secondary schools' : 'Primary schools'}
<span className={styles.groupCount}>{list.length}</span>
</h2>
)}
<SchoolTable schools={list} phase={p} />
</section>
))}
{neighbours.length > 0 && (
<nav className={styles.neighbours} aria-label="Nearby places">
<h2>Nearby</h2>
<ul>
{neighbours.map((n) => (
<li key={n.kind + n.slug}>
<Link href={placeUrl(n.kind, n.slug)} className={styles.inlineLink}>{n.name}</Link>
</li>
))}
</ul>
</nav>
)}
</div>
);
}
+58
View File
@@ -0,0 +1,58 @@
'use client';
import { useEffect, useRef, useState } from 'react';
import { fetchSuggestions, SUGGEST_MIN_QUERY, type Suggestion } from '@/lib/suggest';
/*
* Long enough that a fast typist does not fire a request per character, short
* enough that the list feels attached to the keyboard.
*/
const DEBOUNCE_MS = 200;
export function useSchoolSuggest(query: string, enabled: boolean) {
const [suggestions, setSuggestions] = useState<Suggestion[]>([]);
const [open, setOpen] = useState(false);
const [activeIndex, setActiveIndex] = useState(-1);
// Set when the user dismisses the list, so a re-render does not reopen it.
const dismissed = useRef('');
useEffect(() => {
const q = query.trim();
if (!enabled || q.length < SUGGEST_MIN_QUERY || dismissed.current === q) {
setSuggestions([]);
setOpen(false);
return;
}
/*
* Abort the superseded request on every keystroke. This is correctness,
* not economy: without it a slow response for "st" can land after the fast
* one for "st marys" and replace a correct list with a stale one.
*/
const controller = new AbortController();
const timer = setTimeout(async () => {
const rows = await fetchSuggestions(q, controller.signal);
if (controller.signal.aborted) return;
setSuggestions(rows);
setActiveIndex(-1);
setOpen(rows.length > 0);
}, DEBOUNCE_MS);
return () => {
clearTimeout(timer);
controller.abort();
};
}, [query, enabled]);
return {
suggestions,
open,
activeIndex,
setActiveIndex,
close: () => {
dismissed.current = query.trim();
setOpen(false);
setActiveIndex(-1);
},
};
}
+8
View File
@@ -12,6 +12,12 @@ jest.mock('next/navigation', () => ({
useSearchParams: () => new URLSearchParams(),
}));
// Everything below this line is browser furniture, and this file runs for
// every suite — including the ones that declare `@jest-environment node` to
// test route handlers, where NextRequest needs Fetch API globals jsdom does
// not provide. There is no `window` there, so guard rather than assume one.
if (typeof window !== 'undefined') {
// Mock window.matchMedia
Object.defineProperty(window, 'matchMedia', {
writable: true,
@@ -52,3 +58,5 @@ const localStorageMock = {
clear: jest.fn(),
};
global.localStorage = localStorageMock;
} // end: browser-only globals
+40
View File
@@ -0,0 +1,40 @@
/**
* Reading feature flags.
*
* Server-side only. No flag value reaches the browser bundle, and there is no
* Unleash dependency in package.json — the SDK lives in FastAPI, which already
* owns every other piece of data this app renders.
*
* Flags are declared in backend/flags.py. A purely front-end flag still has to
* be declared there; it is a flat data edit, and the return is that one list
* answers "what flags exist" for the whole system.
*/
export type Flags = Record<string, boolean>;
/*
* Reading flags pins the calling route to this ISR floor: Next uses the LOWEST
* revalidate among a route's fetches to set the whole route's revalidation
* frequency. 300s matches what /school/[slug] already sits at, so a page that
* reads flags is no more dynamic than a school page already is.
*
* It is also what makes a flip propagate without a webhook: five minutes on
* school pages, an hour on place pages, against flags that flip monthly.
*/
export const FLAGS_REVALIDATE = 300;
const API = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL
|| 'http://localhost:8000/api';
/** Every flag and its value. Never throws: an unreadable flag is a dark one. */
export async function getFlags(): Promise<Flags> {
try {
const res = await fetch(`${API}/flags`, {
next: { revalidate: FLAGS_REVALIDATE },
});
if (!res.ok) return {};
return await res.json();
} catch {
return {};
}
}
+84
View File
@@ -0,0 +1,84 @@
/**
* Client for the places API.
*
* Two namespaces, matching the backend: towns and localities share
* /schools/[place]; authorities take /schools/authority/[la]. 67 town names
* collide with an authority name and neither set contains the other, so one
* namespace would publish near-duplicate pages.
*/
import type { School } from '@/lib/types';
export interface PlaceSummary {
kind: string;
slug: string;
name: string;
count: number;
/** Phases that clear the threshold on their own, so the page links
* variants that exist rather than 404s. Absent on the registry listing. */
phases?: string[];
}
export interface PlaceAuthority {
name: string;
/** null when that authority has no page of its own — two English
* authorities hold fewer schools than the threshold. */
slug: string | null;
count: number;
}
export interface PlaceDetail {
place: PlaceSummary & {
parent_authority: string | null;
/** Every authority the place meaningfully sits in, largest first. SW19 is
* mostly Merton but partly Wandsworth. */
authorities?: PlaceAuthority[];
};
schools: School[];
averages: {
rwm_expected_pct: number | null;
attainment_8_score: number | null;
};
}
export function placeUrl(kind: string, slug: string, phase?: string): string {
const base =
kind === 'authority' ? `/schools/authority/${slug}`
: kind === 'outcode' ? `/schools/near/${slug}`
: `/schools/${slug}`;
return phase ? `${base}/${phase}` : base;
}
/**
* An authority name as it appears in a URL.
*
* Only a fallback: the API sends the slug it built, and that is what should
* be used. This mirrors `_slugify` in backend/app.py, collapsed runs and
* trimmed hyphens included, so the two cannot disagree about a name like
* "Bristol, City of".
*/
export function authoritySlug(name: string): string {
return name.toLowerCase().trim()
.replace(/[^\w\s-]/g, '')
.replace(/\s+/g, '-')
.replace(/-+/g, '-')
.replace(/^-|-$/g, '');
}
const API = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL
|| 'http://localhost:8000/api';
export async function fetchPlaces(): Promise<PlaceSummary[]> {
const res = await fetch(`${API}/places`, { next: { revalidate: 604800 } });
if (!res.ok) return [];
return (await res.json()).places ?? [];
}
export async function fetchPlace(
kind: string, slug: string, phase?: string,
): Promise<PlaceDetail | null> {
const q = phase ? `?phase=${encodeURIComponent(phase)}` : '';
const res = await fetch(`${API}/places/${kind}/${slug}${q}`,
{ next: { revalidate: 604800 } });
if (!res.ok) return null;
return res.json();
}
+39
View File
@@ -0,0 +1,39 @@
/**
* Client for /api/suggest.
*
* No `cache: "no-store"`. The compare modal's search uses it, and copying that
* here would discard both the browser cache and the ETag 304s the backend's
* CacheAndETagMiddleware already provides — on the one endpoint where prefix
* queries repeat most.
*/
export interface Suggestion {
urn: number;
school_name: string;
local_authority: string;
postcode: string;
phase: string;
school_type: string;
}
/** Below this the response is thousands of schools and worth no round trip. */
export const SUGGEST_MIN_QUERY = 2;
const API = process.env.NEXT_PUBLIC_API_URL || '/api';
/** Suggestions for `q`. Never throws: no suggestions is a fine outcome. */
export async function fetchSuggestions(
q: string, signal?: AbortSignal,
): Promise<Suggestion[]> {
if (q.trim().length < SUGGEST_MIN_QUERY) return [];
try {
const res = await fetch(`${API}/suggest?q=${encodeURIComponent(q.trim())}`,
{ signal });
if (!res.ok) return [];
const body = await res.json();
return body.suggestions ?? [];
} catch {
// Includes AbortError, which is the normal path on every keystroke.
return [];
}
}
+6 -1
View File
@@ -361,7 +361,12 @@ export interface SchoolDetailsResponse {
* held back as a paid feature and are not part of this public payload — see
* data_loader._admission_distance.
*/
admission_distance: SchoolAdmissionDistance | null;
/**
* Absent — not null — when the admission_distance flag is off. Null means
* "this school has no published cut-off"; absent means "cut-offs are not
* being published at all". They are different claims and the type says so.
*/
admission_distance?: SchoolAdmissionDistance | null;
deprivation: SchoolDeprivation | null;
finance: SchoolFinance | null;
}
+31
View File
@@ -55,6 +55,37 @@ const nextConfig = {
// Headers for caching and security
async headers() {
return [
{
/*
* Keep non-production hosts out of the index.
*
* Staging serves the same image as production off stx., so without
* this it is a full crawlable duplicate of the site.
*
* X-Robots-Tag, NOT a robots.txt Disallow. Disallow blocks crawling,
* which is not the same as blocking indexing — a disallowed URL can
* still be indexed from external links, and worse, blocking the crawl
* means Google never fetches the page and never sees a noindex at all.
* Staging therefore stays crawlable and answers "noindex" when crawled.
*
* Matched on the staging host explicitly rather than "any host that is
* not production". The inverted form is tempting because it would cover
* future environments automatically, but its failure mode is
* deindexing production if the Host header ever arrives rewritten by a
* proxy. This form's failure mode is a new environment being indexable
* until someone adds it here — recoverable, where the other is not.
*
* Any new non-production hostname must be added to this list.
*/
source: '/:path*',
has: [{ type: 'host', value: 'stx.schoolcompare.co.uk' }],
headers: [
{
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
},
],
},
{
source: '/:path*',
headers: [
@@ -0,0 +1,16 @@
locality_slug,locality_name,outcodes,region
battersea,Battersea,SW11,London
canary-wharf,Canary Wharf,E14,London
clapham,Clapham,SW4,London
shoreditch,Shoreditch,EC2A|E1,London
peckham,Peckham,SE15,London
brixton,Brixton,SW2|SW9,London
camden-town,Camden Town,NW1,London
wimbledon,Wimbledon,SW19,London
putney,Putney,SW15,London
fulham,Fulham,SW6,London
chiswick,Chiswick,W4,London
stratford,Stratford,E15,London
walthamstow,Walthamstow,E17,London
tooting,Tooting,SW17,London
dulwich,Dulwich,SE21|SE22,London
1 locality_slug locality_name outcodes region
2 battersea Battersea SW11 London
3 canary-wharf Canary Wharf E14 London
4 clapham Clapham SW4 London
5 shoreditch Shoreditch EC2A|E1 London
6 peckham Peckham SE15 London
7 brixton Brixton SW2|SW9 London
8 camden-town Camden Town NW1 London
9 wimbledon Wimbledon SW19 London
10 putney Putney SW15 London
11 fulham Fulham SW6 London
12 chiswick Chiswick W4 London
13 stratford Stratford E15 London
14 walthamstow Walthamstow E17 London
15 tooting Tooting SW17 London
16 dulwich Dulwich SE21|SE22 London
+1 -1
View File
@@ -12,4 +12,4 @@ slowapi==0.1.9
secure==0.3.0
typesense==0.21.0
numpy==1.26.4
UnleashClient==6.0.1