Compare commits

..
Author SHA1 Message Date
TudorandClaude Opus 5 413d86cc3c chore(flags): wire UNLEASH_URL and the SDK cache volume into the stacks
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 10s
PR Checks / Build Backend (no push) (pull_request) Successful in 27s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m59s
Both variables default to empty, so an environment without Unleash has
every flag off — the correct dark state rather than a boot failure.

The cache volume is the mitigation for the one real regression risk in
this design: the SDK evaluates everything False until it syncs, so a
backend cold-starting with an empty cache while Unleash is unreachable
would make a *released* feature disappear. On a named volume the disk
cache survives a restart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:58:29 +01:00
TudorandClaude Opus 5 4f01fbdedb test(e2e): make the distance journeys fail loudly, not skip quietly
The existing distance journeys all skip when no school has a published
figure, which is right when the feature is off — and wrong when it is
supposed to be on and is silently broken, because that shows up as a
green run full of skips. The new gate fails in exactly that case.

Feature state is read from the data, not from /api/flags: the public
proxy denies that path on purpose, since it names unreleased features.
Presence of the admission_distance key is the observable effect.

Verified against staging, where the feature is currently on: the on-gate
passes, the off-gate skips, the existing eight distance journeys are
unaffected.

One honest caveat — the /api/flags check passes on staging today because
that image predates the endpoint, not because the denylist works. The
denylist itself is covered by the jest unit test; this is defence in
depth and becomes a real assertion once deployed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:57:34 +01:00
TudorandClaude Opus 5 c3ba7aae0d feat(flags): ship last-distance-offered dark behind a flag
One gate, at the source. The frontend needs no change: DistanceSection
already returns null when distance_m is missing, and the admissions
block already conditions on (admissions || admissionDistance). Only 57
local authorities publish cut-offs, so the off-path is the commonest
path on the site and is well covered already.

Absent, not null. /api/schools/ is public and unauthenticated, so a
field left in the payload is a published field — the reasoning already
recorded in c9a1892 when history was withheld. The two are also
different claims: null says this school has no cut-off, absent says
cut-offs are not being published at all. The frontend type now says so.

The feature is on main and live on staging and has never reached
production, which is what makes it the right first consumer: the flag
lets the code promote without the feature appearing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:56:33 +01:00
TudorandClaude Opus 5 54a30de0d8 feat(flags): server-side getFlags for the frontend
Ships without a consumer, deliberately. The first flag needs none — the
backend withholds the field and the page follows — but 'UI elements on
existing pages' is one of the three surfaces this capability exists for,
and a flag layer that cannot gate one is incomplete.

Never throws: an unreadable flag is a dark one, which matches the
backend's fail-closed default. A page that 500s because the flags
endpoint blinked would be a worse outcome than a hidden feature.

Reading flags pins the calling route to a 300s ISR floor, since Next
takes the lowest revalidate among a route's fetches. That matches what
/school/[slug] already sits at, and it is the same property that makes a
flip propagate without a webhook.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:55:32 +01:00
TudorandClaude Opus 5 c30ad1db07 feat(flags): serve /api/flags, and keep the public proxy off it
The endpoint and its exposure control ship together on purpose. The
moment /api/flags exists, app/api/[...path] forwards it — and the
response names every unreleased feature the codebase knows about, along
with whether it is on. Publishing that is the opposite of shipping dark.

Denied on an exact first-segment match, not a prefix, so /api/flagship
does not go down with /api/flags. Next reads the endpoint server-side
over the Docker network, which never transits the public proxy.

jest.setup.js now guards its browser globals. It runs for every suite,
including the one that declares @jest-environment node to exercise the
route handler — NextRequest needs Fetch API globals jsdom lacks, and
there is no window there to define matchMedia on.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:54:26 +01:00
TudorandClaude Opus 5 7424cef7c6 feat(flags): the registry and a fail-closed Unleash client
Unleash holds flag state; it does not hold the list of flags. REGISTRY
is that list, because the SDK evaluates an unknown flag to False and
without a registry that is an undeclared False — indistinguishable from
a typo in a flag name.

Fail-closed throughout, and never raises: an unset UNLEASH_URL, an
unreachable server, a client that throws, an undeclared name — all
False. A flag layer that can 500 a request path or stop the API booting
is worse than one that is switched off.

Every flag defaults to False, with no per-flag override, because a flag
that defaults on is a kill switch and this is deliberately not one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:52:55 +01:00
TudorandClaude Opus 5 01ccbb8e82 feat(flags): add the Unleash stack and its runbook
Its own Portainer stack, belonging to neither application stack: a
staging redeploy must not be able to disturb production's flag state.

One instance serves both. OSS Unleash ships development and production
environments with environment-scoped client tokens, so the same flag
holds independent state in each — which is what lets a feature be on in
staging, where the E2E journeys exercise it, while production stays dark.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:51:35 +01:00
TudorandClaude Opus 5 c339c2f1a1 docs(flags): implementation plan, eight tasks
Task 1 is the Unleash stack and ends with a human step — the Portainer
deploy and the token generation cannot be automated from here. Nothing
else blocks on it: an unset UNLEASH_URL means every flag is False, which
is what local development and CI get, so the whole suite runs without a
flag server existing.

Self-review caught three defects in the plan itself. get_supplementary_data
takes (db, urn), not (urn), and the test DataFrame was minimised to the
point where the endpoint would have failed for reasons unrelated to
flags — both now copy the known-good shape from test_school_details.py.
The proxy test needs the node jest environment, since NextRequest wants
Fetch API globals jsdom does not provide. And the e2e off-state check
hardcoded a URN, so a 404 page would have satisfied it without proving
anything.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:46:15 +01:00
TudorandClaude Opus 5 e2ca3d79f9 docs(flags): drop the webhook — the seven-day premise was wrong
Next uses the LOWEST revalidate among a route's fetches, not the segment
value. School pages fetch school details at 300s and place pages fetch
national averages at 3600s, so the effective ISR period is five minutes
and one hour respectively — not the seven days the segment declares.

A flag flip therefore propagates on its own, well inside the monthly,
by-hand cadence these flags are for. That deletes two webhook
integrations, a revalidate route, a secret-in-query-string scheme, an
idempotency requirement, and the rule that every fetch carry a cache
tag — which was the part most likely to rot as fetches are added.

Two constraints survive: a flag must never gate content on a
force-static page, because app/admissions never revalidates; and a
route-family flag must rebuild the sitemap, deferred with the route case
since no flag in scope touches it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 10:40:35 +01:00
TudorandClaude Opus 5 c2364bf09e docs(flags): design for a ship-dark feature flag layer
Unleash self-hosted in its own Portainer stack, with FastAPI holding the
only SDK and Next reading flags through a tagged fetch.

The two hard parts are consequences of putting flag state in a service
rather than the repo: main stops being the whole truth about what is on,
and a flag can now change without the deploy that would have cleared the
caches. A code-declared registry bounds the first; webhook-driven
revalidateTag handles the second.

Cache tagging is deliberately coarse — every server fetch carries the
flags tag, not just the flags fetch itself. The first consumer proves
why: admission_distance changes the shape of /api/schools/{urn}, so a
narrow purge would leave ~25,000 school pages serving the pre-flip
render for a week, invisibly.

First consumer is the last-distance-offered feature, which is on main
and staging and has never reached production. It needs one gate, at the
API, because the frontend already no-ops on a missing field.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-23 09:47:58 +01:00
tudor 43e0621728 Merge pull request 'fix(places): phase links must stay in their own namespace' (#124) from fix/place-phase-links into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m31s
Reviewed-on: #124
2026-08-22 17:33:12 +00:00
TudorandClaude Opus 5 d1358cc00f fix(places): phase links must stay in their own namespace
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m37s
Every place page built its phase links as /schools/[slug]/[phase], the
shape that belongs to towns alone.

On an authority page that pointed into the town namespace. For 87 of the
151 authorities the target does not exist and the link 404s; for the
other 64 it resolves to the town of the same name — a different set of
schools, which is precisely the near-duplicate the two namespaces were
introduced to prevent. On an outcode page it 404s outright.

Two causes behind it, both a rule written twice and inherited by only
one of the places that needed it.

The authority phase route was in the spec and dropped by the plan, which
built the three bare routes and no fourth. The sitemap is generated from
the place registry, which was right about them all along, so 302
authority phase URLs have been submitted to Google and every one 404s.
Adding the route makes the sitemap true and serves a real query —
admissions are authority-run, so "primary schools in Kent" is how a
parent searches before they have settled on a town.

The outcode variants were the opposite: the registry computed phases for
outcodes although the spec gives them no route, and the sitemap knew to
skip them while the API did not. The registry now decides alone, and the
sitemap's duplicate of that rule is gone.

Also: an authority under the five-school threshold has no page, so the
API sends a null slug for it and the page names it without linking.
Two English authorities are in that position. It was unreachable in
today's data — verified across the EC and TR outcodes — but the thin
place redirect would have sent a reader to a 404 the year it isn't.

The e2e journey now walks every /schools link a page of each family
emits and requires a 200, which is the check that was missing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 17:22:11 +01:00
tudor 865a69b54d Merge pull request 'feat(places): list schools alphabetically on place pages' (#123) from feat/place-alphabetical-sort into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 20s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m27s
Reviewed-on: #123
2026-08-21 23:22:49 +00:00
TudorandClaude Opus 5 9cc87c41bb fix(places): a phase page needs results, not merely publishable schools
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m4s
Asked where schools with no results should sit in an alphabetical list, and
found that some pages were almost entirely made of them.

The per-phase threshold counted schools that were publishable — a result OR
an Ofsted grade — while a phase page exists for its results column.
/schools/kent/primary published with none of its five rows carrying a result;
Minehead had one of seven, Buntingford one of five. Forty-four phase pages
were majority-blank.

It is the same rule as "no page without a local average", which was written
into the spec as a thin-page control and never extended per phase.

The threshold now counts schools with a result for that phase. It gates
whether the page exists; it does not filter rows — a page that publishes still
lists every school of the phase, because someone looking up a school by name
has to find it whether or not it published results.

126 of 1,012 variant pages stop publishing: 62 primary, 64 secondary. Every
one of them was a table with too little in it to be worth a page.

The ordering itself is unchanged: pure A-Z, blanks interleaved. A school sits
where its name says it does, and at roughly a tenth of rows that reads fine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:14:31 +01:00
TudorandClaude Opus 5 8967966eef feat(places): list schools alphabetically on place pages
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m21s
Someone on a place page is usually looking for a school they can name, so the
order should serve scanning for it rather than ranking. /api/rankings keeps
its league-table ordering; this is a place-page decision, not a site-wide one.
Sorted case-insensitively, or a capitalised name would sort ahead of every
lowercase one.

The change made five pieces of copy untrue, so they go with it. The phase
variant titled itself "— Ranked", and all four route families described
themselves as "ranked by SATs and GCSE results". A page that opens by claiming
an order it does not keep is worse than one that claims nothing.

The ItemList markup carried `position` with no declared order, which reads as
a ranking. It now declares ItemListOrderAscending, so the structured data says
what the table does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:09:58 +01:00
tudor 4a9a5c734b Merge pull request 'fix(e2e): three assertions that were wrong about correct behaviour' (#122) from fix/e2e-canonical-and-robots into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m26s
Reviewed-on: #122
2026-08-21 23:07:57 +00:00
TudorandClaude Opus 5 4e82e6c916 fix(e2e): three assertions that were wrong about correct behaviour
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 41s
The staging gate was red on three journeys. All three were faults in the
tests; the site was behaving correctly in each case.

Next normalises canonical URLs against trailingSlash:false, so the homepage
ships "https://www.schoolcompare.co.uk" with no slash while every other route
keeps its path. Both address the same document. The test hardcoded the slash
and so failed only on the root — /rankings and /admissions passed throughout,
which is what made it look like a homepage bug rather than a test bug.
Compared with trailing slashes stripped from both sides.

The robots.txt assertion matched "Disallow: /" anywhere in the file and
tripped over the AI-crawler groups Cloudflare injects — ClaudeBot, GPTBot,
Amazonbot and six others all carry a blanket disallow, deliberately, and none
of them is Googlebot. It now parses the file into user-agent groups and checks
only the "*" group, which is also the thing the test was always trying to say:
Google may crawl the page, so it can see the noindex header.

Both were the same mistake as the doubled brand: asserting a naive string
rather than the semantics, and asserting against what the code assembles
rather than what the page renders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 23:51:34 +01:00
tudor d4340a8fdd Merge pull request 'feat(places): name every authority a place sits in' (#121) from feat/place-multiple-authorities into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #121
2026-08-21 21:56:50 +00:00
TudorandClaude Opus 5 bb2f7a5841 fix(places): address review, and merge places GIAS spells more than one way
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m34s
Two findings from review on #121, plus a third the review prompted.

The cap at three authorities silently dropped the fourth in exactly the case
where the information matters most — a genuinely fragmented place — and
contradicted the stated goal of naming every authority a place sits in. It is
gone. The share rule was always the real limit and already bounds the list at
ten. Measured against the live corpus, one town would have been truncated
today: LONDON, split evenly between Hackney, Lambeth, Westminster and
Lewisham.

parent_authority used mode() while authorities used value_counts(), and on an
exact tie pandas does not guarantee the two pick the same name, so the 301
could have pointed somewhere other than the authority named first on the page.
The parent is now derived from authorities[0]: one computation, one answer.
It also inherits the sentinel filter, so a place can no longer redirect to
/schools/authority/does-not-apply.

Chasing the truncation case surfaced a worse bug. Places were grouped by raw
town value, but the registry is keyed by slug, and GIAS spells the same place
several ways. Five town slugs come from more than one spelling: "London"
(1,819 schools) and "LONDON" (12) both slugify to `london`, so the later group
simply overwrote the earlier one — /schools/london could have shown twelve
schools, silently, depending on row order. Weston-super-Mare was split 14/19
across two spellings and Newcastle-under-Lyme across three. Grouping is now by
slug, and the display name is the most common spelling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:42:18 +01:00
TudorandClaude Opus 5 1cb5314c53 feat(places): name every authority a place sits in
SW19 is mostly Merton but partly Wandsworth, and the page said only Merton.
The cause was one field doing two jobs: _parent_authority takes the modal
authority, which is right for a 301 target and wrong as a statement about
where a place is.

This is not a corner case. A quarter of viable outcodes (425 of 1,760) and a
third of viable towns (263 of 783) cross an authority boundary — Bedford the
town spans Bedford and Central Bedfordshire.

Place now carries `authorities`, every authority holding at least a tenth of
the schools and at least two of them, largest first. parent_authority stays
single and unchanged, because a redirect still needs one target.

The share threshold exists because GIAS carries postcode errors: EN6 lists two
Shropshire schools among fourteen in Hertfordshire, and a bare "any authority
present" rule would print those as though they were real. A place too small or
too fragmented to clear the threshold still names its largest, so the page
never goes silent about where it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:40:58 +01:00
tudor 4cea26b813 Merge pull request 'fix(places): align the measure column's heading with its values' (#120) from fix/place-table-alignment into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #120
2026-08-21 21:21:03 +00:00
TudorandClaude Opus 5 dbb74d9b60 fix(places): align the measure column's heading with its values
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 14s
The heading sat on the right edge of the column and every value on the left.
A specificity collision, not a layout problem: the two were aligned by
different selectors and only one of them won.

  .table td            (0,1,1)  text-align: left    <- won for the value
  .num                 (0,1,0)  text-align: right   <- lost
  .table th:last-child (0,2,1)  text-align: right   <- won for the heading

The heading and the value cell now share one class and one rule, so they
cannot drift apart again whatever else changes around them.

The column also stretched to half the table. It now hugs its content with
width:1% and nowrap, so the school name takes the remaining width — which is
what made the gap read as misalignment on a wide screen, and what crowded the
name column on a narrow one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:15:02 +01:00
tudor 9545aec7f4 Merge pull request 'fix(places): phase-grouped tables, plain-English measures, styled links' (#119) from fix/place-presentation-to-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 56s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #119
2026-08-21 20:57:13 +00:00
Tudor 3365ebcb3a fix(places): phase-grouped tables, plain-English measures, styled links
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Three presentation faults on the place pages, all found by looking at a
rendered page rather than at a test.

An unphased place page showed one primary-only measure for a list holding both
phases: 8 of 27 rows on /schools/brentwood were blank, because secondaries
have no reading-writing-maths score. Picking the other measure would only have
inverted which rows were empty, and putting both in one column would have
mixed a percentage with a 0-90 score. Each phase now gets its own table, so a
blank cell means the school genuinely has no published result — which is worth
saying, and now says "Not published" rather than a bare dash.

"RWM expected" was invented here. The site already names the measure in
METRIC_DEFINITIONS, surfaced at /api/metrics: "Reading, Writing & Maths
Combined %". The heading now reads "Reading, writing & maths" with the full
definition in the tooltip.

Links carried no class at all, so they rendered as default blue underlined
browser links beside a site that styles table links as body colour with a
brand hover. They now follow RankingsView's convention, and running-copy links
take the brand colour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 21:56:06 +01:00
tudor 6d79bd3331 Merge pull request 'fix(places): stop the place titles doubling the brand' (#117) from fix/place-title-brand-doubling into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #117
2026-08-21 20:53:14 +00:00
TudorandClaude Opus 5 24e114dee7 fix(places): stop the place titles doubling the brand
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 27s
Every place page shipped as 'Schools in Brentwood - Compare 27 Schools |
schoolcompare | schoolcompare'. The root layout's title template appends
'| schoolcompare' to any plain-string title, and all four place routes already
carried the brand. W8 opted the other routes out with an absolute title; the
place routes were written afterwards and did not inherit the lesson.

~2,600 titles affected, and the repetition pushed them past Google's
truncation point, so the doubled brand displaced real words in the result.

An e2e journey now asserts no title repeats the brand, across the static
routes and a place page, so this cannot come back on a route added later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 21:43:32 +01:00
tudor 6c5db0c266 Merge pull request 'fix(places): submit and link the phase variants' (#116) from fix/place-phase-variants into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m29s
Reviewed-on: #116
2026-08-21 19:46:55 +00:00
TudorandClaude Opus 5 6f749ed21f fix(places): submit and link the phase variants
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m39s
/schools/[place]/[phase] shipped as routes but reached nothing. The sitemap
emitted one URL per registry entry and the registry had no phase dimension, so
~950 pages were absent from every sitemap — and PlaceView did not link them
either, leaving them reachable by nothing at all.

That is the query shape the baseline actually showed: 'primary schools in
beccles', 'secondary schools in brentwood'. Publishing the routes without a
path in meant building for the demand and then hiding from it.

Place now carries phase_urns so the per-phase threshold can be applied without
re-querying, the sitemap emits a variant wherever a phase clears the threshold
on its own, and the API exposes the qualifying phases so the place page links
only variants that exist. Outcodes are excluded: nobody searches 'primary
schools in SW11' and those routes do not exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 20:43:07 +01:00
tudor d423826840 Merge pull request 'fix(places): a locality collision must not break the sitemap' (#115) from fix/locality-collision-skip into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m38s
Reviewed-on: #115
2026-08-21 19:26:02 +00:00
TudorandClaude Opus 5 d3c63ccc6d fix(places): a locality collision must not break the sitemap
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 34s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 34s
Sitemap regeneration failed on staging. 'richmond' in the curated locality
list collides with the GIAS town Richmond in North Yorkshire (37 schools), the
registry raised, and the admin endpoint 500d — taking down sitemap generation
for all 25,000 school pages over one bad row of curated data.

The guard now skips the colliding locality and logs an error. Skipping still
achieves what the guard was for — a locality never silently shadows a town —
without letting curated data break the site. That matters beyond this bug:
GIAS town names change with no code change here, so a raise could fire
spontaneously in production later.

Also removes four localities that were London boroughs rather than districts.
Hackney, Islington, Greenwich and Ealing are local authorities with 104, 72,
108 and 115 schools and already have authority pages; a locality defined by
two or three outcodes would have been a partial near-duplicate of one — the
thin-content failure the two-namespace design exists to avoid. A test now
guards the whole borough list.

Validated against the live corpus: 15 localities, no town collisions, no
authority duplicates, all 15 clear the threshold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 20:20:52 +01:00
tudor b93eb3a691 Merge pull request 'feat(seo): the location layer — town, locality, authority and outcode pages (W2)' (#114) from feat/w2-location-layer into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m18s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 4s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 11m8s
Reviewed-on: #114
2026-08-21 17:56:05 +00:00
TudorandClaude Opus 5 6b871ce1e9 feat(places): ItemList and BreadcrumbList, and the e2e gate
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 4m33s
ItemList tells Google the page is a ranked set rather than prose;
BreadcrumbList puts the place in a hierarchy. School URLs in the markup are
absolute on the canonical host, since a relative URL in JSON-LD is ambiguous.

Eight journeys covering all four families, the two-namespace guarantee, the
threshold, the canonical, the sitemap and the local-versus-England line — the
last because that comparison is the reason these pages are not a name dropped
into a template.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:19:08 +01:00
TudorandClaude Opus 5 c981d89137 feat(places): town, locality, authority and outcode routes
Every generateStaticParams is gated behind PRERENDER_PLACES and wrapped in the
same try/catch the school route uses. The plan claimed authority pages were
'few enough to always prebuild' — but few enough still means the API must be
reachable at build time, and in CI it is not: the build failed with
ECONNREFUSED rather than degrading to ISR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:17:57 +01:00
TudorandClaude Opus 5 de5e790112 feat(places): place page client and view component
One component for all four families: they differ in what fills the registry,
not in what the page shows, so a second would be a second place to forget the
same change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:15:10 +01:00
TudorandClaude Opus 5 42138fc402 feat(places): submit place and outcode sitemaps
Separate children per family so Search Console reports the location layer's
indexation apart from the school pages' — which is the point of the index
built in W1, and the number the stop condition watches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:13:55 +01:00
TudorandClaude Opus 5 c5af476213 feat(places): /api/places registry and place detail endpoints
The registry is cached for the process and reset by the same admin endpoint
that rebuilds the sitemaps, so places and sitemap always describe the same
corpus rather than drifting apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:13:05 +01:00
TudorandClaude Opus 5 de853b90b3 feat(places): London localities and postcode districts
The GIAS town field puts 1,819 London schools under the single value
'London', so it cannot answer 'schools in Battersea' — a query that appears in
the baseline. No single field can: parliamentary constituency gives Battersea
but not Canary Wharf, admin_ward gives Canary Wharf but not Battersea, and
neither gives Clapham or Shoreditch. So a locality is curated, defined by the
postcode districts it covers, which needs no new ingestion.

A locality may not shadow a published town: the registry raises rather than
silently costing a page that carries real demand. One below the threshold is
logged rather than raising, because a locality can legitimately be too small.

The pipeline seed mirrors the module, with a test guarding the drift — the
same arrangement gias_codes has, and for the same reason: the backend image
does not contain pipeline/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:11:56 +01:00
TudorandClaude Opus 5 759d9f5cea feat(places): registry of towns and authorities
Two namespaces because 67 town names collide with an authority name and
neither set contains the other — postal towns cross authority boundaries, so
Bedford the town holds 104 schools against the authority's 86.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:10:37 +01:00
TudorandClaude Opus 5 555d3f0a7d docs(seo): implementation plan for the W2 location layer
Seven tasks: the place registry, London localities and outcodes, the places
API, per-family sitemaps, the shared place view, the four route families, and
structured data plus the e2e gate.

Two things the plan corrects against the spec. The backend image does not
contain pipeline/, so the curated locality list cannot live only in a dbt
seed — it follows the gias_codes.py precedent instead, canonical in backend
with the seed as a mirror. And NationalAverages is nested by phase rather than
flat, which the first draft read wrongly and would have rendered every page
without its England comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:09:35 +01:00
TudorandClaude Opus 5 ecc847091c docs(seo): design for the W2 location layer
Supersedes the original spec's W2. The Search Console baseline inverted its
ordering: every measured location query is town or district level, none is an
administrative area, and phase is part of the query rather than a filter.

Two problems the original design did not anticipate. 67 viable towns share a
name with a local authority, and the authority is the larger set in only 43 of
them — postal towns cross authority boundaries, so neither can absorb the
other. Two namespaces resolve it by construction. And the GIAS town field
collapses 1,819 London schools into one value, which a curated
locality-to-outcode seed solves without new ingestion.

Sizing is measured against the live 25,185-school corpus rather than
estimated: 783 viable towns, 1,760 outcodes, 154 authorities.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:09:35 +01:00
tudor b187a478c9 Merge pull request 'feat(seo): rewrite the C1 snippets to earn the click (W8)' (#113) from feat/seo-metadata-c1 into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #113
2026-08-20 23:26:12 +00:00
TudorandClaude Opus 5 c0547c45e5 feat(seo): rewrite the C1 snippets to earn the click (W8)
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 11s
The baseline says these pages already rank and are not clicked. 'compare
school performance' sits at position 6.1 with 0.43% CTR; 'compare schools' at
7.2 with 0.87%. The brand query 'school compare' draws 9.16% from the same
neighbourhood of the same results page, which rules out a ranking explanation
— when the snippet gives a reason to click, it gets clicked.

These SERPs are owned by the DfE's own 'Compare school performance' service.
The old title put a lowercase brand nobody searches for in the most valuable
pixels, then a near-paraphrase of that service's name. Beside the government's
own result it read as a lookalike.

Intent in the title, differentiator in the description. Titles now match what
people type, and the descriptions carry the one fact gov.uk does not publish:
how close you had to live to get a place.

/compare deliberately takes the tool phrasing rather than the homepage's, so
the two pages stop competing for one phrase. The root layout's default and
Open Graph copy were saying something different again; they now agree.

No hard school counts in any of it. The corpus moves with every data refresh
and this repo has already shipped one copy bug of that kind.

Tests guard the mechanics — SERP length, intent keyword, the differentiator,
no brand-first title — and leave the wording free to iterate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:21:08 +01:00
tudor 4a3928df9f Merge pull request 'fix(seo): a school is publishable on any year's results, not the latest' (#112) from fix/sitemap-any-year-data into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #112
2026-08-20 23:15:19 +00:00
TudorandClaude Opus 5 07c97a46c5 fix(seo): a school is publishable on any year's results, not the latest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 53s
_school_sitemap_rows tested only the latest year's row, which quietly dropped
every school with results in its history but a null row for the most recent
year — a school that stopped reporting, or whose figures were suppressed for
small-cohort disclosure.

The Mallard Academy (150367) is the case that caught it: real KS2 results for
2015-16 through 2018-19, then null rows from 2022-23 on. Its detail page shows
all four years; the sitemap omitted it. Sampling 40 of the 2,206 excluded
schools found 4 like this, so roughly 220 real pages were being withheld.

Publishable is now a property of the school, computed across every row, while
lastmod still comes from the latest row so the most recent Ofsted date wins.
The field list is a module constant shared with _has_publishable_data so the
two checks cannot drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:04:32 +01:00
tudor bb81337aba Merge pull request 'fix(seo): keep staging out of the search index' (#111) from fix/staging-noindex into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m1s
Reviewed-on: #111
2026-08-20 22:39:38 +00:00
tudor f928a15c1e Merge pull request 'fix(seo): crawl hygiene and a per-family sitemap index (W1)' (#110) from feat/seo-crawl-hygiene-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m21s
Reviewed-on: #110
2026-08-20 22:23:10 +00:00
TudorandClaude Opus 5 2208ad93c1 feat(seo): split the sitemap into a per-family index
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m2s
Search Console reports coverage per submitted sitemap, so one file per page
family is what will make W2's location pages measurable when they land. The
index's lastmod is generation time, which is the correct semantic there —
unlike on a <url>, where it would be a claim we cannot support.

Children sit under /sitemaps/ because Next only treats a whole bracketed path
segment as dynamic; a route folder named sitemap-[...parts] would be read as a
literal static segment and never match. Confirmed by the build output, which
lists /sitemaps/[...parts] as a dynamic route.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
TudorandClaude Opus 5 24f3cb4c65 fix(seo): submit only school pages that have something to show
Drops the schools with neither results nor an Ofsted grade, adds /admissions
which was never listed, replaces the invented priority and changefreq with a
lastmod taken from each school's Ofsted date.

lastmod is omitted where no date is known rather than defaulted to now. An
always-now lastmod is a claim Google learns to distrust; absent honestly
means unknown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
TudorandClaude Opus 5 3816b92d06 fix(seo): noindex parameterised comparisons, keep bare /compare
25,193 schools make ~317 million pairs. The bare page stays indexable as the
landing page for the head term; the parameter space goes noindex, follow so
its outbound links still count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
TudorandClaude Opus 5 51ce2d6373 fix(seo): declare a canonical on every route
The homepage read eleven search params and declared no canonical, so every
filter combination was a crawlable near-duplicate of the page we most want to
rank. Rankings and admissions declared none either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
TudorandClaude Opus 5 5ddc7314fd fix(seo): canonicalise on the www host, which is the one that serves 200
The apex 301s to www at Cloudflare, but metadataBase, the school-page
canonical, robots.txt's Sitemap: line and the sitemap's own <loc> entries all
named the apex. Every one of those pointed Google at a redirect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:15:00 +01:00
50 changed files with 7777 additions and 112 deletions

No files matched your search

+343 -51
View File
@@ -7,6 +7,7 @@ Uses real data from UK Government Compare School Performance downloads.
import hashlib
import re
from contextlib import asynccontextmanager
from datetime import datetime, timezone
from typing import Optional
import numpy as np
@@ -34,6 +35,8 @@ from .data_loader import (
search_schools_typesense,
)
from .data_loader import get_data_info as get_db_info
from . import flags
from .places import build_place_registry
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .utils import clean_for_json, convert_to_native
@@ -48,11 +51,20 @@ PHASE_GROUPS: dict[str, set[str]] = {
"all-through": {"all-through"},
}
BASE_URL = "https://schoolcompare.co.uk"
# Must match SITE_URL in nextjs-app/lib/site.ts. The apex 301s to www, and a
# sitemap <loc> that redirects wastes a crawl on every URL it lists.
BASE_URL = "https://www.schoolcompare.co.uk"
MAX_SLUG_LENGTH = 60
# In-memory sitemap cache
_sitemap_xml: str | None = None
# In-memory sitemap cache: name -> XML. Populated on startup and by the admin
# regenerate endpoint after a pipeline run.
_sitemaps: dict[str, str] | None = None
# Built from the same DataFrame the sitemap uses, so places and sitemap can
# never describe different corpora. Reset by the same admin endpoint.
_place_registry: dict | None = None
VALID_PLACE_KINDS = ("town", "locality", "authority", "outcode")
def _slugify(text: str) -> str:
@@ -70,43 +82,206 @@ def _school_url(urn: int, school_name: str) -> str:
return f"/school/{urn}-{slug}"
def build_sitemap() -> str:
"""Generate sitemap XML from in-memory school data. Returns the XML string."""
# Routes worth submitting that are not a school page. /admissions was missing
# from the sitemap entirely despite being a static, indexable guide.
STATIC_SITEMAP_PATHS = ("/", "/rankings", "/compare", "/admissions")
# A page has something a search result could state if any of these is present
# in any year. Shared by _has_publishable_data and the per-school check in
# _school_sitemap_rows so the two can never drift.
_PUBLISHABLE_FIELDS = ("rwm_expected_pct", "attainment_8_score", "ofsted_grade")
def _has_publishable_data(row) -> bool:
"""True when a school page has something a search result could state.
A school with no results in any year and no Ofsted grade renders an empty
page. Submitting it spends crawl budget and drags the corpus-wide quality
signal down, so it stays out of the sitemap. The page itself still resolves
for anyone who has the URL.
"""
for field in _PUBLISHABLE_FIELDS:
value = row.get(field)
if value is not None and not pd.isna(value):
return True
return False
def _url_element(loc: str, lastmod: str | None = None) -> str:
"""One <url> entry. No priority or changefreq — Google ignores both."""
body = f"<loc>{loc}</loc>"
if lastmod:
body += f"<lastmod>{lastmod}</lastmod>"
return f" <url>{body}</url>"
def _school_sitemap_rows(df) -> list[str]:
"""A <url> element per school that has something to show.
lastmod comes from the school's Ofsted date where there is one and is
omitted otherwise. An always-now lastmod is a claim Google learns to
distrust; an absent one honestly means "unknown".
"""
if df.empty or "urn" not in df.columns or "school_name" not in df.columns:
return []
rows: list[str] = []
seen: set[int] = set()
# Publishable is a property of the SCHOOL, not of its latest row.
#
# The first cut tested the latest year's row alone, which quietly dropped
# every school that has results in its history but a null row for the most
# recent year — a school that stopped reporting, or whose figures were
# suppressed for small-cohort disclosure. The Mallard Academy (150367) is
# the case that caught it: real KS2 results for 2015-16 through 2018-19,
# then null rows for 2022-23 onward. Its page shows all four years; the
# sitemap omitted it. Roughly 220 schools were affected.
publishable_cols = [c for c in _PUBLISHABLE_FIELDS if c in df.columns]
publishable: set[int] = (
set(df.loc[df[publishable_cols].notna().any(axis=1), "urn"].astype(int))
if publishable_cols else set()
)
# Latest row per URN first, so a school's most recent Ofsted date wins.
ordered = df.sort_values("year", ascending=False) if "year" in df.columns else df
for _, row in ordered.iterrows():
urn = int(row["urn"])
if urn in seen:
continue
seen.add(urn)
if urn not in publishable:
continue
lastmod = None
ofsted_date = row.get("ofsted_date")
if ofsted_date is not None and not pd.isna(ofsted_date):
lastmod = pd.Timestamp(ofsted_date).date().isoformat()
rows.append(_url_element(
BASE_URL + _school_url(urn, str(row["school_name"])), lastmod))
return rows
# Sitemaps cap at 50,000 URLs per file. 10,000 keeps a child small enough to
# scan by eye in Search Console, which is the point of splitting at all:
# coverage is reported per submitted sitemap, so one file per page family is
# what makes an indexation problem attributable to a family.
SITEMAP_CHUNK_SIZE = 10_000
# Children are served under /sitemaps/ because Next.js only treats a whole
# bracketed path segment as dynamic — a route folder named "sitemap-[...parts]"
# is read as a literal static segment and never matches.
SITEMAP_CHILD_PREFIX = "/sitemaps"
def get_place_registry() -> dict:
"""The place registry, built once and cached for the process."""
global _place_registry
if _place_registry is None:
_place_registry = build_place_registry(load_school_data())
return _place_registry
def _urlset(rows: list[str]) -> str:
return "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
*rows,
"</urlset>",
])
def _place_url(place) -> str:
"""The canonical path for a place. Two namespaces, per the spec.
Towns and localities share /schools/[place]; authorities take their own
prefix because 67 town names collide with an authority name and neither
set contains the other.
"""
if place.kind == "authority":
return f"/schools/authority/{place.slug}"
if place.kind == "outcode":
return f"/schools/near/{place.slug}"
return f"/schools/{place.slug}"
def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
Phase is part of the query — "primary schools in beccles" — so each
variant is its own indexable page. Submitting only the bare place URL left
~950 of them reachable by nothing: absent from every sitemap, and not
linked from the place page either.
"""
rows: list[str] = []
for p in sorted(get_place_registry().values(), key=lambda p: (p.kind, p.slug)):
if p.kind not in kinds:
continue
rows.append(_url_element(BASE_URL + _place_url(p)))
# Which phases a place publishes is the registry's decision alone —
# outcodes report none, because the spec gives them no phase route.
# Repeating that rule here was how the page and the sitemap came to
# disagree about which URLs exist.
for phase in ("primary", "secondary"):
if p.publishes_phase(phase):
rows.append(_url_element(f"{BASE_URL}{_place_url(p)}/{phase}"))
return rows
def build_sitemaps() -> dict[str, str]:
"""Build the sitemap index and every child, keyed by name."""
df = load_school_data()
static_urls = [
(BASE_URL + "/", "daily", "1.0"),
(BASE_URL + "/rankings", "weekly", "0.8"),
(BASE_URL + "/compare", "weekly", "0.8"),
children: dict[str, str] = {
"static.xml": _urlset(
[_url_element(BASE_URL + path) for path in STATIC_SITEMAP_PATHS]),
}
school_rows = _school_sitemap_rows(df)
# Always emit at least one school child, so the index shape is stable even
# on an empty database.
chunks = [school_rows[i:i + SITEMAP_CHUNK_SIZE]
for i in range(0, len(school_rows), SITEMAP_CHUNK_SIZE)] or [[]]
for n, chunk in enumerate(chunks, start=1):
children[f"schools-{n}.xml"] = _urlset(chunk)
# Separate children per family: Search Console reports coverage per
# submitted sitemap, which is how the location layer's indexation is
# measured apart from the school pages'.
for label, kinds in (("places", ("town", "locality", "authority")),
("outcodes", ("outcode",))):
rows = _place_sitemap_rows(kinds)
chunks = [rows[i:i + SITEMAP_CHUNK_SIZE]
for i in range(0, len(rows), SITEMAP_CHUNK_SIZE)] or [[]]
for n, chunk in enumerate(chunks, start=1):
children[f"{label}-{n}.xml"] = _urlset(chunk)
# On a sitemap index, lastmod means "when this sitemap file last changed",
# so generation time is the correct value here — unlike on a <url>, where
# it would be a claim about content we cannot support.
generated = datetime.now(timezone.utc).date().isoformat()
index_rows = [
f" <sitemap><loc>{BASE_URL}{SITEMAP_CHILD_PREFIX}/{name}</loc>"
f"<lastmod>{generated}</lastmod></sitemap>"
for name in children
]
index = "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
'<sitemapindex xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">',
*index_rows,
"</sitemapindex>",
])
return {**children, "sitemap.xml": index}
lines = ['<?xml version="1.0" encoding="UTF-8"?>',
'<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">']
for url, freq, priority in static_urls:
lines.append(
f" <url><loc>{url}</loc>"
f"<changefreq>{freq}</changefreq>"
f"<priority>{priority}</priority></url>"
)
if not df.empty and "urn" in df.columns and "school_name" in df.columns:
seen = set()
for _, row in df[["urn", "school_name"]].drop_duplicates(subset="urn").iterrows():
urn = int(row["urn"])
name = str(row["school_name"])
if urn in seen:
continue
seen.add(urn)
path = _school_url(urn, name)
lines.append(
f" <url><loc>{BASE_URL}{path}</loc>"
f"<changefreq>monthly</changefreq>"
f"<priority>0.6</priority></url>"
)
lines.append("</urlset>")
return "\n".join(lines)
def build_sitemap() -> str:
"""The sitemap index. Kept for `lifespan` and the admin endpoint."""
return build_sitemaps()["sitemap.xml"]
def clean_filter_values(series: pd.Series) -> list[str]:
@@ -284,7 +459,8 @@ def validate_postcode(postcode: Optional[str]) -> Optional[str]:
@asynccontextmanager
async def lifespan(app: FastAPI):
"""Application lifespan - startup and shutdown events."""
global _sitemap_xml
global _sitemaps
flags.init()
print("Loading school data from marts...")
df = load_school_data()
if df.empty:
@@ -294,9 +470,9 @@ async def lifespan(app: FastAPI):
# Pre-compute the latest-year snapshot so the first search request is fast
await asyncio.to_thread(load_latest_school_data)
try:
_sitemap_xml = build_sitemap()
n = _sitemap_xml.count("<url>")
print(f"Sitemap built: {n} URLs.")
_sitemaps = build_sitemaps()
n = sum(x.count("<url>") for x in _sitemaps.values())
print(f"Sitemaps built: {len(_sitemaps)} files, {n} URLs.")
except Exception as e:
print(f"Warning: sitemap build failed on startup: {e}")
@@ -632,7 +808,13 @@ async def get_school_details(request: Request, urn: int):
"census": supplementary.get("census"),
"admissions": supplementary.get("admissions"),
"admissions_history": supplementary.get("admissions_history") or [],
"admission_distance": supplementary.get("admission_distance"),
# Behind a flag, and withheld at the source rather than rendered-but-
# hidden: this endpoint is public and unauthenticated, so a field left
# in the payload is a published field. The key is absent, not null —
# null would state that this school has no cut-off, which is a
# different claim from "we are not publishing cut-offs".
**({"admission_distance": supplementary.get("admission_distance")}
if flags.is_enabled("admission_distance") else {}),
"sen_detail": supplementary.get("sen_detail"),
"phonics": supplementary.get("phonics"),
"deprivation": supplementary.get("deprivation"),
@@ -1007,6 +1189,100 @@ async def get_rankings(
}
@app.get("/api/places")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def list_places(request: Request):
"""Every published place. The sitemap and the link modules read this."""
registry = get_place_registry()
return {"places": [
{"kind": p.kind, "slug": p.slug, "name": p.name, "count": len(p.urns)}
for p in sorted(registry.values(), key=lambda p: (p.kind, p.slug))
]}
@app.get("/api/places/{kind}/{slug}")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_place(request: Request, kind: str, slug: str,
phase: Optional[str] = None):
"""One place: its schools ranked, and its local averages."""
if kind not in VALID_PLACE_KINDS:
raise HTTPException(status_code=404, detail="No such place")
registry = get_place_registry()
place = registry.get(f"{kind}:{slug}")
if place is None:
raise HTTPException(status_code=404, detail="No such place")
df = load_latest_school_data()
rows = df[df["urn"].isin(place.urns)]
if phase:
wanted = PHASE_GROUPS.get(phase.lower())
if wanted and "phase" in rows.columns:
rows = rows[rows["phase"].fillna("").str.lower().isin(wanted)]
# The metric the page shows, and averages.
metric = "attainment_8_score" if phase == "secondary" else "rwm_expected_pct"
# Alphabetical, not by score. A place page is read by someone looking for
# a school they can name, and scanning for it is what the order should
# serve. /rankings is where the league-table ordering lives, and it keeps
# sorting by metric.
if "school_name" in rows.columns:
rows = rows.sort_values("school_name", key=lambda c: c.str.lower())
averages = {
m: (None if m not in rows.columns or rows[m].dropna().empty
else float(rows[m].dropna().mean()))
for m in ("rwm_expected_pct", "attainment_8_score")
}
cols = [c for c in SCHOOL_COLUMNS + ["latitude", "longitude", "phase",
"rwm_expected_pct", "attainment_8_score",
"total_pupils"]
if c in rows.columns]
return {
"place": {"kind": place.kind, "slug": place.slug, "name": place.name,
"count": len(place.urns),
"parent_authority": place.parent_authority,
# Every authority the place meaningfully sits in. SW19 is
# mostly Merton but partly Wandsworth; naming one asserts
# something false.
#
# The slug is null where that authority has no page of its
# own: City of London and the Isles of Scilly hold fewer
# schools than the threshold. Naming them is still right;
# linking them would be a 404.
"authorities": [
{"name": name,
"slug": (_slugify(name)
if f"authority:{_slugify(name)}" in registry
else None),
"count": n}
for name, n in place.authorities
],
# Only phases that clear the threshold, so the page links
# variants that exist rather than 404s.
"phases": [ph for ph in ("primary", "secondary")
if place.publishes_phase(ph)]},
"schools": clean_for_json(rows[cols]),
"averages": averages,
}
@app.get("/api/flags")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_feature_flags(request: Request):
"""Every declared flag and its current value.
Internal only. The Next proxy denies this path, because the response names
every unreleased feature the codebase knows about — which is exactly what
shipping dark is meant to keep quiet.
"""
return flags.all_flags()
@app.get("/api/data-info")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_data_info(request: Request):
@@ -1087,16 +1363,28 @@ async def robots_txt():
return FileResponse(settings.frontend_dir / "robots.txt", media_type="text/plain")
@app.get("/sitemap.xml")
async def sitemap_xml():
"""Serve sitemap.xml for search engine indexing."""
global _sitemap_xml
if _sitemap_xml is None:
def _serve_sitemap(name: str) -> Response:
global _sitemaps
if _sitemaps is None:
try:
_sitemap_xml = build_sitemap()
_sitemaps = build_sitemaps()
except Exception as e:
raise HTTPException(status_code=503, detail=f"Sitemap unavailable: {e}")
return Response(content=_sitemap_xml, media_type="application/xml")
if name not in _sitemaps:
raise HTTPException(status_code=404, detail="No such sitemap")
return Response(content=_sitemaps[name], media_type="application/xml")
@app.get("/sitemap.xml")
async def sitemap_xml():
"""Serve the sitemap index."""
return _serve_sitemap("sitemap.xml")
@app.get("/sitemaps/{name}")
async def sitemap_child(name: str):
"""Serve a child sitemap (static.xml, or schools-N.xml)."""
return _serve_sitemap(name)
@app.post("/api/admin/regenerate-sitemap")
@@ -1106,10 +1394,14 @@ async def regenerate_sitemap(
_: bool = Depends(verify_admin_api_key),
):
"""Rebuild and cache the sitemap from current school data. Called by Airflow after data updates."""
global _sitemap_xml
_sitemap_xml = build_sitemap()
n = _sitemap_xml.count("<url>")
return {"status": "ok", "urls": n}
global _sitemaps, _place_registry
# Places and sitemap are rebuilt together — they read the same marts, and
# letting them drift apart would submit URLs for places that no longer
# exist.
_place_registry = None
_sitemaps = build_sitemaps()
n = sum(x.count("<url>") for x in _sitemaps.values())
return {"status": "ok", "urls": n, "sitemaps": len(_sitemaps)}
# Mount static files directly (must be after all routes to avoid catching API calls)
+10
View File
@@ -42,6 +42,16 @@ class Settings(BaseSettings):
typesense_url: str = "http://localhost:8108"
typesense_api_key: str = ""
# Feature flags (Unleash). An empty unleash_url disables flags entirely and
# every flag evaluates False — the correct behaviour for local development
# and CI, and the reason no test needs a running Unleash.
unleash_url: str = ""
unleash_api_token: str = ""
unleash_app_name: str = "schoolcompare-backend"
# On a named volume, so a restart during an Unleash outage keeps
# last-known state instead of reverting a released feature to dark.
unleash_cache_directory: str = "/app/.unleash"
# Analytics
ga_measurement_id: Optional[str] = "G-J0PCVT14NY" # Google Analytics 4 Measurement ID
+111
View File
@@ -0,0 +1,111 @@
"""Feature flags: what can be switched, and what is switched right now.
Ship-dark, not a kill switch. Flags let work merge and deploy without becoming
visible; they are expected to flip about monthly, by a person, deliberately.
Nothing here does percentage rollouts or user targeting — the site has no user
identity to target.
Unleash holds the state. It does not hold the list. REGISTRY below is that
list, and it exists for three reasons: the SDK evaluates an unknown flag to
False, so without a registry that is an *undeclared* False, indistinguishable
from a typo; /api/flags needs a key set to return when Unleash is unreachable;
and a flag in the UI but not in the registry is orphaned and should be visibly
so rather than quietly authoritative.
Every flag defaults to False. There is no per-flag default, because a flag that
defaults on is a kill switch, and this is not one.
"""
from __future__ import annotations
import logging
from dataclasses import dataclass
from datetime import date
from .config import settings
logger = logging.getLogger(__name__)
# A flag is temporary scaffolding. See test_a_flag_older_than_the_limit.
MAX_FLAG_AGE_DAYS = 90
@dataclass(frozen=True)
class Flag:
# One string: the registry key, the Unleash flag name, and the JSON key in
# /api/flags. snake_case, matching the API's existing convention. No case
# transformation anywhere, so there is no mapping layer to get wrong.
name: str
description: str # one line: what turning this on reveals
added: date # for the staleness tripwire
REGISTRY: dict[str, Flag] = {
f.name: f for f in (
Flag(
name="admission_distance",
description=(
"The last-distance-offered figure on the Admissions tile and "
"the 'How far away are you?' section on school pages."
),
added=date(2026, 8, 23),
),
)
}
_client = None
def init() -> None:
"""Start the Unleash client, or log why flags are all off.
Called once from the app lifespan. Never raises: a flag system that can
stop the API from booting is worse than one that is switched off.
"""
global _client
if not settings.unleash_url or not settings.unleash_api_token:
logger.warning(
"Unleash is not configured (UNLEASH_URL / UNLEASH_API_TOKEN); "
"every feature flag evaluates to False.")
return
try:
from UnleashClient import UnleashClient
_client = UnleashClient(
url=settings.unleash_url,
app_name=settings.unleash_app_name,
custom_headers={"Authorization": settings.unleash_api_token},
cache_directory=settings.unleash_cache_directory,
refresh_interval=15,
)
_client.initialize_client()
logger.info("Unleash client initialised against %s", settings.unleash_url)
except Exception:
# Fail closed and keep serving. The SDK also evaluates everything False
# until its first successful sync, so this is the same direction.
_client = None
logger.exception("Unleash client failed to start; flags are all False.")
def is_enabled(name: str) -> bool:
"""Whether `name` is on. False for anything unknown, unreachable or broken."""
if name not in REGISTRY:
logger.error(
"undeclared feature flag %r was evaluated; returning False. "
"Add it to backend/flags.py REGISTRY or fix the name.", name)
return False
if _client is None:
return False
try:
return bool(_client.is_enabled(
name, fallback_function=lambda feature_name, context: False))
except Exception:
logger.exception("flag %r failed to evaluate; returning False", name)
return False
def all_flags() -> dict[str, bool]:
"""Every declared flag and its current value. Serves /api/flags."""
return {name: is_enabled(name) for name in REGISTRY}
+54
View File
@@ -0,0 +1,54 @@
"""Curated London localities, defined by the postcode districts they cover.
The GIAS `town` field puts 1,819 London schools under the single value
"London", so it cannot answer "schools in Battersea" — a query that appears in
the Search Console baseline. No single field can: parliamentary constituency
gives Battersea but not Canary Wharf; postcodes.io's admin_ward gives Canary
Wharf but not Battersea; neither gives Clapham or Shoreditch, which are postal
and colloquial rather than administrative.
So this is curated. Where a locality ends is a judgement, not a fact, and a
reviewable file is the honest place for a judgement. No new ingestion is
needed — the corpus already carries postcodes.
This is the canonical copy. `pipeline/transform/seeds/locality_outcodes.csv`
mirrors it for anyone querying the warehouse directly; the backend image does
not contain `pipeline/`, which is why the module rather than the seed is
canonical. Same arrangement as `backend/gias_codes.py`.
A locality whose outcodes hold fewer than MIN_SCHOOLS schools is not
published, so a typo produces no page rather than an empty one. Places that
fail that check are logged at startup, because a locality you meant to publish
quietly not appearing is the failure worth hearing about.
Two rules for anything added here.
**Sub-borough districts only.** A London borough is a local authority and
already has a page at /schools/authority/[la] covering all of its schools; a
locality defined by two or three outcodes would be a partial, near-duplicate
subset of it. Hackney, Islington, Greenwich and Ealing were all in the first
draft for that reason and have been removed.
**The slug must not match a GIAS town.** "Richmond" did — GIAS has a Richmond
in North Yorkshire with 37 schools — so the London one could never publish.
The registry skips any locality that collides and logs it.
"""
# slug -> (display name, outcodes)
LOCALITY_OUTCODES: dict[str, tuple[str, tuple[str, ...]]] = {
"battersea": ("Battersea", ("SW11",)),
"canary-wharf": ("Canary Wharf", ("E14",)),
"clapham": ("Clapham", ("SW4",)),
"shoreditch": ("Shoreditch", ("EC2A", "E1")),
"peckham": ("Peckham", ("SE15",)),
"brixton": ("Brixton", ("SW2", "SW9")),
"camden-town": ("Camden Town", ("NW1",)),
"wimbledon": ("Wimbledon", ("SW19",)),
"putney": ("Putney", ("SW15",)),
"fulham": ("Fulham", ("SW6",)),
"chiswick": ("Chiswick", ("W4",)),
"stratford": ("Stratford", ("E15",)),
"walthamstow": ("Walthamstow", ("E17",)),
"tooting": ("Tooting", ("SW17",)),
"dulwich": ("Dulwich", ("SE21", "SE22")),
}
+314
View File
@@ -0,0 +1,314 @@
"""The place registry: what places the site publishes, and what is in each.
One module owns this question. The pages, the sitemap and the internal-link
modules all read from here, so the threshold and the collision rules exist in
exactly one place and are testable without a browser or a database.
Two namespaces, never one. 67 viable town names collide with a local
authority name, and the authority is the larger set in only 43 of them —
postal towns cross authority boundaries, so neither can absorb the other.
Keys are "<kind>:<slug>" so the collision cannot reappear in the dict.
"""
from __future__ import annotations
import logging
import re
from dataclasses import dataclass, field
logger = logging.getLogger(__name__)
# Five schools with publishable data. Below this a place has nothing to say
# that a list of schools does not, and publishing it is index bloat.
MIN_SCHOOLS = 5
@dataclass(frozen=True)
class Place:
kind: str # "town" | "locality" | "authority" | "outcode"
slug: str
name: str
urns: tuple[int, ...]
parent_authority: str | None # authority NAME, for the 301 target
# Every authority the place meaningfully sits in, largest first. A quarter
# of outcodes and a third of towns straddle a boundary — SW19 is mostly
# Merton but partly Wandsworth — so naming only one asserts something
# false. parent_authority stays single because a redirect needs one
# target; this is what the page shows.
authorities: tuple[tuple[str, int], ...] = ()
# URNs per phase, so the per-phase threshold can be applied without
# re-querying. A place with 30 primaries and 2 secondaries publishes a
# primary variant and no secondary one.
phase_urns: dict[str, tuple[int, ...]] = field(default_factory=dict)
def publishes_phase(self, phase: str) -> bool:
return len(self.phase_urns.get(phase, ())) >= MIN_SCHOOLS
@property
def key(self) -> str:
return f"{self.kind}:{self.slug}"
def _publishable_urns(df) -> set[int]:
"""URNs with something a page could state, deduplicated across years."""
from backend.app import _PUBLISHABLE_FIELDS
cols = [c for c in _PUBLISHABLE_FIELDS if c in df.columns]
if not cols:
return set()
return set(df.loc[df[cols].notna().any(axis=1), "urn"].astype(int))
# The measure a phase page is built around. A page with no results in this
# column has nothing a list of school names does not already give.
_PHASE_METRIC = {
"primary": "rwm_expected_pct",
"secondary": "attainment_8_score",
}
def _phase_urns(group, publishable: set[int]) -> dict[str, tuple[int, ...]]:
"""URNs per phase, counting only schools with a result for that phase.
Not merely "publishable". A school with an Ofsted grade and no results is
worth a page of its own and belongs in the place list, but it cannot
populate a phase page's results column — and the threshold is there to ask
whether that column will have anything in it.
Counting publishable schools instead let /schools/kent/primary publish
with none of its five rows carrying a result, and left 44 phase pages
majority-blank. It is the same rule as "no page without a local average",
which was never extended per phase.
All-through schools count toward both phases, matching the PHASE_GROUPS
mapping the search filters already use.
"""
from backend.app import PHASE_GROUPS
if "phase" not in group.columns:
return {}
lowered = group["phase"].fillna("").str.lower()
out: dict[str, tuple[int, ...]] = {}
for phase in ("primary", "secondary"):
wanted = PHASE_GROUPS.get(phase, set())
subset = group[lowered.isin(wanted)]
# The page lists every school of the phase; the threshold counts only
# those carrying a result, so a mostly-empty table never publishes.
metric = _PHASE_METRIC[phase]
with_result = (
{int(u) for u in subset.loc[subset[metric].notna(), "urn"]}
if metric in subset.columns else set()
)
if len(with_result & publishable) < MIN_SCHOOLS:
continue
urns = tuple(sorted({int(u) for u in subset["urn"]} & publishable))
if urns:
out[phase] = urns
return out
# A place is described by an authority when it holds at least a tenth of the
# schools, and at least two. GIAS carries occasional postcode errors — EN6
# lists two Shropshire schools among fourteen in Hertfordshire — and a bare
# "any authority present" rule would print those as though they were real.
# There is deliberately no cap on how many are named. An earlier cut stopped
# at three, which silently dropped the fourth in exactly the case where the
# information matters most — a genuinely fragmented place. The share rule is
# the only limit, and it already bounds the list at ten.
_AUTHORITY_MIN_SHARE = 0.10
_AUTHORITY_MIN_SCHOOLS = 2
def _authorities(group) -> tuple[tuple[str, int], ...]:
"""Authorities this place meaningfully sits in, largest first."""
from backend.app import EXCLUDED_FILTER_VALUES
if "local_authority" not in group.columns:
return ()
counts = group["local_authority"].dropna().value_counts()
total = int(counts.sum())
if not total:
return ()
kept = [
(str(name), int(n)) for name, n in counts.items()
if str(name) not in EXCLUDED_FILTER_VALUES
and n >= _AUTHORITY_MIN_SCHOOLS
and n / total >= _AUTHORITY_MIN_SHARE
]
# A place too small or too fragmented for the share rule still names its
# largest authority, or the page would say nothing about where it is.
if not kept:
for name, n in counts.items():
if str(name) not in EXCLUDED_FILTER_VALUES:
return ((str(name), int(n)),)
return ()
return tuple(kept)
def _parent_authority(authorities: tuple[tuple[str, int], ...]) -> str | None:
"""The 301 target: the largest authority a place sits in.
Derived from `authorities` rather than computed separately. The first cut
used `mode()` here while `authorities` used `value_counts()`, and on an
exact tie pandas does not guarantee the two pick the same name — so the
redirect could have pointed somewhere other than the authority the page
named first. One computation, one answer.
Deriving it also inherits the sentinel filter, so a place can no longer
redirect to /schools/authority/does-not-apply.
"""
return authorities[0][0] if authorities else None
def _group(df, column: str, kind: str, publishable: set[int]) -> dict[str, Place]:
"""One Place per distinct SLUG in `column` that clears the threshold.
Grouped by slug, not by raw value, because GIAS spells the same place
several ways and they all resolve to one URL. Five town slugs come from
more than one spelling: "London" (1,819 schools) and "LONDON" (12) both
slugify to `london`; Weston-super-Mare is split 14/19 across two
spellings; Newcastle-under-Lyme across three.
Grouping by raw value meant the later group simply overwrote the earlier
one in this dict — so /schools/london could have shown twelve schools
instead of 1,819, silently and depending on row order.
The display name is the most common spelling, which is the one a reader
expects to see.
"""
from backend.app import _slugify
if column not in df.columns:
return {}
working = df.assign(_slug=df[column].map(
lambda v: _slugify(str(v).strip()) if isinstance(v, str) and v.strip() else None))
working = working[working["_slug"].notna() & (working["_slug"] != "")]
out: dict[str, Place] = {}
for slug, group in working.groupby("_slug"):
slug = str(slug)
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
continue
spellings = group[column].dropna().value_counts()
if spellings.empty:
continue
name = str(spellings.index[0]).strip()
authorities = () if kind == "authority" else _authorities(group)
place = Place(
kind=kind, slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable),
)
out[place.key] = place
return out
# "SW11 2AA" -> "SW11". Two letters max, one or two digits, optional letter.
_OUTCODE_RE = re.compile(r"^([A-Z]{1,2}\d{1,2}[A-Z]?)\s")
def _outcode(postcode) -> str | None:
if not isinstance(postcode, str):
return None
m = _OUTCODE_RE.match(postcode.upper().strip())
return m.group(1) if m else None
def _outcode_places(df, publishable: set[int]) -> dict[str, Place]:
"""One Place per postcode district clearing the threshold.
These carry no phase variants: nobody searches "primary schools in SW11",
so the spec gives them no /primary or /secondary route. `phase_urns` is
left empty rather than computed and then filtered downstream — the
registry is the one place that decides which phases a place publishes,
and the page links whatever it reports.
Computing them here put a link to a route that does not exist on every one
of the 1,720 outcode pages.
"""
if "postcode" not in df.columns:
return {}
working = df.assign(_oc=df["postcode"].map(_outcode))
working = working[working["_oc"].notna()]
out: dict[str, Place] = {}
for oc, group in working.groupby("_oc"):
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
continue
authorities = _authorities(group)
place = Place(kind="outcode", slug=str(oc).lower(), name=str(oc),
urns=urns, parent_authority=_parent_authority(authorities),
authorities=authorities)
out[place.key] = place
return out
def _locality_places(df, publishable: set[int],
town_slugs: set[str]) -> dict[str, Place]:
"""One Place per curated locality clearing the threshold."""
from backend.localities import LOCALITY_OUTCODES
if "postcode" not in df.columns:
return {}
working = df.assign(_oc=df["postcode"].map(_outcode))
out: dict[str, Place] = {}
for slug, (name, outcodes) in LOCALITY_OUTCODES.items():
if slug in town_slugs:
# Skip, do not raise. The guard exists so a locality never
# silently shadows a town — skipping achieves that, and the error
# log makes it loud.
#
# Raising here took down sitemap generation for all 25,000 school
# pages when "richmond" met the GIAS town Richmond in North
# Yorkshire. Worse, GIAS town names change without any code change,
# so a raise means curated data can break the site spontaneously.
# A curation mistake must cost one page, not the sitemap.
logger.error(
"locality %r collides with the published town of the same "
"slug and has been skipped; rename it or remove it", slug)
continue
group = working[working["_oc"].isin(outcodes)]
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
# Not an error — a locality can legitimately be too small. Logged
# because one you meant to publish quietly vanishing is the
# failure worth hearing about.
logger.warning(
"locality %s (%s) has %d publishable schools, below the "
"threshold of %d - not published",
slug, ", ".join(outcodes), len(urns), MIN_SCHOOLS)
continue
authorities = _authorities(group)
place = Place(kind="locality", slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable))
out[place.key] = place
return out
def build_place_registry(df) -> dict[str, Place]:
"""Every place the site publishes, keyed by "<kind>:<slug>"."""
if df.empty or "urn" not in df.columns:
return {}
publishable = _publishable_urns(df)
registry: dict[str, Place] = {}
registry.update(_group(df, "local_authority", "authority", publishable))
towns = _group(df, "town", "town", publishable)
registry.update(towns)
town_slugs = {p.slug for p in towns.values()}
registry.update(_locality_places(df, publishable, town_slugs))
registry.update(_outcode_places(df, publishable))
return registry
+163
View File
@@ -0,0 +1,163 @@
"""Tests for the feature flag layer (spec 2026-08-23).
None of these need a running Unleash. That is the point: an unset UNLEASH_URL
means every flag is False, which is what local development and CI get.
"""
from datetime import date, timedelta
from backend import flags
def test_every_declared_flag_is_keyed_by_its_own_name():
# One string is the registry key, the Unleash flag name and the JSON key.
# A mismatch here would mean the UI toggles a flag the code never reads.
for key, flag in flags.REGISTRY.items():
assert key == flag.name
def test_flag_names_are_snake_case():
# Matches the API's existing convention (admission_distance,
# rwm_expected_pct) so no case transformation exists to get wrong.
for name in flags.REGISTRY:
assert name == name.lower()
assert "-" not in name and " " not in name
def test_an_unconfigured_client_evaluates_every_flag_false(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
for name in flags.REGISTRY:
assert flags.is_enabled(name) is False
def test_an_undeclared_flag_is_false_rather_than_an_error(monkeypatch):
# A typo'd flag name must not raise in a request path. It is logged as an
# error, because an undeclared flag is always a bug.
monkeypatch.setattr(flags, "_client", None)
assert flags.is_enabled("no_such_flag") is False
def test_an_exploding_client_is_false_rather_than_a_500(monkeypatch):
class Boom:
def is_enabled(self, *a, **kw):
raise RuntimeError("unleash is on fire")
monkeypatch.setattr(flags, "_client", Boom())
name = next(iter(flags.REGISTRY))
assert flags.is_enabled(name) is False
def test_all_flags_reports_every_declared_flag(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
assert set(flags.all_flags()) == set(flags.REGISTRY)
assert all(v is False for v in flags.all_flags().values())
def test_a_flag_older_than_the_limit_fails_this_test():
"""A tripwire, not an assertion about correctness.
Flags are temporary scaffolding and the failure mode of every flag system
is accumulation. This fails on the day a flag turns 90, on whatever PR
happens to be open — which is the point: someone has to decide.
To fix: delete the flag and the branches that read it, or, if it genuinely
still needs to exist, move its `added` date and say why in the commit.
"""
stale = [
f.name for f in flags.REGISTRY.values()
if date.today() - f.added > timedelta(days=flags.MAX_FLAG_AGE_DAYS)
]
assert not stale, (
f"Flags older than {flags.MAX_FLAG_AGE_DAYS} days: {stale}. "
"Remove the flag and the code branches it guards, or move its `added` "
"date deliberately."
)
def _client():
from fastapi.testclient import TestClient
from backend import app as app_module
return TestClient(app_module.app, raise_server_exceptions=False)
def test_the_flags_endpoint_lists_every_declared_flag(monkeypatch):
monkeypatch.setattr(flags, "_client", None)
body = _client().get("/api/flags").json()
assert set(body) == set(flags.REGISTRY)
def test_the_flags_endpoint_answers_false_when_unleash_is_unreachable(monkeypatch):
# The endpoint must still answer. A frontend that cannot read flags renders
# everything dark, which is right; one that gets a 500 renders nothing.
monkeypatch.setattr(flags, "_client", None)
res = _client().get("/api/flags")
assert res.status_code == 200
assert all(v is False for v in res.json().values())
def _school_payload(monkeypatch, *, flag_on: bool):
"""Fetch one school's payload with the distance flag forced on or off.
The DataFrame shape is copied from test_school_details.py rather than
minimised: the endpoint reads a wide set of GIAS columns, and a trimmed
frame fails for reasons that have nothing to do with flags.
"""
import numpy as np
import pandas as pd
from fastapi.testclient import TestClient
from backend import app as app_module
df = pd.DataFrame([{
"urn": 150275,
"school_name": "West London Performing Arts Academy",
"phase": "Secondary",
"school_type": "Special post 16 institution",
"trust_name": None,
"religious_denomination": "Does not apply",
"gender": None,
"age_range": "16-25",
"admissions_policy": None,
"capacity": np.nan,
"gias_total_pupils": np.nan,
"headteacher_name": None,
"website": None,
"ofsted_grade": np.nan,
"local_authority": "Ealing",
"address": "268 Northfield Avenue, London, W5 4UB",
"postcode": "W5 4UB",
"latitude": 51.4986,
"longitude": -0.3148,
"year": np.nan,
"total_pupils": np.nan,
"eligible_pupils": np.nan,
"rwm_expected_pct": np.nan,
}])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
# Two arguments: get_supplementary_data(db, urn). See backend/app.py.
monkeypatch.setattr(
app_module, "get_supplementary_data",
lambda db, urn: {"admission_distance": {"distance_m": 772.49,
"year": 2024}})
monkeypatch.setattr(flags, "is_enabled", lambda name: flag_on)
client = TestClient(app_module.app, raise_server_exceptions=False)
res = client.get("/api/schools/150275")
assert res.status_code == 200, res.text
return res.json()
def test_the_distance_field_is_absent_when_the_flag_is_off(monkeypatch):
"""Absent, not null, and withheld at the source.
/api/schools/ is public and unauthenticated. Leaving a withheld field in
the payload while declining to render it hands the record to anyone who
opens the network tab — the reasoning already recorded in c9a1892.
"""
body = _school_payload(monkeypatch, flag_on=False)
assert "admission_distance" not in body
def test_the_distance_field_is_present_when_the_flag_is_on(monkeypatch):
body = _school_payload(monkeypatch, flag_on=True)
assert body["admission_distance"]["distance_m"] == 772.49
+420
View File
@@ -0,0 +1,420 @@
"""Tests for the place registry (spec 2026-08-21).
The registry is built from the in-memory school DataFrame, so these build a
small frame directly rather than touching a database.
"""
import numpy as np
import pandas as pd
import pytest
from backend.places import MIN_SCHOOLS, build_place_registry
def _df(rows: list[dict]) -> pd.DataFrame:
base = {
"year": 202425, "ofsted_grade": 2.0, "ofsted_date": None,
"rwm_expected_pct": 60.0, "attainment_8_score": np.nan,
"phase": "Primary", "postcode": "AA1 1AA",
}
return pd.DataFrame([{**base, **r} for r in rows])
def _town(n: int, town: str, la: str, start: int = 100000, **kw) -> list[dict]:
"""`start` offsets the URNs so two calls can describe different schools —
the Bedford case needs two authorities' worth of distinct URNs in one
town."""
return [
{"urn": start + i, "school_name": f"{town} School {i}",
"town": town, "local_authority": la, **kw}
for i in range(n)
]
def test_town_clearing_the_threshold_is_published():
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
assert "town:brentwood" in reg
assert reg["town:brentwood"].name == "Brentwood"
assert len(reg["town:brentwood"].urns) == MIN_SCHOOLS
def test_town_below_the_threshold_is_not_published():
reg = build_place_registry(_df(_town(MIN_SCHOOLS - 1, "Crosby", "Sefton")))
assert "town:crosby" not in reg
def test_a_town_below_threshold_still_names_its_authority():
# The route layer needs somewhere to 301 to.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS - 1, "Crosby", "Sefton") + _town(MIN_SCHOOLS, "Bootle", "Sefton")))
assert "authority:sefton" in reg
def test_town_and_authority_of_the_same_name_are_separate_places():
# 67 real collisions. Neither set contains the other: Bedford the town has
# 104 schools, Bedford the authority 86, because postal towns cross
# authority boundaries.
rows = (_town(MIN_SCHOOLS, "Bedford", "Bedford")
+ _town(MIN_SCHOOLS, "Bedford", "Central Bedfordshire", start=200000))
reg = build_place_registry(_df(rows))
town, authority = reg["town:bedford"], reg["authority:bedford"]
assert set(town.urns) != set(authority.urns)
assert len(town.urns) == MIN_SCHOOLS * 2 # both authorities' schools
assert len(authority.urns) == MIN_SCHOOLS # only this authority's
def test_schools_without_publishable_data_do_not_count_toward_the_threshold():
rows = _town(MIN_SCHOOLS, "Ghosttown", "Nowhere")
for r in rows:
r["rwm_expected_pct"] = np.nan
r["ofsted_grade"] = np.nan
reg = build_place_registry(_df(rows))
assert "town:ghosttown" not in reg
def test_blank_town_is_ignored():
rows = _town(MIN_SCHOOLS, "", "Essex")
reg = build_place_registry(_df(rows))
assert not any(k.startswith("town:") for k in reg)
def test_a_school_is_counted_once_even_with_several_years_of_rows():
rows = []
for year in (202324, 202425):
rows += [{**r, "year": year} for r in _town(MIN_SCHOOLS, "Beccles", "Suffolk")]
reg = build_place_registry(_df(rows))
assert len(reg["town:beccles"].urns) == MIN_SCHOOLS
def test_locality_groups_schools_by_outcode(monkeypatch):
# The GIAS town field collapses 1,819 London schools into "London", so a
# locality is defined by its postcode districts instead.
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"battersea": ("Battersea", ("SW11",))})
rows = _town(MIN_SCHOOLS, "London", "Wandsworth")
for r in rows:
r["postcode"] = "SW11 2AA"
reg = build_place_registry(_df(rows))
assert reg["locality:battersea"].name == "Battersea"
assert len(reg["locality:battersea"].urns) == MIN_SCHOOLS
def test_locality_below_the_threshold_is_not_published(monkeypatch):
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"nowhere": ("Nowhere", ("ZZ99",))})
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "London", "Wandsworth")))
assert "locality:nowhere" not in reg
def test_a_locality_may_not_shadow_a_viable_town(monkeypatch, caplog):
"""A colliding locality is skipped loudly, and the town survives.
This used to raise, which took down sitemap generation for all 25,000
school pages the first time a curated slug met a real GIAS town. Curated
data must not be able to break the site — and GIAS town names change with
no code change at all, so the raise could fire spontaneously.
"""
import logging
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"brentwood": ("Brentwood", ("CM13",))})
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
with caplog.at_level(logging.ERROR):
reg = build_place_registry(_df(rows))
assert "locality:brentwood" not in reg # skipped
assert "town:brentwood" in reg # the town is untouched
assert "brentwood" in caplog.text # and it was loud about it
def test_a_locality_collision_does_not_break_the_rest_of_the_registry(monkeypatch):
# The whole point of skipping rather than raising.
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"brentwood": ("Brentwood", ("CM13",))})
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
reg = build_place_registry(_df(rows))
assert "authority:essex" in reg
assert "outcode:cm13" in reg
def test_outcode_places_are_built_from_postcodes():
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:cm13"].name == "CM13"
assert len(reg["outcode:cm13"].urns) == MIN_SCHOOLS
def test_malformed_postcodes_do_not_create_places():
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "not a postcode"
reg = build_place_registry(_df(rows))
assert not any(k.startswith("outcode:") for k in reg)
def test_every_curated_locality_is_structurally_valid():
# Guards the hand-maintained file: real slug, real name, real outcodes.
import re
from backend.localities import LOCALITY_OUTCODES
assert LOCALITY_OUTCODES, "the curated locality list must not be empty"
for slug, (name, outcodes) in LOCALITY_OUTCODES.items():
assert re.fullmatch(r"[a-z0-9-]+", slug), slug
assert name.strip() == name and name, slug
assert outcodes, f"{slug} has no outcodes"
for oc in outcodes:
assert re.fullmatch(r"[A-Z]{1,2}\d{1,2}[A-Z]?", oc), (slug, oc)
def test_the_pipeline_seed_mirrors_the_canonical_module():
"""Two copies with no drift guard is worse than one copy.
backend/localities.py is canonical because the backend image does not
contain pipeline/. The seed exists so the warehouse can join on the same
definitions, and this is what stops the two diverging — the same
arrangement assert_gias_code_names_match_seed.sql gives gias_codes.
"""
import csv
from pathlib import Path
from backend.localities import LOCALITY_OUTCODES
seed_path = (Path(__file__).resolve().parents[2]
/ "pipeline/transform/seeds/locality_outcodes.csv")
assert seed_path.exists(), f"missing seed mirror at {seed_path}"
seed = {
row["locality_slug"]: (row["locality_name"],
tuple(row["outcodes"].split("|")))
for row in csv.DictReader(seed_path.open())
}
assert seed == LOCALITY_OUTCODES
def test_no_curated_locality_names_a_london_borough():
"""Boroughs are authorities and already have a page.
A locality defined by two or three outcodes inside a borough would be a
partial, near-duplicate subset of that authority page — the exact
thin-content failure the two-namespace design exists to avoid. Hackney,
Islington, Greenwich and Ealing were all in the first draft.
Hardcoded rather than read from the corpus because this must fail in CI,
where there is no database.
"""
from backend.localities import LOCALITY_OUTCODES
boroughs = {
"barking-and-dagenham", "barnet", "bexley", "brent", "bromley",
"camden", "croydon", "ealing", "enfield", "greenwich", "hackney",
"hammersmith-and-fulham", "haringey", "harrow", "havering",
"hillingdon", "hounslow", "islington", "kensington-and-chelsea",
"kingston-upon-thames", "lambeth", "lewisham", "merton", "newham",
"redbridge", "richmond-upon-thames", "southwark", "sutton",
"tower-hamlets", "waltham-forest", "wandsworth", "westminster",
}
named = boroughs & set(LOCALITY_OUTCODES)
assert not named, (
f"these are boroughs, not districts: {sorted(named)} - they already "
"have an authority page covering every school"
)
def test_a_place_names_every_authority_it_straddles():
"""SW19 is mostly Merton but partly Wandsworth.
A quarter of viable outcodes and a third of viable towns cross an
authority boundary, so naming only the largest asserts something false.
"""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
names = [n for n, _ in reg["outcode:sw19"].authorities]
assert names == ["Merton", "Wandsworth"] # largest first
assert dict(reg["outcode:sw19"].authorities)["Wandsworth"] == 7
def test_the_redirect_target_stays_a_single_authority():
# parent_authority and authorities do different jobs: a 301 needs one
# target, the page needs the truth.
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:sw19"].parent_authority == "Merton"
def test_a_stray_authority_below_the_share_threshold_is_not_named():
# GIAS carries postcode errors — EN6 lists two Shropshire schools among
# fourteen in Hertfordshire. Printing those as though real would be worse
# than omitting them.
rows = (_town(30, "Barnet", "Hertfordshire", start=300000)
+ _town(1, "Barnet", "Shropshire", start=400000))
for r in rows:
r["postcode"] = "EN6 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:en6"].authorities] == ["Hertfordshire"]
def test_a_sentinel_authority_is_never_named():
rows = (_town(20, "London", "Merton", start=300000)
+ _town(6, "London", "Does not apply", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:sw19"].authorities] == ["Merton"]
def test_a_place_always_names_at_least_one_authority():
# Even when every authority is below the share threshold, the page has to
# say where the place is.
rows = []
for i, la in enumerate(["A", "B", "C", "D", "E", "F", "G"]):
rows += _town(1, "Fragmented", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
place = reg.get("town:fragmented")
assert place is not None
assert len(place.authorities) == 1
def test_every_qualifying_authority_is_named_with_no_cap():
"""An earlier cut stopped at three, dropping the fourth silently.
That truncation bit exactly where the information matters most — a
genuinely fragmented place — and nothing recorded it.
"""
rows = []
for i, la in enumerate(["Hackney", "Lambeth", "Westminster", "Lewisham"]):
rows += _town(3, "Fourway", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
assert len(reg["town:fourway"].authorities) == 4
def test_the_redirect_target_is_the_authority_named_first():
"""They were computed separately — mode() against value_counts() — and on
an exact tie pandas does not guarantee the two agree."""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
place = build_place_registry(_df(rows))["outcode:sw19"]
assert place.parent_authority == place.authorities[0][0]
def test_a_place_never_redirects_to_a_sentinel_authority():
# Deriving the parent from `authorities` inherits its sentinel filter.
rows = (_town(6, "Someplace", "Does not apply", start=300000)
+ _town(5, "Someplace", "Essex", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:someplace"].parent_authority == "Essex"
def test_spellings_of_one_place_are_merged_not_overwritten():
"""GIAS spells the same place several ways, and they share a URL.
"London" (1,819 schools) and "LONDON" (12) both slugify to `london`.
Grouping by raw value let the later group overwrite the earlier one, so
the page could have shown twelve schools instead of 1,819 — silently, and
depending on row order.
"""
rows = (_town(6, "Weston-super-Mare", "North Somerset", start=300000)
+ _town(5, "Weston-Super-Mare", "North Somerset", start=400000))
reg = build_place_registry(_df(rows))
assert len(reg["town:weston-super-mare"].urns) == 11
def test_the_merged_place_takes_its_most_common_spelling():
rows = (_town(9, "Newcastle-under-Lyme", "Staffordshire", start=300000)
+ _town(5, "NEWCASTLE-UNDER-LYME", "Staffordshire", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:newcastle-under-lyme"].name == "Newcastle-under-Lyme"
def test_a_phase_page_needs_results_not_merely_publishable_schools():
"""/schools/kent/primary published with none of its five rows scored.
The threshold counted schools that were publishable — a result OR an
Ofsted grade — while the page exists for its results column. Forty-four
phase pages were majority-blank; one had no results at all.
"""
rows = _town(MIN_SCHOOLS, "Kent", "Kent")
for r in rows:
r["rwm_expected_pct"] = np.nan # Ofsted only, no results
reg = build_place_registry(_df(rows))
assert "town:kent" in reg # the place still publishes
assert not reg["town:kent"].publishes_phase("primary")
def test_a_phase_page_publishes_once_enough_schools_carry_a_result():
rows = _town(MIN_SCHOOLS, "Beccles", "Suffolk")
reg = build_place_registry(_df(rows))
assert reg["town:beccles"].publishes_phase("primary")
def test_a_publishing_phase_page_still_lists_its_unscored_schools():
"""The threshold gates whether the page exists; it does not filter rows.
A parent looking up a school by name has to find it whether or not it
published results.
"""
scored = _town(MIN_SCHOOLS, "Beccles", "Suffolk", start=300000)
unscored = _town(2, "Beccles", "Suffolk", start=400000)
for r in unscored:
r["rwm_expected_pct"] = np.nan
reg = build_place_registry(_df(scored + unscored))
place = reg["town:beccles"]
assert place.publishes_phase("primary")
assert len(place.phase_urns["primary"]) == MIN_SCHOOLS + 2
def test_the_secondary_threshold_counts_its_own_metric():
# A town full of scored primaries must not thereby publish a secondary page.
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
reg = build_place_registry(_df(rows))
assert not reg["town:brentwood"].publishes_phase("secondary")
def test_an_outcode_publishes_no_phase_variants():
"""There is no /schools/near/[outcode]/[phase] route, by design.
Nobody searches "primary schools in SW11", so the spec gives outcodes no
phase variants. The registry computed them anyway, and the place page —
which links whatever phases the registry reports — put two 404s on every
outcode page in the site.
This is the single rule now: a kind with no phase route reports no phases,
so neither the page nor the sitemap can offer one.
"""
rows = [{"urn": 500000 + i, "school_name": f"SW11 School {i}",
"town": "London", "local_authority": "Wandsworth",
"postcode": "SW11 1AA"} for i in range(MIN_SCHOOLS + 3)]
reg = build_place_registry(_df(rows))
place = reg["outcode:sw11"]
assert place.phase_urns == {}
assert not place.publishes_phase("primary")
assert not place.publishes_phase("secondary")
def test_an_authority_still_publishes_phase_variants():
"""Authorities keep theirs — "primary schools in Kent" is a real query,
and /schools/authority/[la]/[phase] is the route that serves it."""
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Maidstone", "Kent")))
assert reg["authority:kent"].publishes_phase("primary")
+139
View File
@@ -0,0 +1,139 @@
"""Tests for the places API (spec 2026-08-21)."""
import numpy as np
import pandas as pd
import pytest
from fastapi.testclient import TestClient
def _schools_df() -> pd.DataFrame:
base = {
"local_authority": "Essex", "school_type": "Academy",
"phase": "Primary", "year": 202425, "ofsted_grade": 2.0,
"ofsted_date": None, "attainment_8_score": np.nan,
"town": "Brentwood", "postcode": "CM13 1AA", "status": "Open",
"address": "1 Test Street", "latitude": 51.6, "longitude": 0.3,
}
return pd.DataFrame([
{**base, "urn": 100000 + i, "school_name": f"Brentwood School {i}",
"rwm_expected_pct": 50.0 + i}
for i in range(6)
])
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _schools_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_registry_lists_each_published_place(client):
body = client.get("/api/places").json()
slugs = {(p["kind"], p["slug"]) for p in body["places"]}
assert ("town", "brentwood") in slugs
assert ("authority", "essex") in slugs
assert ("outcode", "cm13") in slugs
def test_registry_carries_a_count_per_place(client):
body = client.get("/api/places").json()
town = next(p for p in body["places"] if p["slug"] == "brentwood")
assert town["count"] == 6
def test_place_detail_returns_its_schools_alphabetically(client):
"""A place page is read by someone looking for a school they can name.
Scanning for it is what the order should serve, so the list is A-Z.
/api/rankings is where the league-table ordering lives.
"""
body = client.get("/api/places/town/brentwood").json()
assert body["place"]["name"] == "Brentwood"
names = [s["school_name"] for s in body["schools"]]
assert names == sorted(names, key=str.lower)
def test_place_ordering_ignores_case(client):
body = client.get("/api/places/town/brentwood").json()
names = [s["school_name"] for s in body["schools"]]
# A capitalised name must not sort ahead of every lowercase one.
assert names == sorted(names, key=str.lower)
def test_the_rankings_endpoint_still_ranks_by_metric(client):
# Alphabetical is a place-page decision, not a site-wide one.
body = client.get("/api/rankings?metric=rwm_expected_pct&phase=primary").json()
scores = [r["rwm_expected_pct"] for r in body.get("rankings", [])
if r.get("rwm_expected_pct") is not None]
assert scores == sorted(scores, reverse=True)
def test_place_detail_carries_the_local_average(client):
body = client.get("/api/places/town/brentwood").json()
# 50..55 inclusive
assert body["averages"]["rwm_expected_pct"] == pytest.approx(52.5)
def test_phase_filter_narrows_the_school_list(client):
body = client.get("/api/places/town/brentwood?phase=secondary").json()
assert body["schools"] == []
def test_unknown_place_404s(client):
assert client.get("/api/places/town/atlantis").status_code == 404
def test_unknown_kind_404s(client):
assert client.get("/api/places/planet/mars").status_code == 404
def _straddling_df() -> pd.DataFrame:
"""Eight schools in CM13: six in Essex, which has a page, and two in an
authority too small to have one.
Two, not one: the registry ignores an authority holding a single school in
a place, because GIAS carries occasional postcode errors."""
df = _schools_df()
extra = df.iloc[:2].copy()
extra["urn"] = [200000, 200001]
extra["school_name"] = ["Scilly School 0", "Scilly School 1"]
extra["local_authority"] = "Isles Of Scilly"
return pd.concat([df, extra], ignore_index=True)
@pytest.fixture()
def straddling_client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _straddling_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _straddling_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_an_outcode_reports_no_phases_because_it_has_no_phase_route(client):
body = client.get("/api/places/outcode/cm13").json()
assert body["place"]["phases"] == []
def test_an_authority_reports_the_phases_it_publishes(client):
body = client.get("/api/places/authority/essex").json()
assert body["place"]["phases"] == ["primary"]
def test_an_authority_without_a_page_is_named_but_carries_no_slug(straddling_client):
"""Two English authorities — City of London and the Isles of Scilly — hold
fewer than the five schools a page needs, so they have no page.
Naming them is still right: the page says where the place is. Linking them
would not be. A null slug is what tells the page to print the name plainly
rather than invent a URL that 404s.
"""
body = straddling_client.get("/api/places/outcode/cm13").json()
by_name = {a["name"]: a for a in body["place"]["authorities"]}
assert by_name["Essex"]["slug"] == "essex"
assert by_name["Isles Of Scilly"]["slug"] is None
+306
View File
@@ -0,0 +1,306 @@
"""Tests for sitemap generation (spec 2026-08-20, workstream W1).
The sitemap is built from the in-memory school DataFrame, so these inject a
small frame via monkeypatch rather than touching a database.
"""
import numpy as np
import pandas as pd
import pytest
def _schools_df() -> pd.DataFrame:
"""Two schools: one with results, one with neither results nor Ofsted."""
base = {
"local_authority": "Testshire",
"school_type": "Academy",
"phase": "Primary",
"year": 202425,
"ofsted_date": None,
}
return pd.DataFrame(
[
{**base, "urn": 100001, "school_name": "Alpha Primary",
"rwm_expected_pct": 62.0, "attainment_8_score": np.nan,
"ofsted_grade": 2.0},
{**base, "urn": 100002, "school_name": "Ghost Primary",
"rwm_expected_pct": np.nan, "attainment_8_score": np.nan,
"ofsted_grade": np.nan},
]
)
@pytest.fixture()
def sitemap(monkeypatch) -> str:
"""The sitemap index."""
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
return app_module.build_sitemap()
@pytest.fixture()
def schools_child(monkeypatch) -> str:
"""The first school child sitemap, where school URLs actually live."""
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
return app_module.build_sitemaps()["schools-1.xml"]
@pytest.fixture()
def static_child(monkeypatch) -> str:
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
return app_module.build_sitemaps()["static.xml"]
def test_every_loc_uses_the_www_host(sitemaps):
# The apex 301s to www. A <loc> that redirects burns a crawl per URL.
# Checked across every file, index included, not just one.
#
# A child can legitimately be empty — this fixture holds two schools and no
# town clearing the threshold — so the presence check applies only to files
# that carry URLs. The absence check applies to all of them.
for name, xml in sitemaps.items():
assert "https://schoolcompare.co.uk" not in xml, name
if "<loc>" in xml:
assert "https://www.schoolcompare.co.uk" in xml, name
def test_school_with_results_is_listed(schools_child):
assert "/school/100001-alpha-primary" in schools_child
def test_school_with_no_results_and_no_ofsted_is_omitted(schools_child):
# Nothing for a search result to say about it. Submitting it spends crawl
# budget and drags the corpus-wide quality signal down.
#
# Asserted against the child, not the index: the index carries no school
# URLs at all, so it would pass this trivially and prove nothing.
assert "/school/100002" not in schools_child
def test_no_invented_priority_or_changefreq(sitemaps):
# Google ignores both. They were noise dressed as signal.
for name, xml in sitemaps.items():
assert "<priority>" not in xml, name
assert "<changefreq>" not in xml, name
def test_ofsted_date_becomes_lastmod(monkeypatch):
from backend import app as app_module
import datetime
def _df():
base = _schools_df()
base.loc[base["urn"] == 100001, "ofsted_date"] = datetime.date(2024, 3, 14)
return base
monkeypatch.setattr(app_module, "load_school_data", _df)
xml = app_module.build_sitemaps()["schools-1.xml"]
assert "<lastmod>2024-03-14</lastmod>" in xml
def test_no_lastmod_invented_when_date_unknown(monkeypatch):
# An always-now lastmod is a claim Google learns to distrust. Absent
# honestly means unknown.
from backend import app as app_module
def _df():
df = _schools_df()
df["ofsted_date"] = None
return df
monkeypatch.setattr(app_module, "load_school_data", _df)
# The child only. The index legitimately carries a lastmod, because there
# it means "when this sitemap file changed", which we do know.
xml = app_module.build_sitemaps()["schools-1.xml"]
assert "<lastmod>" not in xml
def test_static_routes_are_listed(static_child):
for path in ("/", "/rankings", "/compare", "/admissions"):
assert f"<loc>https://www.schoolcompare.co.uk{path}</loc>" in static_child
@pytest.fixture()
def sitemaps(monkeypatch) -> dict:
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
return app_module.build_sitemaps()
def test_index_lists_each_child(sitemaps):
index = sitemaps["sitemap.xml"]
assert "<sitemapindex" in index
assert "https://www.schoolcompare.co.uk/sitemaps/static.xml" in index
assert "https://www.schoolcompare.co.uk/sitemaps/schools-1.xml" in index
def test_index_carries_no_url_elements(sitemaps):
# A sitemap index holds <sitemap> entries only; mixing in <url> is invalid.
assert "<url>" not in sitemaps["sitemap.xml"]
def test_index_does_not_list_itself(sitemaps):
assert "<loc>https://www.schoolcompare.co.uk/sitemap.xml</loc>" not in sitemaps["sitemap.xml"]
def test_static_child_holds_the_static_routes(sitemaps):
static = sitemaps["static.xml"]
for path in ("/", "/rankings", "/compare", "/admissions"):
assert f"<loc>https://www.schoolcompare.co.uk{path}</loc>" in static
def test_school_child_holds_the_schools(sitemaps):
assert "/school/100001-alpha-primary" in sitemaps["schools-1.xml"]
def test_children_are_chunked_under_the_limit(monkeypatch):
# Sitemaps cap at 50,000 URLs per file. Chunk at 10,000 so a child stays
# small enough to eyeball in Search Console.
from backend import app as app_module
import pandas as _pd
rows = [
{"urn": 200000 + i, "school_name": f"School {i}", "year": 202425,
"rwm_expected_pct": 60.0, "attainment_8_score": None,
"ofsted_grade": 2.0, "ofsted_date": None}
for i in range(10_001)
]
monkeypatch.setattr(app_module, "load_school_data", lambda: _pd.DataFrame(rows))
maps = app_module.build_sitemaps()
assert maps["schools-1.xml"].count("<url>") == 10_000
assert maps["schools-2.xml"].count("<url>") == 1
def test_build_sitemap_still_returns_the_index(sitemap):
# lifespan and the admin endpoint call build_sitemap(); keep it working.
assert "<sitemapindex" in sitemap
def test_school_with_results_in_an_earlier_year_is_still_listed(monkeypatch):
"""Regression: The Mallard Academy (150367).
Real KS2 results 2015-16 to 2018-19, then null rows from 2022-23 onward
because the school stopped reporting. The first cut tested the latest
year's row alone and dropped it, along with ~220 others, even though its
detail page shows all four years of results.
"""
from backend import app as app_module
import pandas as _pd
base = {"local_authority": "Testshire", "school_type": "Academy",
"phase": "Primary", "ofsted_date": None, "ofsted_grade": np.nan,
"attainment_8_score": np.nan, "urn": 150367,
"school_name": "Mallard Academy"}
df = _pd.DataFrame([
{**base, "year": 201819, "rwm_expected_pct": 67.0},
{**base, "year": 202324, "rwm_expected_pct": np.nan},
{**base, "year": 202425, "rwm_expected_pct": np.nan},
])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
xml = app_module.build_sitemaps()["schools-1.xml"]
assert "/school/150367-mallard-academy" in xml
def test_school_with_no_results_in_any_year_is_still_omitted(monkeypatch):
"""The fix must not turn into "list everything"."""
from backend import app as app_module
import pandas as _pd
base = {"local_authority": "Testshire", "school_type": "Academy",
"phase": "Primary", "ofsted_date": None, "ofsted_grade": np.nan,
"attainment_8_score": np.nan, "rwm_expected_pct": np.nan,
"urn": 100002, "school_name": "Ghost Primary"}
df = _pd.DataFrame([{**base, "year": y} for y in (202324, 202425)])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
assert "/school/100002" not in app_module.build_sitemaps()["schools-1.xml"]
def _places_df() -> pd.DataFrame:
base = {
"local_authority": "Essex", "school_type": "Academy",
"phase": "Primary", "year": 202425, "ofsted_grade": 2.0,
"ofsted_date": None, "attainment_8_score": np.nan,
"town": "Brentwood", "postcode": "CM13 1AA",
}
return pd.DataFrame([
{**base, "urn": 100000 + i, "school_name": f"Brentwood School {i}",
"rwm_expected_pct": 60.0}
for i in range(6)
])
@pytest.fixture()
def place_sitemaps(monkeypatch) -> dict:
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _places_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return app_module.build_sitemaps()
def test_place_children_are_listed_in_the_index(place_sitemaps):
index = place_sitemaps["sitemap.xml"]
assert "/sitemaps/places-1.xml" in index
assert "/sitemaps/outcodes-1.xml" in index
def test_town_and_authority_urls_use_their_own_namespaces(place_sitemaps):
xml = place_sitemaps["places-1.xml"]
assert "<loc>https://www.schoolcompare.co.uk/schools/brentwood</loc>" in xml
assert "<loc>https://www.schoolcompare.co.uk/schools/authority/essex</loc>" in xml
def test_outcode_urls_live_in_their_own_child(place_sitemaps):
assert "/schools/near/cm13" in place_sitemaps["outcodes-1.xml"]
assert "/schools/near/cm13" not in place_sitemaps["places-1.xml"]
def test_place_urls_carry_no_priority_or_changefreq(place_sitemaps):
for name in ("places-1.xml", "outcodes-1.xml"):
assert "<priority>" not in place_sitemaps[name]
assert "<changefreq>" not in place_sitemaps[name]
def test_phase_variants_are_submitted_where_the_phase_clears_the_threshold(place_sitemaps):
# "primary schools in beccles" is the query shape the baseline showed, so
# each variant is its own page and has to be submitted. Emitting only the
# bare place URL left ~950 of them reachable by nothing.
xml = place_sitemaps["places-1.xml"]
assert "<loc>https://www.schoolcompare.co.uk/schools/brentwood/primary</loc>" in xml
def test_a_phase_below_its_own_threshold_is_not_submitted(place_sitemaps):
# The fixture is six primaries and no secondaries.
xml = place_sitemaps["places-1.xml"]
assert "/schools/brentwood/secondary" not in xml
def test_outcodes_get_no_phase_variants(place_sitemaps):
# Nobody searches "primary schools in CM13"; the routes do not exist.
xml = place_sitemaps["outcodes-1.xml"]
assert "/primary" not in xml and "/secondary" not in xml
def test_authority_phase_variants_are_submitted_in_their_own_namespace(place_sitemaps):
"""302 of these were already in the sitemap, and every one 404'd.
The spec gives authorities a phase route; the plan built the bare
authority route and dropped it. Nothing noticed because the sitemap was
written from the registry, which was right, while the routes were written
by hand. This test fails if the URL ever leaves the sitemap; the e2e
journey fails if the route ever leaves the app.
"""
xml = place_sitemaps["places-1.xml"]
assert ("<loc>https://www.schoolcompare.co.uk"
"/schools/authority/essex/primary</loc>") in xml
# And never in the town namespace, which is a different set of schools.
assert "/schools/essex/primary" not in xml
+9
View File
@@ -16,6 +16,8 @@
# ADMIN_API_KEY — Backend admin API key
# TYPESENSE_API_KEY — Typesense admin API key
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: development
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs)
# STAGING_DB_IP — macvlan IP for staging Postgres (default 10.0.1.190)
# STAGING_FRONTEND_IP — macvlan IP for staging frontend (default 10.0.1.151)
@@ -55,6 +57,12 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- unleash_cache:/app/.unleash
depends_on:
sc_database:
condition: service_healthy
@@ -212,3 +220,4 @@ volumes:
postgres_data:
typesense_data:
airflow_logs:
unleash_cache:
+73
View File
@@ -0,0 +1,73 @@
# Portainer Stack Definition for School Compare — UNLEASH (feature flags)
#
# Deploy as a *separate* Portainer stack ("schoolcompare-unleash"), alongside
# the production and staging stacks. It deliberately belongs to neither: a
# staging redeploy must not be able to disturb production's flag state, and a
# production redeploy must not disturb staging's.
#
# One instance serves both environments. Open-source Unleash ships with
# `development` and `production` environments and environment-scoped client
# tokens, so the same flag holds independent state in each — which is what
# lets a feature be on in staging, where the E2E journeys exercise it, while
# production stays dark.
#
# Portainer environment variables (set in Portainer UI -> Stack -> Environment):
# UNLEASH_DB_PASSWORD — PostgreSQL password for the Unleash database
# UNLEASH_ADMIN_PASSWORD — initial admin password for the Unleash UI
# UNLEASH_IP — macvlan IP for the Unleash server (default 10.0.1.152)
services:
# ── PostgreSQL (Unleash's own; nothing else uses it) ──────────────────
unleash_db:
container_name: sc_unleash_postgres
image: postgres:16-alpine
environment:
POSTGRES_USER: unleash
POSTGRES_PASSWORD: ${UNLEASH_DB_PASSWORD}
POSTGRES_DB: unleash
volumes:
- unleash_postgres_data:/var/lib/postgresql/data
networks:
- unleash
healthcheck:
test: ["CMD-SHELL", "pg_isready -U unleash"]
interval: 10s
timeout: 5s
retries: 5
start_period: 10s
restart: unless-stopped
# ── Unleash server (UI + client API on 4242) ──────────────────────────
unleash:
container_name: sc_unleash
image: unleashorg/unleash-server:6
environment:
DATABASE_URL: postgres://unleash:${UNLEASH_DB_PASSWORD}@unleash_db:5432/unleash
DATABASE_SSL: "false"
INIT_ADMIN_API_TOKENS: ""
UNLEASH_DEFAULT_ADMIN_PASSWORD: ${UNLEASH_ADMIN_PASSWORD}
depends_on:
unleash_db:
condition: service_healthy
networks:
unleash: {}
macvlan:
ipv4_address: ${UNLEASH_IP:-10.0.1.152}
healthcheck:
test: ["CMD-SHELL", "wget -qO- http://localhost:4242/health || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 30s
restart: unless-stopped
networks:
unleash:
driver: bridge
macvlan:
external:
name: macvlan
volumes:
unleash_postgres_data:
+9
View File
@@ -7,6 +7,8 @@
# ADMIN_API_KEY — Backend admin API key
# TYPESENSE_API_KEY — Typesense admin API key
# TYPESENSE_SEARCH_KEY — Typesense search-only key (exposed to frontend)
# UNLEASH_URL — http://<unleash-ip>:4242/api (empty = all flags off)
# UNLEASH_API_TOKEN — Unleash *client* token, environment: production
# AIRFLOW_ADMIN_USER — Airflow admin username (password auto-generated, see api-server logs)
services:
@@ -44,6 +46,12 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- unleash_cache:/app/.unleash
depends_on:
sc_database:
condition: service_healthy
@@ -201,3 +209,4 @@ volumes:
postgres_data:
typesense_data:
airflow_logs:
unleash_cache:
+4
View File
@@ -36,6 +36,10 @@ services:
ADMIN_API_KEY: ${ADMIN_API_KEY:-changeme}
TYPESENSE_URL: http://typesense:8108
TYPESENSE_API_KEY: ${TYPESENSE_API_KEY:-changeme}
# Unset means every feature flag is False — the correct dark state for an
# environment with no Unleash, not a failure.
UNLEASH_URL: ${UNLEASH_URL:-}
UNLEASH_API_TOKEN: ${UNLEASH_API_TOKEN:-}
volumes:
- ./data:/app/data:ro
depends_on:
+51
View File
@@ -150,3 +150,54 @@ token Gitea Actions provides automatically (`secrets.GITEA_TOKEN` — no setup
needed), and fails the check only when a finding is rated
**severe** (would break prod, leak data, or corrupt data). Minor findings are
informational and never block a merge.
## Feature flags (Unleash)
Flag state lives in a self-hosted Unleash instance, deployed as its own
Portainer stack from `docker-compose.portainer.unleash.yml`. It is separate
from the application stacks on purpose — redeploying staging must not be able
to disturb production's flags.
The flags themselves are declared in `backend/flags.py`. Unleash holds the
state; the registry holds the list. A flag in the UI that is not in the
registry is orphaned and nothing reads it.
### First-time setup
1. Deploy the stack in Portainer. Set `UNLEASH_DB_PASSWORD`,
`UNLEASH_ADMIN_PASSWORD` and (optionally) `UNLEASH_IP`.
2. Log in to the UI at `http://<UNLEASH_IP>:4242` as `admin`.
3. Create one **client** API token per environment:
- `schoolcompare-staging`, environment **development**
- `schoolcompare-prod`, environment **production**
Client tokens, not admin tokens — the backend only reads.
4. Put each token in the matching Portainer stack's `UNLEASH_API_TOKEN`
variable, and set `UNLEASH_URL` to `http://<UNLEASH_IP>:4242/api`.
5. Redeploy the application stacks.
### Turning a feature on
Toggle the flag in the environment you want. Flags appear in the Unleash UI
after the backend has evaluated them once, so a newly declared flag shows up
shortly after the deploy that introduced it.
A flip reaches school pages within about five minutes and place pages within
the hour. Next's ISR does the propagating — it revalidates a route at the
*lowest* `revalidate` among that route's fetches, which is 300s for
`/school/[slug]` and 3600s for the place pages. There is no webhook, and
adding one would only be worth it if flips ever needed to be instant.
### When Unleash is unreachable
Every flag evaluates to `False` and the site serves as though nothing were
switched on. That is deliberate — an unfinished feature staying hidden is the
safe direction — but it means a *released* feature disappears if a backend
container cold-starts with an empty cache while Unleash is down. The SDK's
disk cache is on a named volume so restarts keep last-known state, and flags
are removed from the code within 90 days (enforced by a test), which bounds
how long any feature is exposed to this.
If `UNLEASH_URL` is unset, every flag is `False` and no connection is
attempted. That is the correct behaviour for local development and CI, and it
means the test suites need no flag server.
File diff suppressed because it is too large. Load diff
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,278 @@
# W2: The Location Layer — Design
Date: 2026-08-21
Status: awaiting review
Supersedes: workstream W2 in `2026-08-20-seo-programme-design.md`
Scope note: this covers four page families in one spec. Splitting them — towns
and authorities first, outcodes and localities after — was proposed and
declined in favour of building the layer in one pass. The decomposition
argument was that the curated locality seed needs human review and would hold
up 783 pages of measured demand behind it; that risk is accepted here, and the
implementation plan should sequence the seed early enough that review time
does not become the critical path.
## Problem
Location intent is the largest unserved demand the site has. In the 16-month
Search Console baseline it draws **874 impressions, one click, average
position 49.5**. The site does not compete.
Unlike named-school queries — which the same baseline showed to be
navigational and unwinnable, since a parent typing "audley junior school"
wants that school's own website — location queries have no incumbent owner.
Nobody owns "primary schools in Brentwood" the way a school owns its name.
The cause is structural: the site has no page about a place. Every competitor
ranking above it does.
## What the demand actually looks like
Every location query in the baseline is **town or district level**. Not one is
an administrative area:
| Query | Impressions | Position |
|-------|-------------|----------|
| colleges in solihull | 112 | 51.2 |
| schools in ramsey | 64 | 42.5 |
| schools in crosby | 57 | 47.7 |
| primary schools in beccles | 44 | 40.9 |
| private schools in battersea | 41 | 71.9 |
| secondary schools in brentwood | 37 | 56.1 |
| secondary schools in canary wharf | 30 | 35.9 |
Three patterns follow directly, and they drive the whole design.
**Towns, not authorities.** The superseded W2 put `/schools/[la]` first and
towns second. The data inverts that. Brentwood appears four times in different
phrasings; Beccles twice. Both are towns, not authorities.
**Phase is part of the query**, not a filter applied afterwards: "primary
schools in beccles", "secondary schools in brentwood", "colleges in solihull".
**London is searched by district** — Battersea, Canary Wharf — and the GIAS
`town` field cannot serve it at all.
## Measured sizing
Counted against the live corpus of 25,185 schools, not estimated.
| Family | Viable (≥5 schools) | Below threshold |
|--------|--------------------|-----------------|
| Towns | **783** | 907 → redirect to authority |
| Outcodes | **1,760** | 305 |
| Local authorities | 154 | — |
| London localities | ~100–150 (curated) | — |
With phase variants — 783 town pages plus roughly 700 primary and 250
secondary variants, 154 authorities across three variants, 1,760 outcodes and
the curated localities — the total lands near **4,000 pages**. Phase variants
need their own threshold: there are 17,426 primaries but only 4,456 secondaries nationally,
so most towns will support a primary page and not a secondary one.
## Two design problems this spec exists to solve
### 1. Town and authority names collide, and neither contains the other
67 viable towns share a name with a local authority. The obvious fix — let the
authority absorb the town, since it sounds like a superset — **does not work**:
| Place | Schools in the town | Schools in the authority |
|-------|--------------------|-----------------------|
| Bedford | 104 | 86 |
| Birmingham | 520 | 518 |
| Derby | 157 | 119 |
| Doncaster | 152 | 145 |
The authority is the larger set in only 43 of the 67. Postal towns cross
authority boundaries, so these are overlapping sets that happen to share a
name. Publishing both into one namespace produces near-duplicate pages, which
is the specific failure that sinks programmatic SEO.
**Resolution: two namespaces.**
```
/schools/[place] towns and London localities
/schools/[place]/primary
/schools/[place]/secondary
/schools/authority/[la] local authorities
/schools/authority/[la]/primary
/schools/authority/[la]/secondary
/schools/near/[outcode]
```
Outcodes carry no phase variants: nobody searches "primary schools in SW11",
so the variants would be pages without demand.
Every collision disappears by construction. `/schools/[place]` keeps the clean
URL for the pattern that carries the demand; authorities get a namespace whose
purpose is genuinely different — admissions are authority-run, and the
authority page is the one that can speak to catchment policy and LA averages.
A place page and an authority page of the same name must each say plainly
which set of schools they cover, or they read as duplicates to a reader even
when they differ in fact.
### 2. London has no locality field
`town` collapses **1,819 London schools into the single value "London"**. A
page listing all of them is useless, and borough pages do not help because
people search "Battersea", not "Wandsworth".
No single field solves it:
| Search term | `parliamentary_constituency` | postcodes.io `admin_ward` |
|-------------|------------------------------|---------------------------|
| Battersea | **Battersea** ✓ | Northcote / Wandsworth Town ✗ |
| Canary Wharf | Poplar and Limehouse ✗ | **Canary Wharf** ✓ |
| Vauxhall | Vauxhall and Camberwell Green ✗ | **Vauxhall** ✓ |
And neither covers Clapham, Shoreditch or Peckham, which are postal and
colloquial rather than administrative.
**Resolution: a curated seed mapping locality to outcodes.**
```
pipeline/transform/seeds/locality_outcodes.csv
locality_slug,locality_name,outcodes,region
battersea,Battersea,"SW11|SW8",London
canary-wharf,Canary Wharf,"E14",London
clapham,Clapham,"SW4|SW9",London
```
This needs **no new ingestion** — the corpus already has postcodes. It puts
the fuzzy, contested part of the problem in a reviewable file rather than in
derived logic, which suits it: locality boundaries are a judgement, not a
fact. The repo already uses dbt seeds for curated reference data
(`la_code_names.csv`, `gias_code_names.csv`), so this follows an established
pattern.
The seed generalises past London. Any colloquial place — Jesmond, Chorlton,
Clifton — can be defined by its outcodes without a schema change.
**Constraint:** a locality slug may not collide with a viable town slug. The
place registry enforces this and fails the build rather than silently
shadowing a town.
## Architecture
### The place registry
One module owns the question "what places do we publish, and what is in each".
Everything else reads from it: the pages, the sitemap, the internal links.
```
backend/places.py
Place = { kind: "town"|"locality"|"authority"|"outcode",
slug, name, urn_list, parent_authority | None }
build_place_registry(df) -> dict[str, Place]
place_schools(slug, phase=None) -> list[School]
```
Built once at startup from the same DataFrame the sitemap uses, and rebuilt by
the existing `/api/admin/regenerate-sitemap` path after a pipeline run.
Registry construction is where the threshold, the collision rules and the
seed's uniqueness constraint are enforced — in one place, testable without a
browser or a database.
### API
```
GET /api/places the registry: slug, kind, name, count
GET /api/places/{slug}?phase= aggregate + ranked schools for one place
```
`/api/places` is what the sitemap and the internal-link modules enumerate.
### Routes
Next App Router, ISR with the same 7-day revalidate the school pages use.
`generateStaticParams` gated behind an env flag, matching
`PRERENDER_SCHOOLS`, because 3,900 more routes cannot be statically built in
CI on every deploy.
## What each page must contain
A place page that is a name substituted into a template is the thing Google's
helpful-content stance exists to demote. Each page carries computed local
facts that exist nowhere else on the site:
- **H1** matching the query: "Primary schools in Brentwood"
- **Counts framed usefully**: "29 schools, 4 rated Outstanding"
- **A ranked table** of the top 20 on the phase's headline metric —
`rwm_expected_pct` for primary, `attainment_8_score` for secondary, and for
an unphased place page the metric matching whichever phase it holds more of
- **The local average against the England average** — the one number a parent
cannot get from a list
- **Ofsted grade distribution** for the place
- **A map**
- **Links to neighbouring places** and to the parent authority
- **An FAQ block**, feeding `FAQPage` structured data
- **A link to every school page in scope** — this is what finally de-orphans
the 23,000 school pages the original spec identified as near-orphans
## Thin-page controls
Three, and they are the difference between a location layer and index bloat:
1. **Five schools with current data minimum.** Below it, 301 to the parent
authority. This drops 907 towns and 305 outcodes.
2. **Per-phase thresholds.** A town with 30 primaries and 2 secondaries
publishes a primary page and no secondary page.
3. **No page without a local average.** If a place has too few schools with
results to compute one, it has nothing to say that a list does not, and it
falls back to the authority.
## Sitemap
Two new children in the existing index: `/sitemaps/places-{n}.xml` and
`/sitemaps/outcodes-{n}.xml`. Per-family children are why the index was built
in W1 — Search Console reports coverage per submitted sitemap, so indexation
of the location layer is measurable separately from the school pages.
## Testing
Per `CLAUDE.md`, user-facing behaviour extends `e2e/tests/journeys.spec.ts` in
the same PR.
**Unit (registry, no DB):** threshold enforcement; a sub-threshold town
resolves to its authority; a locality slug colliding with a town fails the
build; Bedford's town and authority pages hold different URN sets; per-phase
thresholds.
**Backend:** `/api/places` shape; `/api/places/{slug}` aggregate correctness
against a fixture; unknown slug 404s.
**e2e:** a known town, authority, locality and outcode page each render with
the expected count; a below-threshold town 301s; every place page declares a
canonical and appears in the sitemap; `/schools/bedford` and
`/schools/authority/bedford` both resolve and state which set they cover.
## Risks
**Index bloat** is the failure mode of every programmatic SEO programme. The
three controls above are the answer, and the per-family sitemap is how we find
out early if they were not enough.
**Helpful-content exposure.** Templated location pages are exactly what
Google's stance targets. The mitigation is that every page carries real
computed local data — counts, distributions, local-versus-national comparison
— rather than a name dropped into boilerplate. If indexation of the places
sitemap stalls below roughly half, that is the signal to stop and rethink
rather than to add more pages.
**Build cost.** ~4,000 additional ISR routes on top of 23,000 school pages.
The env-flag gate on `generateStaticParams` keeps CI viable.
**Curation drift.** The locality seed is hand-maintained and will go stale as
places change. It is small and reviewable, and a dbt test asserts every seed
outcode matches at least one school so a typo fails the pipeline rather than
publishing an empty page.
## Out of scope
Catchment-area estimation. It is a strong driver for this cluster and
`fact_admissions` carries the distances, but it is a modelling problem with
real accuracy risk and deserves its own design.
@@ -0,0 +1,294 @@
# Feature Flags — Design
**Date:** 2026-08-23
**Status:** approved for planning
**First consumer:** the last-distance-offered feature (`admission_distance`)
## Goal
Let work merge to `main` and deploy to production without becoming visible,
so that releasing a feature stops being the same event as deploying it.
The site has no way to do this today. A feature is either on `main` and live,
or it is on a branch. That forces long-lived branches for anything not ready,
and it makes every promotion to production an all-or-nothing decision about
everything queued behind it.
This is a **ship-dark** capability, not a kill switch. Flags are expected to
flip on the order of once a month, by a person, deliberately. Nothing here is
designed for flipping something off in seconds under pressure, and nothing
here does percentage rollouts, user targeting or A/B tests — the site has no
user identity to target.
## Decision: Unleash
Flag state is held in a self-hosted [Unleash](https://www.getunleash.io/)
instance (Apache-2.0), not in the repository.
A lighter option was considered and rejected by the project owner: a typed
registry in each runtime with environment-variable overrides set in the
Portainer stack files, which would have needed no new container and kept flag
state in git. The argument for Unleash is that it provides a UI and an audit
log without a deploy, and that flags are expected to become an ongoing
operational tool rather than an occasional one.
Two consequences follow from choosing a service, and this design exists mostly
to handle them:
1. **Flag state lives outside the repository.** `main` is no longer the whole
truth about what is switched on. The registry in §2 exists to bound that.
2. **A flag can change without a deploy**, so nothing else clears the caches
that a deploy would have cleared. §4 establishes how long a flip takes to
become visible, and why that is short enough to need no extra mechanism.
Also considered: Flagsmith (heavier — Django, Postgres and Redis), GrowthBook
(requires MongoDB), and Flipt v2 (the closest conceptual fit, git-native, but
now under the Fair Core Licence — source-available, not OSI open source).
## 1. Topology
A third Portainer stack, `docker-compose.portainer.unleash.yml`, holding
`unleashorg/unleash-server` and its own PostgreSQL 16. It is on the macvlan so
both application stacks can reach it, and it belongs to neither of them — a
staging redeploy must not be able to disturb production's flag state, and vice
versa.
One instance serves both environments. Open-source Unleash ships with
`development` and `production` environments and environment-scoped client
tokens, so the same flag holds independent state in each: staging's FastAPI
carries a `development` token, production's carries a `production` one.
That property is what makes ship-dark testable. A feature can be **on in
staging and off in production** for as long as it takes, which means the `e2e/`
journeys exercise it against staging while production stays unchanged.
## 2. The registry
Unleash supplies flag *state* and the toggle UI. It does not supply the list of
flags. `backend/flags.py` declares every flag the code knows about:
```python
@dataclass(frozen=True)
class Flag:
name: str # identical in the registry, in Unleash, and in JSON
description: str # one line: what turning this on reveals
added: date # for the staleness test in §8
```
**Every flag defaults to `False`.** There is no per-flag default field, because
a flag that defaults on is not a ship-dark flag — it is a kill switch, and this
design does not offer one. A single unconditional default also means the
fallback path has no branching to get wrong.
Three reasons the registry is not optional:
- The Unleash SDK evaluates an unknown flag to `False`. Without a registry that
is an *undeclared* false — indistinguishable from a typo in a flag name.
- `/api/flags` needs a key set to return when Unleash is unreachable. It cannot
enumerate flags it has never heard of.
- A flag present in the Unleash UI but absent from the registry is orphaned,
and should be visibly so rather than quietly authoritative.
**Naming.** One string, used unchanged as the registry key, the Unleash flag
name, and the JSON key in `/api/flags`. It is snake_case, matching the API's
existing convention (`admission_distance`, `rwm_expected_pct`) and the mirrored
types in `nextjs-app/lib/types.ts`. No case transformation anywhere, so there
is no mapping layer to get wrong.
## 3. Read paths
### Backend
`backend/flags.py` wraps `UnleashClient` behind `is_enabled(name: str) -> bool`.
Fail-closed is the default rather than something added: the Python SDK
evaluates every flag to `False` until it has synchronised with the server. An
unfinished feature therefore stays hidden when Unleash is unreachable, which is
the correct direction for ship-dark.
The SDK's fcache directory is mounted on a named volume so a container restart
during an Unleash outage keeps last-known state rather than reverting a
released feature to dark. The registry default remains `False`, so the worst
case is a feature disappearing, never one appearing.
### Frontend
`nextjs-app/lib/flags.ts` exposes `getFlags(): Promise<Flags>`, a single
server-side fetch of `/api/flags` returning a typed record. Server components
only — no flag value reaches the browser bundle, and `package.json` gains no
Unleash dependency. The Unleash client library stays entirely inside the
service that already owns every other piece of data the frontend renders.
The cost, named plainly: a purely front-end flag must still be declared in a
Python file. It is a flat data edit rather than programming, and the return is
one list, so nobody has to ask which service knows about a given flag.
### `/api/flags` must not be publicly reachable
`nextjs-app/app/api/[...path]/route.ts` proxies **everything** under `/api/` to
FastAPI. Left alone, `https://www.schoolcompare.co.uk/api/flags` would return
`{"admission_distance": false, ...}` — publishing the name and state of every
unreleased feature, which defeats the purpose of shipping dark.
The proxy therefore gains a denylist, and `flags` is on it: a request for a
denied path returns 404 rather than being forwarded. Next's own `getFlags()` is
unaffected because it calls `FASTAPI_URL` directly across the Docker network
and never transits the public proxy.
This is a general hole rather than a flags-specific one — the proxy will
forward any future internal endpoint too — so the denylist is written as a
named constant with a comment saying what belongs on it.
## 4. Propagation
**Time-based revalidation is sufficient. There is no webhook.**
An earlier draft of this section specified two Unleash webhooks and a
`revalidateTag('flags')` purge, on the premise that pages cache for seven days.
That premise was wrong, and checking it removed the most complex part of the
design.
Next uses the **lowest** `revalidate` among a route's fetches to set the
revalidation frequency of the whole route — the segment-level
`export const revalidate` does not override a lower value inside it. Measured
against this codebase:
| Page family | Segment | Lowest fetch | Effective |
|---|---|---|---|
| `/school/[slug]` | 604800 | `fetchSchoolDetails` at 300 | **5 minutes** |
| `/schools/*` | 604800 | `fetchNationalAverages` at 3600 | **1 hour** |
The Unleash SDK polls every 15 seconds, so a flip reaches school pages within
about five minutes and place pages within the hour, unaided. Flags flip
monthly, by hand, deliberately. That is fast enough.
What this removes: two webhook integrations, a `/api/revalidate-flags` route, a
shared-secret-in-a-query-string scheme, an idempotency requirement against
duplicate and out-of-order delivery, and a rule that every fetch in
`nextjs-app/lib/` carry a cache tag. None of it has to be built, maintained, or
kept correct as new fetches are added.
**If instant flips are ever wanted**, the webhook is the way to add them, and it
is purely additive — nothing in this design has to change first.
### Two constraints this leaves behind
**Never flag content on a `force-static` page.** `app/admissions/page.tsx`
declares `export const dynamic = 'force-static'`, so it is baked at build time
and never revalidates. A flag gating anything on such a page would not take
effect until the next deploy, silently. If a flag ever needs to reach one, that
page must first move to ISR.
**A route-family flag still needs the sitemap rebuilt.** The sitemap is held in
memory and rebuilt only at startup or via `POST /api/admin/regenerate-sitemap`.
No flag in scope touches the sitemap (§6), so this is deferred with the route
case rather than solved now — but a route flag must not ship without it, or the
sitemap will advertise URLs that `notFound()`.
## 5. What "off" means, per surface
| Surface | Off |
|---|---|
| Route | `notFound()`, **and** absent from the sitemap, **and** absent from nav |
| UI element | Not rendered; surrounding page byte-identical to today |
| API field | Key **absent**, not `null` |
| API endpoint | 404, not 403 |
The three parts of the route rule move together or not at all. Submitting URLs
to Google that return 404 is the bug fixed in PR #124, and a flag is a new way
to reintroduce it.
An API field is withheld **at the source**, never rendered-but-hidden. The
precedent is already set in this codebase by commit `c9a1892`: `/api/schools/`
is public and unauthenticated, so leaving a withheld field in the payload hands
the record to anyone who opens the network tab.
## 6. First consumer: `admission_distance`
The last-distance-offered feature is merged to `main` and live on staging.
Production has never received it: `/api/schools/100010` on production carries
no `admission_distance` key, and no Distance section renders.
It needs **exactly one gate** — `backend/app.py:809`, where the field is
attached to the school payload:
```python
"admission_distance": (
supplementary.get("admission_distance")
if flags.is_enabled("admission_distance") else None
),
```
The frontend follows with no change. `DistanceSection` already returns `null`
when `admission_distance?.distance_m == null`, and `PrimarySchoolSections`
already conditions the admissions block on `(admissions || admissionDistance)`.
The off-state is the commonest state on the site — only 57 local authorities
publish cut-off distances at all — so it is well covered by construction.
The flag does not touch the sitemap: school pages exist either way.
Intended lifecycle: default off, so production receives the code dark on the
next promotion; on in the `development` environment so staging keeps testing
it; flipped on in `production` when the owner chooses.
**This flag exercises two of the three surfaces** in §5 — API field and UI
element. No route case ships with it. The route rule is specified but unproven
until a route-shaped flag exists, and should be treated as such.
## 7. Testing
**Backend unit.** The registry is well-formed; an unknown flag evaluates
`False`; `/api/flags` returns every declared flag with its default when the
SDK is unreachable; `admission_distance` is absent from the school payload when
the flag is off and present when on.
**Frontend unit.** `getFlags()` returns declared defaults when `/api/flags`
fails, rather than throwing and taking the page with it.
**E2E.** Journeys read `/api/flags` and gate flag-dependent assertions on it,
matching the `test.skip` shape the suite already uses.
One trap to avoid, worth stating because the existing distance journeys walk
straight into it: they already skip when no school has a published figure, so
with the flag off they would skip silently and the suite would go green. The
gate must be explicit — **if `/api/flags` reports `admission_distance` on, then
a school with a cut-off must be found**, converting a silent skip into a real
assertion.
## 8. Lifecycle
A flag is temporary scaffolding, and the failure mode of every flag system is
accumulation.
The registry records the date each flag was added, and a backend test fails any
flag older than **90 days**. Removing a flag means deleting the registry entry,
the branches that read it, and the flag in the Unleash UI.
Unleash SDK usage metrics stay enabled, so the UI shows which flags are still
being evaluated — the evidence needed to retire one safely.
## 9. Risks
**Production gains a homelab dependency.** If Unleash is unreachable when a
production container cold-starts with an empty cache, every flag evaluates
`False` and any feature currently switched on disappears. The fcache volume
covers restarts; the 90-day lifecycle rule bounds how long any feature is
exposed to this. It is a real regression risk and the reason flags must be
retired rather than left on indefinitely.
**Flag state is not in git.** `main` no longer tells you what production is
showing. The registry lists what *can* be flagged; only the Unleash UI says
what *is*. This is inherent to the choice of a service.
**A large promotion backlog exists.** Production is running the
pre-SEO-programme build — no place pages, and a sitemap still declaring the
apex host. The first promotion after this work ships that entire backlog. The
flag isolates the distance feature from it and nothing else.
## Out of scope
- Percentage rollouts, user targeting, A/B testing, and Unleash strategies
beyond simple on/off. Flags are booleans.
- Pipeline and dbt flags. Airflow and dbt are not flag consumers.
- Client-side flag evaluation. Flags are server-side only.
- Automatic flag removal. The staleness test reports; a person deletes.
+489 -9
View File
@@ -1226,6 +1226,27 @@ const CUTOFF_CANDIDATE_URNS = [
101099, 100553, 102574, 100769, // mixed
];
/**
* Whether the last-distance-offered feature is switched on here.
*
* Read from the data rather than from /api/flags, which the public proxy
* denies on purpose — the endpoint names unreleased features. The observable
* effect is the field's presence: the flag is off iff no candidate school
* carries an `admission_distance` key at all.
*
* The distinction that matters: `admission_distance: null` means this school
* has no published cut-off, and the key being ABSENT means cut-offs are not
* being published at all.
*/
async function distanceFeatureIsOn(page: Page): Promise<boolean> {
for (const urn of CUTOFF_CANDIDATE_URNS) {
const res = await page.request.get(`/api/schools/${urn}`);
if (!res.ok()) continue;
if ('admission_distance' in (await res.json())) return true;
}
return false;
}
async function schoolWithCutoff(page: Page) {
for (const urn of CUTOFF_CANDIDATE_URNS) {
const res = await page.request.get(`/api/schools/${urn}`);
@@ -1237,6 +1258,59 @@ async function schoolWithCutoff(page: Page) {
return null;
}
test('when the distance feature is on, a school with a cut-off is findable', async ({ page }) => {
/*
* The gate that stops the other distance journeys passing vacuously.
*
* They all skip when schoolWithCutoff() finds nothing, which is right when
* the feature is off — but it means a feature that is *supposed* to be on
* and is silently broken shows up as a green run full of skips. This test
* fails in that case.
*/
test.skip(!(await distanceFeatureIsOn(page)),
'the admission_distance flag is off in this environment');
expect(await schoolWithCutoff(page),
'the distance feature is on, but no candidate school has a cut-off — '
+ 'the flag is on and the data or the query behind it is broken')
.not.toBeNull();
});
test('with the distance feature off, the section is absent rather than empty', async ({ page }) => {
// Shipping dark means the page renders as it did before the feature existed,
// not as a feature with its content removed.
test.skip(await distanceFeatureIsOn(page),
'the admission_distance flag is on in this environment');
// A school that exists, found rather than hardcoded — a 404 page would
// satisfy the absent-heading assertion without proving anything.
//
// A plain loop, not Array.find: find's predicate is synchronous, so an async
// one returns a Promise, every Promise is truthy, and it would always hand
// back the first URN whether or not that school exists.
let urn: number | null = null;
for (const candidate of CUTOFF_CANDIDATE_URNS) {
if ((await page.request.get(`/api/schools/${candidate}`)).ok()) {
urn = candidate;
break;
}
}
expect(urn, 'no candidate school resolves in this environment').not.toBeNull();
await page.goto(`/school/${urn}`);
await expect(page.locator('h1')).toBeVisible();
await expect(page.getByRole('heading', { name: /How far away are you\?/ }))
.toHaveCount(0);
});
test('/api/flags is not reachable from the public internet', async ({ page }) => {
// It names every unreleased feature and whether it is on. Next reads it
// server-side over the Docker network; the public proxy must deny it.
const res = await page.request.get('/api/flags');
expect(res.status()).toBe(404);
});
test('a published cut-off distance is shown with the year it belongs to', async ({ page }) => {
const found = await schoolWithCutoff(page);
test.skip(found === null, 'no school in the sample has a published cut-off distance yet');
@@ -1587,18 +1661,130 @@ test('a Welsh school URL 404s while an English one still resolves', async ({ pag
expect(welsh?.status(), 'a Welsh school should no longer resolve').toBe(404);
});
test('the sitemap submits no Welsh or overseas school', async ({ page }) => {
async function sitemapChildren(page: Page): Promise<string[]> {
const res = await page.request.get('/sitemap.xml');
expect(res.ok()).toBeTruthy();
const xml = await res.text();
const index = await res.text();
expect(index).toContain('<sitemapindex');
return [...index.matchAll(/<loc>([^<]+)<\/loc>/g)].map((m) => m[1]);
}
const urlCount = (xml.match(/<url>/g) ?? []).length;
expect(urlCount, 'sitemap looks empty or truncated').toBeGreaterThan(1000);
test('the sitemap index names children that all resolve', async ({ page }) => {
const index = await (await page.request.get('/sitemap.xml')).text();
// An index holds <sitemap> entries only; mixing in <url> is invalid.
expect(index).not.toContain('<url>');
// 401559 (Cardiff) and 402426 (ACT Schools, Cardiff) were both submitted
// before the England-only filter landed.
expect(xml).not.toContain('/school/401559');
expect(xml).not.toContain('/school/402426');
const locs = await sitemapChildren(page);
expect(locs.length).toBeGreaterThanOrEqual(2);
for (const loc of locs) {
expect(loc.startsWith('https://www.schoolcompare.co.uk/sitemaps/')).toBeTruthy();
const child = await page.request.get(new URL(loc).pathname);
expect(child.ok(), `${loc} should resolve`).toBeTruthy();
expect(await child.text()).toContain('<urlset');
}
});
test('the sitemap submits no Welsh or overseas school', async ({ page }) => {
const locs = await sitemapChildren(page);
let total = 0;
for (const loc of locs) {
const xml = await (await page.request.get(new URL(loc).pathname)).text();
total += (xml.match(/<url>/g) ?? []).length;
// 401559 (Adamsdown, Cardiff) and 402426 (ACT Schools, Cardiff) were both
// submitted before the England-only filter landed.
expect(xml).not.toContain('/school/401559');
expect(xml).not.toContain('/school/402426');
}
expect(total, 'sitemap looks empty or truncated').toBeGreaterThan(1000);
});
test('the sitemap invents no priority or changefreq', async ({ page }) => {
const [first] = await sitemapChildren(page);
expect(first).toBeTruthy();
const xml = await (await page.request.get(new URL(first).pathname)).text();
// Google ignores both. They were noise dressed as signal.
expect(xml).not.toContain('<priority>');
expect(xml).not.toContain('<changefreq>');
});
/*
* Canonical URLs (spec 2026-08-20, W1).
*
* Every indexable route declares exactly one canonical, on the www host, with
* no query string. The homepage's eleven search params filter a result set
* rather than making a new document, so they all collapse onto "/".
*/
const CANONICAL_ROUTES: Array<[string, string]> = [
['/', 'https://www.schoolcompare.co.uk/'],
['/rankings', 'https://www.schoolcompare.co.uk/rankings'],
['/admissions', 'https://www.schoolcompare.co.uk/admissions'],
];
/**
* Next normalises canonical URLs against `trailingSlash: false`, so the root
* ships as `https://www.schoolcompare.co.uk` with no slash while every other
* route keeps its path. Both forms address the same document, and which one
* Next emits is its business, not something worth pinning a test to.
*
* The first cut hardcoded the slash and failed only on the homepage — the
* same gap as the doubled brand: it asserted the metadata object rather than
* what the page actually renders.
*/
function sameUrl(a: string | null, b: string): boolean {
const strip = (u: string) => u.replace(/\/+$/, '');
return strip(a ?? '') === strip(b);
}
for (const [path, expected] of CANONICAL_ROUTES) {
test(`${path} declares exactly one canonical, on the www host`, async ({ page }) => {
await page.goto(path);
const hrefs = await page.locator('link[rel="canonical"]').evaluateAll(
(els) => els.map((e) => e.getAttribute('href')));
expect(hrefs, `${path} should declare one canonical`).toHaveLength(1);
expect(sameUrl(hrefs[0], expected),
`${path} canonical was ${hrefs[0]}, expected ${expected}`).toBe(true);
});
}
test('a filtered homepage still canonicalises to the bare root', async ({ page }) => {
await page.goto('/?search=primary&phase=primary&sort=name&page=2');
const href = await page.locator('link[rel="canonical"]').first()
.getAttribute('href');
expect(sameUrl(href, 'https://www.schoolcompare.co.uk/'),
`filtered homepage canonical was ${href}`).toBe(true);
});
test('a school page canonicalises to its own slug on the www host', async ({ page }) => {
const res = await page.request.get('/api/schools?search=primary&per_page=1');
expect(res.ok()).toBeTruthy();
const [first] = (await res.json()).schools ?? [];
expect(first, 'no school available').toBeTruthy();
await page.goto(`/school/${first.urn}-x`);
const href = await page.locator('link[rel="canonical"]').first()
.getAttribute('href');
expect(href).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\/school\/\d+-/);
});
test('a bare /compare is indexable, a parameterised one is not', async ({ page }) => {
await page.goto('/compare');
await expect(page.locator('meta[name="robots"]')).toHaveCount(0);
const [a, b] = await twoPrimaryUrns(page);
await page.goto(`/compare?urns=${a},${b}`);
const robots = await page.locator('meta[name="robots"]').first()
.getAttribute('content');
expect(robots).toContain('noindex');
expect(robots).toContain('follow');
// noindex but follow: the links out to each school page still count, so the
// canonical must still be present and point at the bare path.
const canonical = await page.locator('link[rel="canonical"]').first()
.getAttribute('href');
expect(canonical).toBe('https://www.schoolcompare.co.uk/compare');
});
/*
@@ -1619,10 +1805,33 @@ test('staging answers noindex, and stays crawlable so the noindex is seen', asyn
// The other half, and the reason this is one test rather than two: a
// Disallow would stop Google fetching the page at all, so it would never
// see the noindex above. The two only work together.
//
// Scoped to the `*` group. The first cut matched `Disallow: /` anywhere in
// the file and tripped over the AI-crawler groups Cloudflare injects —
// ClaudeBot, GPTBot, Amazonbot and friends all carry a blanket disallow,
// deliberately, and none of them is Googlebot.
const robots = await (await page.request.get('/robots.txt')).text();
expect(robots).not.toMatch(/^\s*Disallow:\s*\/\s*$/mi);
expect(blocksEverything(robots, '*'),
'the * group must not disallow the whole site, or the noindex is never seen')
.toBe(false);
});
/** True when `agent`'s group in a robots.txt disallows the entire site. */
function blocksEverything(robots: string, agent: string): boolean {
let current: string | null = null;
let blocked = false;
for (const raw of robots.split('\n')) {
const line = raw.split('#')[0].trim();
if (!line) continue;
const [key, ...rest] = line.split(':');
const value = rest.join(':').trim();
const k = key.trim().toLowerCase();
if (k === 'user-agent') current = value;
else if (current === agent && k === 'disallow' && value === '/') blocked = true;
}
return blocked;
}
test('a school page on staging is noindexed too, not just the homepage', async ({ page }) => {
const list = await page.request.get('/api/schools?search=primary&per_page=1');
const [first] = (await list.json()).schools ?? [];
@@ -1631,3 +1840,274 @@ test('a school page on staging is noindexed too, not just the homepage', async (
const res = await page.request.get(`/school/${first.urn}-x`);
expect(res.headers()['x-robots-tag']).toContain('noindex');
});
/*
* W8 — the C1 pages must ship a description, and it must differentiate.
*
* Baseline was 0.43% CTR at position 6.1 on "compare school performance",
* against 9.16% for the brand query from the same neighbourhood. The SERP is
* owned by the DfE's own service, so a description that paraphrases it earns
* nothing. Google may rewrite a snippet, but it cannot use one we never sent.
*/
test('every C1 page ships a description, and none opens its title with the brand', async ({ page }) => {
for (const path of ['/', '/compare', '/rankings', '/admissions']) {
await page.goto(path);
const desc = await page.locator('meta[name="description"]').first()
.getAttribute('content');
expect(desc, `${path} must ship a description`).toBeTruthy();
expect(desc!.length, `${path} description too short to be worth reading`)
.toBeGreaterThan(100);
const title = await page.title();
expect(title.toLowerCase().startsWith('schoolcompare'),
`${path} spends its most valuable pixels on the brand`).toBe(false);
}
});
test('the homepage snippet names what gov.uk does not publish', async ({ page }) => {
await page.goto('/');
const desc = await page.locator('meta[name="description"]').first()
.getAttribute('content');
// Admissions distance is the one fact the DfE service has no equivalent for.
expect(desc).toMatch(/close you had to live|distance/i);
});
/*
* The location layer (spec 2026-08-21, W2).
*
* Location intent sat at position 49.5 with one click across the whole 16-month
* baseline — the site published no page about a place. These assert the four
* families render, stay in their own namespaces, and reach the sitemap.
*/
async function firstPlaceOfKind(page: Page, kind: string) {
const res = await page.request.get('/api/places');
expect(res.ok()).toBeTruthy();
const { places } = await res.json();
const hit = places.find((p: { kind: string }) => p.kind === kind);
expect(hit, `no ${kind} in the registry`).toBeTruthy();
return hit as { kind: string; slug: string; name: string; count: number };
}
for (const [kind, prefix, article] of [
['town', '/schools/', 'a'],
['authority', '/schools/authority/', 'an'],
['outcode', '/schools/near/', 'an'],
] as const) {
test(`${article} ${kind} page renders with its school count`, async ({ page }) => {
const place = await firstPlaceOfKind(page, kind);
await page.goto(`${prefix}${place.slug}`);
await expect(page.locator('h1')).toContainText(place.name, { ignoreCase: true });
await expect(page.locator('a[href^="/school/"]').first()).toBeVisible();
});
}
test('a town and an authority sharing a name are different pages', async ({ page }) => {
// 67 real collisions, and the authority is the larger set in only 43 — so
// one namespace would have published near-duplicates.
const { places } = await (await page.request.get('/api/places')).json();
const townSlugs = new Set(
places.filter((p: { kind: string }) => p.kind === 'town')
.map((p: { slug: string }) => p.slug));
const clash = places.find((p: { kind: string; slug: string }) =>
p.kind === 'authority' && townSlugs.has(p.slug));
test.skip(!clash, 'no town/authority name collision in this environment');
const townRes = await page.request.get(`/api/places/town/${clash.slug}`);
const laRes = await page.request.get(`/api/places/authority/${clash.slug}`);
expect(townRes.ok() && laRes.ok()).toBeTruthy();
const townUrns = (await townRes.json()).schools.map((s: { urn: number }) => s.urn).sort();
const laUrns = (await laRes.json()).schools.map((s: { urn: number }) => s.urn).sort();
expect(townUrns).not.toEqual(laUrns);
});
test('a place below the threshold has no page', async ({ page }) => {
// Crosby holds one school; publishing it would be a page with nothing to say.
const res = await page.request.get('/api/places/town/crosby');
expect(res.status()).toBe(404);
});
test('place pages declare a canonical and reach the sitemap', async ({ page }) => {
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
const canonical = await page.locator('link[rel="canonical"]').first()
.getAttribute('href');
expect(canonical).toBe(`https://www.schoolcompare.co.uk/schools/${place.slug}`);
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toContain(`/schools/${place.slug}`);
});
test('the place page ships ItemList structured data that parses', async ({ page }) => {
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
const raw = await page.locator('script[type="application/ld+json"]').first()
.textContent();
const parsed = JSON.parse(raw!);
const types = (parsed['@graph'] ?? []).map((n: { '@type': string }) => n['@type']);
expect(types).toContain('ItemList');
expect(types).toContain('BreadcrumbList');
});
test('a place page states the local average against England', async ({ page }) => {
// The one number a list cannot give, and the reason these pages are not
// a name dropped into a template.
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
await expect(page.getByTestId('local-vs-england')).toContainText(/across England/i);
});
test('a place page links its phase variants, and they resolve', async ({ page }) => {
// "primary schools in beccles" is the query shape the baseline showed. The
// first cut submitted only the bare place URL and linked nothing, leaving
// ~950 variant pages reachable by nothing at all.
const res = await page.request.get('/api/places');
const { places } = await res.json();
const town = places.find((p: { kind: string }) => p.kind === 'town');
expect(town).toBeTruthy();
const detail = await (await page.request.get(`/api/places/town/${town.slug}`)).json();
test.skip(!(detail.place.phases ?? []).length, 'no phase clears the threshold here');
await page.goto(`/schools/${town.slug}`);
const phase = detail.place.phases[0];
const link = page.locator(`a[href="/schools/${town.slug}/${phase}"]`).first();
await expect(link).toBeVisible();
await link.click();
await expect(page.locator('h1')).toContainText(new RegExp(`${phase} schools in`, 'i'));
});
test('phase variants are submitted in the places sitemap', async ({ page }) => {
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toMatch(/\/schools\/[a-z0-9-]+\/primary</);
});
test('authority phase variants are submitted, and in their own namespace', async ({ page }) => {
// 302 of these were in the sitemap for weeks and every one 404'd: the spec
// called for the route, the plan built the bare authority page and dropped
// it, and the sitemap — written from the registry — kept submitting them.
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toMatch(/\/schools\/authority\/[a-z0-9-]+\/primary</);
});
test('every place link a place page emits resolves', async ({ page }) => {
/*
* The guard that was missing. Each family built its own links, so a URL
* shape belonging to one namespace was used by all four: an authority page
* offered "Primary schools in Barnet" pointing at /schools/barnet/primary,
* the *town*. For 87 of 151 authorities that 404'd; for the other 64 it
* quietly served a different set of schools under the same name.
*
* Only /schools links are followed. The per-school links are the same
* component the school-page journeys already cover, and there are hundreds
* of them on a page.
*/
for (const kind of ['town', 'authority', 'outcode'] as const) {
const place = await firstPlaceOfKind(page, kind);
const prefix = kind === 'authority' ? '/schools/authority/'
: kind === 'outcode' ? '/schools/near/' : '/schools/';
await page.goto(`${prefix}${place.slug}`);
const hrefs = [...new Set(
await page.locator('a[href^="/schools"]').evaluateAll(
(els) => els.map((e) => e.getAttribute('href')!)))];
expect(hrefs.length, `${kind} page links no other place`).toBeGreaterThan(0);
for (const href of hrefs) {
const res = await page.request.get(href);
expect(res.status(), `${kind} page links ${href}`).toBe(200);
}
}
});
test('an outcode page offers no phase link, because no such page exists', async ({ page }) => {
// Nobody searches "primary schools in SW11", so the spec gives outcodes no
// phase route. The registry computed the variants anyway and the page
// linked them, putting two 404s on each of 1,720 outcode pages.
const place = await firstPlaceOfKind(page, 'outcode');
const detail = await (await page.request.get(
`/api/places/outcode/${place.slug}`)).json();
expect(detail.place.phases).toEqual([]);
await page.goto(`/schools/near/${place.slug}`);
await expect(page.getByRole('navigation', { name: 'By phase' })).toHaveCount(0);
});
test('no page title repeats the brand', async ({ page }) => {
// The root layout appends '| schoolcompare' to a plain-string title. Any
// route whose title already carries the brand must opt out with
// `absolute`, or it ships '... | schoolcompare | schoolcompare' — which is
// how ~2,600 place pages first went out.
const res = await page.request.get('/api/places');
const { places } = await res.json();
const town = places.find((p: { kind: string }) => p.kind === 'town');
for (const path of ['/', '/rankings', '/admissions', `/schools/${town.slug}`]) {
await page.goto(path);
const title = await page.title();
const brands = (title.match(/schoolcompare/gi) ?? []).length;
expect(brands, `${path} repeats the brand: ${title}`).toBeLessThanOrEqual(1);
}
});
test('a place straddling a boundary names every authority it sits in', async ({ page }) => {
// A quarter of outcodes and a third of towns cross an authority boundary —
// SW19 is mostly Merton but partly Wandsworth. Naming only the largest
// asserts something false about the place.
const { places } = await (await page.request.get('/api/places')).json();
const outcode = places.find((p: { kind: string }) => p.kind === 'outcode');
expect(outcode).toBeTruthy();
// Find any place the registry reports as straddling.
let straddling: { kind: string; slug: string } | null = null;
for (const p of places.filter((p: { kind: string }) => p.kind === 'outcode').slice(0, 40)) {
const d = await (await page.request.get(`/api/places/outcode/${p.slug}`)).json();
if ((d.place.authorities ?? []).length > 1) { straddling = p; break; }
}
test.skip(!straddling, 'no straddling outcode found in the sample');
const detail = await (await page.request.get(
`/api/places/outcode/${straddling!.slug}`)).json();
await page.goto(`/schools/near/${straddling!.slug}`);
for (const a of detail.place.authorities) {
if (a.slug) {
await expect(page.locator(`a[href="/schools/authority/${a.slug}"]`).first())
.toBeVisible();
} else {
// No page of its own — City of London and the Isles of Scilly are
// under the threshold. Named, deliberately not linked.
await expect(page.locator('header p')).toContainText(a.name);
await expect(page.getByRole('link', { name: a.name })).toHaveCount(0);
}
}
});
test('a place page lists its schools alphabetically', async ({ page }) => {
// Someone on a place page is usually looking for a school they can name,
// so the order should serve scanning for it. /rankings is where the
// league-table ordering lives.
const { places } = await (await page.request.get('/api/places')).json();
const town = places.find((p: { kind: string; count: number }) =>
p.kind === 'town' && p.count >= 5);
expect(town).toBeTruthy();
await page.goto(`/schools/${town.slug}`);
const names = await page.locator('a[href^="/school/"]').allTextContents();
expect(names.length).toBeGreaterThan(1);
const sorted = [...names].sort((a, b) =>
a.toLowerCase().localeCompare(b.toLowerCase()));
expect(names).toEqual(sorted);
});
test('the rankings page still orders by score, not name', async ({ page }) => {
// Alphabetical is a place-page decision, not a site-wide one.
const res = await page.request.get('/api/rankings?metric=rwm_expected_pct&phase=primary');
expect(res.ok()).toBeTruthy();
const scores = ((await res.json()).rankings ?? [])
.map((r: { rwm_expected_pct: number | null }) => r.rwm_expected_pct)
.filter((v: number | null) => v != null);
expect(scores).toEqual([...scores].sort((a: number, b: number) => b - a));
});
@@ -0,0 +1,31 @@
/**
* The /api/* proxy is public. Anything it forwards is on the internet.
*
* @jest-environment node
*/
// The docblock above is load-bearing. jest.config.js sets jsdom globally, and
// NextRequest/NextResponse need the Web Fetch API globals that only the node
// environment provides — under jsdom this suite fails on import, not on an
// assertion.
import { NextRequest } from 'next/server';
import { GET } from '@/app/api/[...path]/route';
function request(path: string) {
return new NextRequest(`http://localhost:3000/api/${path}`);
}
describe('public API proxy', () => {
it('refuses to forward internal-only paths', async () => {
// /api/flags names every unreleased feature and its state. Forwarding it
// publishes the thing shipping dark exists to keep quiet.
const res = await GET(request('flags'), { params: Promise.resolve({ path: ['flags'] }) });
expect(res.status).toBe(404);
});
it('does not deny a path that merely starts with the same letters', async () => {
// A prefix match would take /api/flagship down with /api/flags.
const res = await GET(
request('flagship'), { params: Promise.resolve({ path: ['flagship'] }) });
expect(res.status).not.toBe(404);
});
});
+130
View File
@@ -0,0 +1,130 @@
import { metadata as homeMetadata } from '@/app/page';
import { metadata as rankingsMetadata } from '@/app/rankings/page';
import { metadata as admissionsMetadata } from '@/app/admissions/page';
import { generateMetadata as compareMetadata } from '@/app/compare/page';
describe('canonical URLs', () => {
it('the homepage canonicalises to the bare root', () => {
// page.tsx reads eleven search params. Without this, every filter
// combination is a crawlable near-duplicate of the one page we want to
// rank for "compare schools".
expect(homeMetadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/');
});
it('rankings canonicalises to the bare path', () => {
expect(rankingsMetadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/rankings');
});
it('admissions canonicalises to the bare path', () => {
expect(admissionsMetadata.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/admissions');
});
});
describe('/compare indexability', () => {
it('the bare compare page is indexable and canonical to itself', async () => {
// This is the landing page for the "compare schools" head term.
const meta = await compareMetadata({ searchParams: Promise.resolve({}) });
expect(meta.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/compare');
expect(meta.robots).toBeUndefined();
});
it('a comparison of specific schools is noindex, follow', async () => {
// ~317 million pairs before triples. Indexing the parameter space would
// swamp everything else in the corpus.
const meta = await compareMetadata({
searchParams: Promise.resolve({ urns: '100001,100002' }),
});
expect(meta.robots).toEqual({ index: false, follow: true });
});
it('a parameterised comparison still canonicalises to the bare path', async () => {
// follow:true plus a canonical means the outbound links to each school
// page still pass value even though this URL is not indexed.
const meta = await compareMetadata({
searchParams: Promise.resolve({ urns: '100001,100002' }),
});
expect(meta.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/compare');
});
});
/*
* W8 — snippet copy for the C1 cluster.
*
* The baseline (GSC, 16 months to 2026-08-20) showed these pages ranking on
* page one and converting at a tenth of the normal rate: "compare school
* performance" at position 6.1 with 0.43% CTR, against 9.16% for the brand
* query from the same neighbourhood. The SERP is dominated by the DfE's own
* "Compare school performance" service, so the job of this copy is to say
* what that service does not offer, without losing intent match on the title.
*
* These tests guard the mechanics that make a snippet work — length, intent
* keyword, differentiator, no brand-first — not the exact wording, which
* should stay free to iterate.
*/
// Google truncates titles near 60 characters and descriptions near 155.
const TITLE_MAX = 60;
const DESC_MIN = 110;
const DESC_MAX = 155;
type Meta = { title?: unknown; description?: unknown };
const titleOf = (m: Meta): string => {
const t = m.title as string | { absolute?: string } | undefined;
return typeof t === 'string' ? t : (t?.absolute ?? '');
};
describe('C1 snippet copy', () => {
const pages: Array<[string, Meta, RegExp]> = [
['home', homeMetadata as Meta, /compare schools/i],
['rankings', rankingsMetadata as Meta, /league table/i],
['admissions', admissionsMetadata as Meta, /admission/i],
];
for (const [name, meta, intent] of pages) {
it(`${name}: title carries the search intent and fits the SERP`, () => {
const t = titleOf(meta);
expect(t).toMatch(intent);
expect(t.length).toBeLessThanOrEqual(TITLE_MAX);
});
it(`${name}: title does not open with the brand`, () => {
// The measured 0.43% CTR came from a brand-first title. The most
// valuable pixels go to the thing the searcher typed.
expect(titleOf(meta).toLowerCase().startsWith('schoolcompare')).toBe(false);
});
it(`${name}: description is long enough to be worth reading, short enough to survive`, () => {
const d = meta.description as string;
expect(d.length).toBeGreaterThanOrEqual(DESC_MIN);
expect(d.length).toBeLessThanOrEqual(DESC_MAX);
});
}
it('the homepage description names what gov.uk does not publish', () => {
// Admissions distance is the one fact the DfE service has no equivalent
// for. If it ever leaves this description, the snippet is competing with
// gov.uk on gov.uk's own ground.
expect(homeMetadata.description).toMatch(/close you had to live|distance/i);
});
it('/compare targets the tool phrasing rather than repeating the homepage', () => {
// Two pages chasing one phrase is how a site competes with itself.
return compareMetadata({ searchParams: Promise.resolve({}) }).then((m) => {
expect(m.title).toMatch(/comparison tool/i);
expect(m.title).not.toBe(titleOf(homeMetadata as Meta));
});
});
it('no C1 page claims a school count that will drift', () => {
// The corpus moves with every data refresh; this repo has already shipped
// one copy bug of that kind ("three schools" against MAX_SCHOOLS = 5).
for (const [, meta] of pages) {
expect(meta.description as string).not.toMatch(/\b\d{2},\d{3}\b|\b\d{2},000\b/);
}
});
});
@@ -0,0 +1,42 @@
import { generateMetadata as placeMeta } from '@/app/schools/[place]/page';
jest.mock('@/lib/places', () => ({
...jest.requireActual('@/lib/places'),
fetchPlace: jest.fn(async (kind: string, slug: string) =>
slug === 'atlantis' ? null : ({
place: { kind, slug, name: 'Brentwood', count: 29,
parent_authority: 'Essex' },
schools: [], averages: { rwm_expected_pct: 63, attainment_8_score: null },
})),
fetchPlaces: jest.fn(async () => []),
}));
describe('place page metadata', () => {
it('titles the page the way the place is searched', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'brentwood' }) });
expect((m.title as { absolute: string }).absolute).toMatch(/schools in brentwood/i);
});
it('canonicalises to its own path on the www host', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'brentwood' }) });
expect(m.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/schools/brentwood');
});
it('opts out of the layout template, which would double the brand', () => {
// The root layout appends '| schoolcompare' to a plain string title, and
// these titles already carry it — every place page shipped reading
// '... | schoolcompare | schoolcompare' until this was made absolute.
return placeMeta({ params: Promise.resolve({ place: 'brentwood' }) })
.then((m) => {
expect(typeof m.title).toBe('object');
expect((m.title as { absolute: string }).absolute)
.not.toMatch(/schoolcompare.*schoolcompare/);
});
});
it('an unknown place gets a not-found title rather than inventing one', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'atlantis' }) });
expect(m.title).toMatch(/not found/i);
});
});
@@ -0,0 +1,348 @@
import { render, screen } from '@testing-library/react';
import { PlaceView } from '@/components/places/PlaceView';
import type { PlaceDetail } from '@/lib/places';
const detail: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 29,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', rwm_expected_pct: 82,
ofsted_grade: 1, phase: 'Primary' } as never,
{ urn: 2, school_name: 'Beta Primary', rwm_expected_pct: 44,
ofsted_grade: 3, phase: 'Primary' } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
describe('PlaceView', () => {
it('leads with an H1 that matches how the place is searched', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByRole('heading', { level: 1 }))
.toHaveTextContent(/primary schools in brentwood/i);
});
it('states the count so the page says something before the table', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByText(/29 schools/i)).toBeInTheDocument();
});
it('compares the local average against England, which a list cannot', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByTestId('local-vs-england')).toHaveTextContent('63');
expect(screen.getByTestId('local-vs-england')).toHaveTextContent('61');
});
it('links every school in scope, which is what de-orphans them', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getAllByRole('link', { name: /Primary$/ })).toHaveLength(2);
});
it('links to the parent authority so the place sits in a hierarchy', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByRole('link', { name: /Essex/i }))
.toHaveAttribute('href', '/schools/authority/essex');
});
it('shows the Ofsted distribution, not just a count of Outstanding', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByTestId('ofsted-distribution')).toBeInTheDocument();
});
it('links to neighbouring places so the page is not a dead end', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[{ kind: 'town', slug: 'romford', name: 'Romford', count: 40 }]} />);
expect(screen.getByRole('link', { name: /Romford/ }))
.toHaveAttribute('href', '/schools/romford');
});
it('says nothing about an average it does not have', () => {
render(<PlaceView detail={{ ...detail, averages:
{ rwm_expected_pct: null, attainment_8_score: null } }}
phase="primary" englandAverage={61} neighbours={[]} />);
expect(screen.queryByTestId('local-vs-england')).not.toBeInTheDocument();
});
});
describe('PlaceView structured data', () => {
function jsonLd() {
const { container } = render(<PlaceView detail={detail} phase="primary"
englandAverage={61} neighbours={[]} />);
const el = container.querySelector('script[type="application/ld+json"]');
return JSON.parse(el!.textContent!);
}
it('declares the page as a ranked list, not prose', () => {
const types = jsonLd()['@graph'].map((n: { '@type': string }) => n['@type']);
expect(types).toContain('ItemList');
expect(types).toContain('BreadcrumbList');
});
it('gives every listed school an absolute URL on the canonical host', () => {
const list = jsonLd()['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListElement).toHaveLength(2);
for (const item of list.itemListElement) {
expect(item.url).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\/school\//);
}
});
});
describe('PlaceView phase variants', () => {
it('links the phase variants that exist', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /Primary schools in Brentwood/i }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('links no variant for a phase below its own threshold', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.queryByRole('link', { name: /Secondary schools in Brentwood/i }))
.not.toBeInTheDocument();
});
it('does not link sideways from a variant page to itself', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.queryByRole('link', { name: /Primary schools in Brentwood/i }))
.not.toBeInTheDocument();
});
});
describe('PlaceView presentation', () => {
// /schools/brentwood shipped with 8 of 27 rows blank: an unphased page shows
// one primary-only measure for a list that also holds secondaries.
const mixed: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 4,
parent_authority: 'Essex', phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta High', phase: 'Secondary',
rwm_expected_pct: null, attainment_8_score: 47 } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: 45 },
};
it('gives each phase its own table rather than one column of blanks', () => {
render(<PlaceView detail={mixed} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('heading', { name: /^Primary schools/ })).toBeInTheDocument();
expect(screen.getByRole('heading', { name: /^Secondary schools/ })).toBeInTheDocument();
expect(screen.getByText('82%')).toBeInTheDocument();
expect(screen.getByText('47')).toBeInTheDocument();
});
it('names the measure in plain words, not jargon', () => {
// The first cut said "RWM expected", which appears nowhere else on the site.
render(<PlaceView detail={mixed} englandAverage={61} neighbours={[]} />);
expect(screen.getByText('Reading, writing & maths')).toBeInTheDocument();
expect(screen.getByText('Attainment 8')).toBeInTheDocument();
expect(screen.queryByText(/RWM expected/i)).not.toBeInTheDocument();
});
it('says a missing result is unpublished rather than showing a bare dash', () => {
const noResult: PlaceDetail = {
...mixed,
schools: [{ urn: 3, school_name: 'New Primary', phase: 'Primary',
rwm_expected_pct: null, attainment_8_score: null } as never],
};
render(<PlaceView detail={noResult} englandAverage={61} neighbours={[]} />);
expect(screen.getByText('Not published')).toBeInTheDocument();
});
it('styles school links to the site convention rather than browser default', () => {
const { container } = render(<PlaceView detail={mixed} englandAverage={61}
neighbours={[]} />);
const link = container.querySelector('a[href^="/school/"]');
expect(link?.className).toBeTruthy();
});
it('a phased page shows one table and no phase headings', () => {
render(<PlaceView detail={mixed} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.queryByRole('heading', { name: /^Secondary schools/ }))
.not.toBeInTheDocument();
});
});
describe('PlaceView table alignment', () => {
const aligned: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('aligns the measure heading and its values with the same class', () => {
// They were aligned by two different selectors whose specificity did not
// match: `.table th:last-child` (0,2,1) won and went right, while `.num`
// (0,1,0) lost to `.table td` (0,1,1) and stayed left. Sharing one class
// is what makes them impossible to drift apart.
const { container } = render(<PlaceView detail={aligned} englandAverage={61}
neighbours={[]} />);
const th = container.querySelectorAll('th')[1];
const td = container.querySelectorAll('tbody td')[1];
expect(th.className).toBeTruthy();
expect(td.className).toBe(th.className);
});
it('leaves the school-name column unclassed so it takes the spare width', () => {
const { container } = render(<PlaceView detail={aligned} englandAverage={61}
neighbours={[]} />);
expect(container.querySelectorAll('th')[0].className).toBe('');
});
});
describe('PlaceView authorities', () => {
const straddling: PlaceDetail = {
place: { kind: 'outcode', slug: 'sw19', name: 'SW19', count: 33,
parent_authority: 'Merton', phases: ['primary'],
authorities: [
{ name: 'Merton', slug: 'merton', count: 26 },
{ name: 'Wandsworth', slug: 'wandsworth', count: 7 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names every authority the place straddles, not just the largest', () => {
// SW19 is mostly Merton but partly Wandsworth. Naming one asserts
// something false about a quarter of outcodes.
render(<PlaceView detail={straddling} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' }))
.toHaveAttribute('href', '/schools/authority/merton');
expect(screen.getByRole('link', { name: 'Wandsworth' }))
.toHaveAttribute('href', '/schools/authority/wandsworth');
});
it('joins them readably rather than as a bare list', () => {
// Asserted on the summary line's whole text: a loose /and/ matcher also
// hits "Wandsworth".
const { container } = render(<PlaceView detail={straddling}
englandAverage={61} neighbours={[]} />);
const summary = container.querySelector('header p');
expect(summary?.textContent).toContain('Merton and Wandsworth');
});
it('falls back to the single parent when the field is absent', () => {
// A cached API response predating the authorities field must not blank
// the line entirely.
const legacy = { ...straddling,
place: { ...straddling.place, authorities: undefined } };
render(<PlaceView detail={legacy} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' })).toBeInTheDocument();
});
});
describe('PlaceView list ordering', () => {
const detail3: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 40, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta Primary', phase: 'Primary',
rwm_expected_pct: 90, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 65, attainment_8_score: null },
};
it('renders schools in the order the API sent them, not by score', () => {
// The API sorts alphabetically now; the component must not re-sort.
render(<PlaceView detail={detail3} englandAverage={61} neighbours={[]} />);
const links = screen.getAllByRole('link', { name: /Primary$/ });
expect(links.map((l) => l.textContent))
.toEqual(['Alpha Primary', 'Beta Primary']);
});
it('declares the list as ascending rather than implying a ranking', () => {
// An ItemList carrying `position` reads as a ranking unless it says
// otherwise, and the table is A-Z.
const { container } = render(<PlaceView detail={detail3} englandAverage={61}
neighbours={[]} />);
const ld = JSON.parse(
container.querySelector('script[type="application/ld+json"]')!.textContent!);
const list = ld['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListOrder).toBe('https://schema.org/ItemListOrderAscending');
});
});
describe('PlaceView phase links', () => {
const authority: PlaceDetail = {
place: { kind: 'authority', slug: 'barnet', name: 'Barnet', count: 156,
parent_authority: null, phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('keeps an authority phase link in the authority namespace', () => {
// The link was built as `/schools/${slug}/${phase}` for every kind, so an
// authority page pointed into the town namespace. For 87 of 151
// authorities that 404'd; for the other 64 it silently landed on the town
// page of the same name — a different set of schools, and exactly the
// duplicate the two namespaces exist to prevent. Barnet is one of the 64.
render(<PlaceView detail={authority} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /^Primary schools in Barnet$/ }))
.toHaveAttribute('href', '/schools/authority/barnet/primary');
expect(screen.getByRole('link', { name: /^Secondary schools in Barnet$/ }))
.toHaveAttribute('href', '/schools/authority/barnet/secondary');
});
it('still uses the bare namespace for a town', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /^Primary schools in Brentwood$/ }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('offers no phase link when the place publishes none', () => {
// Outcodes are the case: no phase route exists for them, so the registry
// reports no phases and the nav does not render.
const outcode = { ...detail,
place: { ...detail.place, kind: 'outcode', slug: 'cm13', name: 'CM13',
phases: [] } };
render(<PlaceView detail={outcode} englandAverage={61} neighbours={[]} />);
expect(screen.queryByRole('navigation', { name: 'By phase' }))
.not.toBeInTheDocument();
});
});
describe('PlaceView unlinkable authorities', () => {
const withUnpublished: PlaceDetail = {
place: { kind: 'outcode', slug: 'tr21', name: 'TR21', count: 8,
parent_authority: 'Cornwall', phases: [],
authorities: [
{ name: 'Cornwall', slug: 'cornwall', count: 6 },
{ name: 'Isles Of Scilly', slug: null, count: 2 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names an authority with no page without linking it', () => {
// City of London and the Isles of Scilly hold fewer schools than a page
// needs. Saying where the place is stays right; linking there would 404.
const { container } = render(<PlaceView detail={withUnpublished}
englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Cornwall' })).toBeInTheDocument();
expect(screen.queryByRole('link', { name: 'Isles Of Scilly' }))
.not.toBeInTheDocument();
expect(container.querySelector('header p')?.textContent)
.toContain('Isles Of Scilly');
});
});
+34
View File
@@ -0,0 +1,34 @@
import { getFlags } from '@/lib/flags';
// jsdom provides no global fetch, so there is nothing for jest.spyOn to attach
// to — assign it and restore the original afterwards. This is the first test
// here to mock fetch; later ones should follow this shape.
const realFetch = global.fetch;
function mockFetch(impl: () => Promise<unknown>) {
global.fetch = jest.fn(impl) as unknown as typeof fetch;
}
describe('getFlags', () => {
afterEach(() => { global.fetch = realFetch; });
it('returns the flags the API reports', async () => {
mockFetch(async () => ({
ok: true,
json: async () => ({ admission_distance: true }),
}));
await expect(getFlags()).resolves.toEqual({ admission_distance: true });
});
it('returns no flags rather than throwing when the API is down', async () => {
// A page that cannot read flags must render everything dark, not 500.
// Fail-closed is the same direction as the backend's default.
mockFetch(async () => { throw new Error('ECONNREFUSED'); });
await expect(getFlags()).resolves.toEqual({});
});
it('returns no flags rather than throwing on a non-200', async () => {
mockFetch(async () => ({ ok: false, status: 503 }));
await expect(getFlags()).resolves.toEqual({});
});
});
+27
View File
@@ -0,0 +1,27 @@
import { SITE_URL, absoluteUrl } from '@/lib/site';
describe('SITE_URL', () => {
it('is the www host, which is the one that serves a 200', () => {
// The apex 301s to www at Cloudflare. A canonical pointing at a redirect
// is a wasted signal, so every absolute URL we emit must already be www.
expect(SITE_URL).toBe('https://www.schoolcompare.co.uk');
});
it('has no trailing slash, so joins never double up', () => {
expect(SITE_URL.endsWith('/')).toBe(false);
});
});
describe('absoluteUrl', () => {
it('joins a rooted path', () => {
expect(absoluteUrl('/rankings')).toBe('https://www.schoolcompare.co.uk/rankings');
});
it('joins a path missing its leading slash', () => {
expect(absoluteUrl('rankings')).toBe('https://www.schoolcompare.co.uk/rankings');
});
it('maps the site root to a bare trailing slash', () => {
expect(absoluteUrl('/')).toBe('https://www.schoolcompare.co.uk/');
});
});
+6 -2
View File
@@ -1,12 +1,16 @@
import { absoluteUrl } from '@/lib/site';
import type { Metadata } from 'next';
import { AdmissionsView } from '@/components/AdmissionsView';
export const dynamic = 'force-static';
export const metadata: Metadata = {
title: 'School Admissions Guide',
// Deadlines and offer days are what gets searched, and what this page is
// genuinely best at — the countdowns are live.
title: { absolute: 'School Admissions Deadlines & Offer Days | schoolcompare' },
description:
'Understand the Primary and Secondary school admissions process in England, with live countdowns to every key deadline and National Offer Day.',
'Every key date for primary and secondary school admissions in England, with live countdowns to the application deadline and National Offer Day.',
alternates: { canonical: absoluteUrl('/admissions') },
};
export default function AdmissionsPage() {
+19
View File
@@ -26,8 +26,27 @@ function backendBase(): string {
const STRIPPED_RESPONSE_HEADERS = ['content-encoding', 'content-length', 'transfer-encoding', 'connection'];
const METHODS_WITH_BODY = new Set(['POST', 'PUT', 'PATCH', 'DELETE']);
/*
* API paths this public proxy must not forward.
*
* Matched on the first segment, exactly — a prefix match would take
* /api/flagship down with /api/flags.
*
* `flags` is here because GET /api/flags names every unreleased feature the
* codebase knows about, along with whether it is on. Publishing that defeats
* the point of shipping dark. Next reads it server-side via FASTAPI_URL, on
* the Docker network, which never transits this route.
*
* Anything else internal-only belongs here too.
*/
const INTERNAL_ONLY_SEGMENTS = new Set(['flags']);
async function handler(req: NextRequest, ctx: { params: Promise<{ path: string[] }> }) {
const { path } = await ctx.params;
if (INTERNAL_ONLY_SEGMENTS.has(path[0])) {
return NextResponse.json({ detail: 'Not Found' }, { status: 404 });
}
const target = `${backendBase()}/${path.join('/')}${req.nextUrl.search}`;
const headers = new Headers(req.headers);
+31 -7
View File
@@ -5,6 +5,7 @@
import { fetchComparison, fetchMetrics } from '@/lib/api';
import { ComparisonView } from '@/components/ComparisonView';
import { absoluteUrl } from '@/lib/site';
import type { Metadata } from 'next';
interface ComparePageProps {
@@ -14,13 +15,36 @@ interface ComparePageProps {
}>;
}
export const metadata: Metadata = {
title: 'Compare Schools',
description:
'Compare schools in England side by side — Ofsted inspections, KS2 and GCSE results against the England average, admissions odds and school community.',
keywords:
'school comparison, compare schools, Ofsted comparison, school admissions, KS2 comparison, primary school performance',
};
/**
* Indexability depends on the query string, so this cannot be a static export.
*
* Bare /compare is the landing page for the "compare schools" head term and
* stays indexable. /compare?urns=… is an unbounded parameter space — 25,193
* schools make ~317 million pairs — so it goes noindex. It stays `follow` and
* keeps a canonical to the bare path, so the links out to each school page
* still count.
*/
export async function generateMetadata(
{ searchParams }: ComparePageProps,
): Promise<Metadata> {
const { urns } = await searchParams;
const base: Metadata = {
// Deliberately not the homepage's phrase. Two pages chasing "compare
// schools" is how a site competes with itself; this one takes the tool
// phrasing instead.
title: 'School Comparison Tool — Up to Five at Once | schoolcompare',
description:
'Put up to five English schools in one table: SATs and GCSE results against the England average, Ofsted grades, and the distance places were offered.',
keywords:
'school comparison, compare schools, Ofsted comparison, school admissions, KS2 comparison, primary school performance',
alternates: { canonical: absoluteUrl('/compare') },
};
if (!urns) return base;
return { ...base, robots: { index: false, follow: true } };
}
// Dynamic via searchParams; remove force-dynamic so internal data fetches
// can still use Next.js's per-call revalidate cache.
+12 -8
View File
@@ -5,6 +5,7 @@ import { Navigation } from '@/components/Navigation';
import { Footer } from '@/components/Footer';
import { ComparisonToast } from '@/components/ComparisonToast';
import { ComparisonProvider } from '@/context/ComparisonProvider';
import { SITE_URL } from '@/lib/site';
import './globals.css';
// Manrope carries headings and key messaging — the guideline's "friendly,
@@ -47,29 +48,32 @@ export const metadata: Metadata = {
statusBarStyle: 'default',
},
title: {
default: 'schoolcompare | Compare School Performance',
default: 'Compare Schools Side by Side | schoolcompare',
template: '%s | schoolcompare',
},
description: 'Compare primary and secondary school SATs and GCSE performance across England',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place. Free, no sign-up.',
keywords: 'school comparison, KS2 results, KS4 results, primary school, secondary school, England schools, SATs results, GCSE results',
authors: [{ name: 'schoolcompare' }],
manifest: '/manifest.json',
// No `icons` key on purpose: setting it here would override the file
// conventions. app/icon.svg and app/apple-icon.tsx are the source, and
// app/opengraph-image.tsx supplies og:image and twitter:image.
metadataBase: new URL('https://schoolcompare.co.uk'),
metadataBase: new URL(SITE_URL),
openGraph: {
type: 'website',
title: 'schoolcompare | Compare School Performance',
description: 'Compare primary and secondary school SATs and GCSE performance across England',
url: 'https://schoolcompare.co.uk',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
url: SITE_URL,
siteName: 'schoolcompare',
},
twitter: {
// summary_large_image now that there is an image worth showing.
card: 'summary_large_image',
title: 'schoolcompare | Compare School Performance',
description: 'Compare primary and secondary school SATs and GCSE performance across England',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
},
};
+20 -2
View File
@@ -3,6 +3,7 @@
* Main landing page with school search and browsing
*/
import { absoluteUrl } from '@/lib/site';
import type { Metadata } from 'next';
import { fetchSchools, fetchFilters, fetchDataInfo } from '@/lib/api';
import { formatAcademicYear } from '@/lib/utils';
@@ -33,8 +34,25 @@ interface HomePageProps {
* saying the brand twice.
*/
export const metadata: Metadata = {
title: { absolute: 'schoolcompare | Compare every school in England' },
description: 'Search and compare school performance across England',
/*
* Intent in the title, differentiator in the description.
*
* These queries are owned by the DfE's own "Compare school performance"
* service, and the old title — brand first, then a near-paraphrase of that
* service's name — gave a searcher no reason to pick us over it. It drew
* 0.43% CTR at position 6.1 while the brand query drew 9.16% from the same
* neighbourhood, so the ranking was never the problem.
*
* The title now matches what people type. The description carries the one
* fact gov.uk does not publish: how close you had to live to get a place.
*/
title: { absolute: 'Compare Schools Side by Side | schoolcompare' },
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place. Free, no sign-up.',
// This page reads eleven search params. They filter a result set; they do
// not make a new document. Collapsing every combination onto "/" stops the
// homepage competing with itself for its own head terms.
alternates: { canonical: absoluteUrl('/') },
};
// The page reads searchParams, which makes rendering dynamic by default.
+9 -2
View File
@@ -5,6 +5,7 @@
import { fetchRankings, fetchFilters, fetchMetrics } from '@/lib/api';
import { RankingsView } from '@/components/RankingsView';
import { absoluteUrl } from '@/lib/site';
import type { Metadata } from 'next';
interface RankingsPageProps {
@@ -17,9 +18,15 @@ interface RankingsPageProps {
}
export const metadata: Metadata = {
title: 'School Rankings',
description: 'Top-ranked schools by SATs and GCSE performance across England',
// 'School Rankings' matched nothing anyone types. League tables is the
// phrase parents actually search, and it spikes each results day.
title: { absolute: 'Primary & Secondary School League Tables | schoolcompare' },
description:
'Rank English schools by SATs results, GCSEs, Progress 8 or Attainment 8, and filter by local authority or year. Built from the DfE’s own figures.',
keywords: 'school rankings, top schools, best schools, KS2 rankings, KS4 rankings, school league tables',
// Param forms (?metric=&local_authority=&year=&phase=) collapse here for
// now. W3 replaces them with real indexable paths.
alternates: { canonical: absoluteUrl('/rankings') },
};
// Dynamic via searchParams; remove force-dynamic so internal data fetches
+2 -1
View File
@@ -4,6 +4,7 @@
*/
import { MetadataRoute } from 'next';
import { absoluteUrl } from '@/lib/site';
export default function robots(): MetadataRoute.Robots {
return {
@@ -14,6 +15,6 @@ export default function robots(): MetadataRoute.Robots {
disallow: ['/api/', '/_next/'],
},
],
sitemap: 'https://schoolcompare.co.uk/sitemap.xml',
sitemap: absoluteUrl('/sitemap.xml'),
};
}
+3 -2
View File
@@ -15,6 +15,7 @@ import {
} from '@/lib/schoolSections';
import { parseSchoolSlug, schoolUrl } from '@/lib/utils';
import type { NationalAverages } from '@/lib/types';
import { absoluteUrl } from '@/lib/site';
import type { Metadata } from 'next';
/**
@@ -97,7 +98,7 @@ export async function generateMetadata({ params }: SchoolPageProps): Promise<Met
title,
description,
type: 'website',
url: `https://schoolcompare.co.uk${canonicalPath}`,
url: absoluteUrl(canonicalPath),
siteName: 'schoolcompare',
},
twitter: {
@@ -106,7 +107,7 @@ export async function generateMetadata({ params }: SchoolPageProps): Promise<Met
description,
},
alternates: {
canonical: `https://schoolcompare.co.uk${canonicalPath}`,
canonical: absoluteUrl(canonicalPath),
},
};
} catch {
@@ -0,0 +1,63 @@
/**
* Phase variants of a place page.
*
* Phase is part of the query — "primary schools in beccles", "secondary
* schools in brentwood" — not a filter applied afterwards, so each gets its
* own indexable path. A place with no schools of the phase has no page: the
* per-phase threshold, not an error.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ place: string; phase: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
const PHASES = ['primary', 'secondary'] as const;
type Phase = (typeof PHASES)[number];
const isPhase = (v: string): v is Phase => (PHASES as readonly string[]).includes(v);
async function resolve(slug: string, phase: Phase) {
return (await fetchPlace('town', slug, phase))
?? (await fetchPlace('locality', slug, phase));
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { place: slug, phase } = await params;
if (!isPhase(phase)) return { title: 'Place Not Found' };
const detail = await resolve(slug, phase);
if (!detail || detail.schools.length === 0) return { title: 'Place Not Found' };
const word = phase === 'secondary' ? 'Secondary' : 'Primary';
const { name } = detail.place;
return {
// Not "Ranked": the table is alphabetical, so the word would be a claim
// the page does not keep.
title: { absolute: `${word} Schools in ${name} | schoolcompare` },
description:
`Every ${phase} school in ${name}, with results, Ofsted grades and the local `
+ `average against England.`,
alternates: { canonical: absoluteUrl(`/schools/${slug}/${phase}`) },
};
}
export default async function PlacePhasePage({ params }: Props) {
const { place: slug, phase } = await params;
if (!isPhase(phase)) notFound();
const detail = await resolve(slug, phase);
if (!detail || detail.schools.length === 0) notFound();
const national = await fetchNationalAverages().catch(() => null);
const englandAverage = phase === 'secondary'
? national?.secondary?.attainment_8_score ?? null
: national?.primary?.rwm_expected_pct ?? null;
return <PlaceView detail={detail} phase={phase}
englandAverage={englandAverage} neighbours={[]} />;
}
+99
View File
@@ -0,0 +1,99 @@
/**
* Town and locality pages.
*
* A place below the five-school threshold is not in the registry, so
* fetchPlace returns null and the request 404s rather than rendering a page
* with nothing to say.
*/
import { notFound, redirect } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces, authoritySlug } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ place: string }> }
// ISR: place aggregates change only when the pipeline runs.
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ place: string }>> {
// Off by default: ~2,000 place routes cannot be built in CI on every deploy.
// Matches the PRERENDER_SCHOOLS gate on the school route.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'town' || p.kind === 'locality')
.map((p) => ({ place: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
async function resolve(slug: string) {
return (await fetchPlace('town', slug)) ?? (await fetchPlace('locality', slug));
}
/** Other towns in the same authority — the cheapest honest definition of
* "nearby", and enough to stop each place page being a dead end. */
async function neighboursOf(detail: { place: { slug: string; parent_authority: string | null } }) {
if (!detail.place.parent_authority) return [];
const all = await fetchPlaces();
return all
.filter((p) => p.kind === 'town' && p.slug !== detail.place.slug)
.slice(0, 12);
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { place: slug } = await params;
const detail = await resolve(slug);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
// absolute: the root layout's template appends '| schoolcompare' to a
// plain string, and this title already carries it. Without this every
// place title read '... | schoolcompare | schoolcompare'.
title: { absolute: `Schools in ${name} — Compare ${count} Schools | schoolcompare` },
description:
`Every school in ${name}, with SATs and GCSE results, Ofsted grades, the local `
+ `average against England, and how close you had to live to get a place.`,
alternates: { canonical: absoluteUrl(`/schools/${slug}`) },
};
}
export default async function PlacePage({ params }: Props) {
const { place: slug } = await params;
const detail = await resolve(slug);
if (!detail) notFound();
// Global constraint: no page without a local average. A place with too few
// schools carrying results has nothing to say that a list does not, so it
// defers to its authority rather than publishing a thin page.
if (detail.averages.rwm_expected_pct == null
&& detail.averages.attainment_8_score == null) {
// The API's own slug, which is null when that authority is itself under
// the threshold and has no page. Re-slugifying the name here would send
// the reader to a 404 instead of telling them the place has no page.
const target = detail.place.authorities?.[0]?.slug
?? (detail.place.parent_authority
? authoritySlug(detail.place.parent_authority)
: null);
if (target) redirect(`/schools/authority/${target}`);
notFound();
}
const national = await fetchNationalAverages().catch(() => null);
// NationalAverages is nested by phase — { primary: {...}, secondary: {...} }
// — not flat. Reading it flat silently yields undefined and the page renders
// with no comparison, which is the one thing that makes it not a list.
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={await neighboursOf(detail)}
/>
);
}
@@ -0,0 +1,65 @@
/**
* Phase variants of an authority page.
*
* The spec called for these; the plan built the bare authority route and
* dropped them. Nothing caught it, because the sitemap is written from the
* place registry — which was right about them all along — while the routes
* were written by hand. 302 authority phase URLs were submitted to Google and
* every one 404'd, and every authority page linked to a phase page in the
* *town* namespace, which is a different set of schools entirely.
*
* "Primary schools in Kent" is the query these serve, and it is a real one:
* admissions are authority-run, so the authority is the unit a parent thinks
* in when they have not settled on a town.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ la: string; phase: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
const PHASES = ['primary', 'secondary'] as const;
type Phase = (typeof PHASES)[number];
const isPhase = (v: string): v is Phase => (PHASES as readonly string[]).includes(v);
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { la, phase } = await params;
if (!isPhase(phase)) return { title: 'Place Not Found' };
const detail = await fetchPlace('authority', la, phase);
if (!detail || detail.schools.length === 0) return { title: 'Place Not Found' };
const word = phase === 'secondary' ? 'Secondary' : 'Primary';
const { name } = detail.place;
return {
// "Local Authority" stays in the title for the same reason it is on the
// bare authority page: 67 town names collide with an authority name, and
// a reader landing on both needs to know which set each covers.
title: { absolute: `${word} Schools in ${name} — Local Authority | schoolcompare` },
description:
`Every ${phase} school in the ${name} local authority, with results, Ofsted `
+ `grades and the authority average against England.`,
alternates: { canonical: absoluteUrl(`/schools/authority/${la}/${phase}`) },
};
}
export default async function AuthorityPhasePage({ params }: Props) {
const { la, phase } = await params;
if (!isPhase(phase)) notFound();
const detail = await fetchPlace('authority', la, phase);
if (!detail || detail.schools.length === 0) notFound();
const national = await fetchNationalAverages().catch(() => null);
const englandAverage = phase === 'secondary'
? national?.secondary?.attainment_8_score ?? null
: national?.primary?.rwm_expected_pct ?? null;
return <PlaceView detail={detail} phase={phase}
englandAverage={englandAverage} neighbours={[]} />;
}
@@ -0,0 +1,67 @@
/**
* Local authority pages.
*
* A separate namespace from /schools/[place] because 67 town names collide
* with an authority name and neither set contains the other — Bedford the
* town holds 104 schools, Bedford the authority 86, because postal towns
* cross authority boundaries. The title says "Local Authority" so a reader
* landing on both knows which set each covers.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ la: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ la: string }>> {
// Gated like every other prerender in this app. There are only ~154
// authorities, but "few enough to always build" still means the API must be
// reachable at build time, and in CI it is not — the build fails with
// ECONNREFUSED rather than degrading. The catch is the same fallback the
// school route uses.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'authority')
.map((p) => ({ la: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { la } = await params;
const detail = await fetchPlace('authority', la);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
title: { absolute: `Schools in ${name} — Local Authority | schoolcompare` },
description:
`All ${count} schools in the ${name} local authority, with SATs and GCSE results, `
+ `Ofsted grades and the authority average against England.`,
alternates: { canonical: absoluteUrl(`/schools/authority/${la}`) },
};
}
export default async function AuthorityPage({ params }: Props) {
const { la } = await params;
const detail = await fetchPlace('authority', la);
if (!detail) notFound();
const national = await fetchNationalAverages().catch(() => null);
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={[]}
/>
);
}
@@ -0,0 +1,62 @@
/**
* Postcode district pages.
*
* No phase variants: nobody searches "primary schools in SW11", so the
* variants would be pages without demand. These exist to catch
* "schools near <postcode>" and to give London districts a geographic page
* where the GIAS town field cannot.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ outcode: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ outcode: string }>> {
// 1,760 of these; same CI budget argument as the town routes.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'outcode')
.map((p) => ({ outcode: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { outcode } = await params;
const detail = await fetchPlace('outcode', outcode);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
title: { absolute: `Schools near ${name} | schoolcompare` },
description:
`${count} schools in the ${name} postcode district, with results, Ofsted grades `
+ `and how close you had to live to get a place.`,
alternates: { canonical: absoluteUrl(`/schools/near/${outcode}`) },
};
}
export default async function OutcodePage({ params }: Props) {
const { outcode } = await params;
const detail = await fetchPlace('outcode', outcode);
if (!detail) notFound();
const national = await fetchNationalAverages().catch(() => null);
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={[]}
/>
);
}
+2 -26
View File
@@ -1,32 +1,8 @@
/**
* Runtime proxy for /sitemap.xml → the FastAPI backend's generated sitemap.
*
* Like the /api/* proxy, this reads FASTAPI_URL at request time rather than
* baking the backend host into the build, so one image works in every
* environment. robots.ts points crawlers here.
*/
import { NextResponse } from 'next/server';
import { proxySitemap } from '@/lib/sitemapProxy';
export const dynamic = 'force-dynamic';
export const runtime = 'nodejs';
function backendOrigin(): string {
const base = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL || 'http://localhost:8000/api';
return base.replace(/\/api$/, '');
}
export async function GET() {
let upstream: Response;
try {
upstream = await fetch(`${backendOrigin()}/sitemap.xml`, { cache: 'no-store' });
} catch {
return new NextResponse('Sitemap temporarily unavailable', { status: 502 });
}
const body = await upstream.text();
return new NextResponse(body, {
status: upstream.status,
headers: { 'content-type': upstream.headers.get('content-type') || 'application/xml' },
});
return proxySitemap('/sitemap.xml');
}
@@ -0,0 +1,24 @@
import { NextResponse } from 'next/server';
import { proxySitemap } from '@/lib/sitemapProxy';
export const dynamic = 'force-dynamic';
export const runtime = 'nodejs';
/**
* Children are /sitemaps/static.xml and /sitemaps/schools-{n}.xml. The name is
* validated here rather than passed through, so this route cannot be used to
* reach arbitrary backend paths.
*/
const CHILD = /^(static|schools-\d+|places-\d+|outcodes-\d+)\.xml$/;
export async function GET(
_request: Request,
{ params }: { params: Promise<{ parts: string[] }> },
) {
const { parts } = await params;
const name = parts.join('/');
if (!CHILD.test(name)) {
return new NextResponse('Not found', { status: 404 });
}
return proxySitemap(`/sitemaps/${name}`);
}
@@ -0,0 +1,210 @@
/* Tokens only — see globals.css. Follows RankingsView's conventions, and in
particular its link treatment: table links take --text-primary with no
underline and a brand-coloured hover, not the browser default. The first
cut used bare <Link> with no class at all, which rendered as default blue
underlined links and read as unstyled beside the rest of the site. */
.container {
width: 100%;
min-width: 0;
}
.header {
margin-bottom: 1.5rem;
}
.header h1 {
font-size: 2.25rem;
font-weight: 700;
color: var(--text-primary);
margin-bottom: 0.5rem;
font-family: var(--font-display);
text-wrap: balance;
}
.summary {
font-size: 1rem;
color: var(--text-secondary);
margin: 0;
line-height: 1.6;
}
/* Links in running copy: brand colour, underline on hover only. */
.inlineLink {
color: var(--brand);
text-decoration: none;
transition: color 0.2s ease;
}
.inlineLink:hover {
color: var(--brand-strong);
text-decoration: underline;
}
/* Phase variants are separate indexable pages, so the bare place page has to
link them — a sitemap entry alone leaves them with no internal path in. */
.phaseLinks {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 0.75rem;
margin: 0 0 1.25rem;
}
.phaseLink {
display: inline-block;
padding: 0.4rem 0.875rem;
border: 1px solid var(--border-strong);
border-radius: 999px;
font-size: 0.875rem;
font-weight: 500;
color: var(--text-primary);
text-decoration: none;
transition: border-color 0.2s ease, color 0.2s ease;
}
.phaseLink:hover {
border-color: var(--brand);
color: var(--brand-strong);
}
/* The one number a list cannot give you, so it gets its own band. */
.compare {
background: var(--bg-secondary);
border: 1px solid var(--border);
border-radius: 8px;
padding: 0.875rem 1.125rem;
margin: 0 0 1.5rem;
color: var(--text-primary);
font-size: 1rem;
}
.ofsted {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1.25rem;
list-style: none;
padding: 0;
margin: 0 0 1.5rem;
font-size: 0.9375rem;
color: var(--text-secondary);
}
.group {
margin-bottom: 2rem;
}
.groupHeading {
display: flex;
align-items: baseline;
gap: 0.625rem;
font-size: 1.25rem;
font-weight: 600;
color: var(--text-primary);
font-family: var(--font-display);
margin: 0 0 0.75rem;
}
.groupCount {
font-size: 0.8125rem;
font-weight: 500;
color: var(--text-secondary);
background: var(--bg-secondary);
border-radius: 999px;
padding: 0.125rem 0.5rem;
}
/* Wide content scrolls in its own container so the page body never does. */
.tableWrap {
overflow-x: auto;
border: 1px solid var(--border);
border-radius: 8px;
background: var(--bg-card);
}
.table {
width: 100%;
border-collapse: collapse;
font-size: 0.9375rem;
}
.table th,
.table td {
padding: 0.75rem 1rem;
text-align: left;
border-bottom: 1px solid var(--border);
}
.table th {
background: var(--bg-secondary);
color: var(--text-secondary);
font-weight: 600;
font-size: 0.8125rem;
}
.table tbody tr:last-child td {
border-bottom: none;
}
/*
* Header and value share one class and one rule, so they cannot drift apart.
*
* The first cut aligned them with two different selectors: `.table th:last-child`
* at (0,2,1) beat the element rule and went right, while `.num` at (0,1,0) lost
* to `.table td` at (0,1,1) and stayed left. The heading and its numbers sat on
* opposite edges of the column.
*
* width:1% with nowrap makes the measure column hug its content so the school
* name takes the remaining width — without it the two columns split evenly and
* the gap between heading and value reads as misalignment on a wide screen.
*/
.table th.num,
.table td.num {
text-align: right;
font-variant-numeric: tabular-nums;
width: 1%;
white-space: nowrap;
}
/* The measure is spelled out; the tooltip carries the definition. */
.metricHead {
text-decoration: none;
cursor: help;
border-bottom: 1px dotted var(--border-strong);
}
/* Table links: site convention is body colour, brand on hover. */
.schoolLink {
color: var(--text-primary);
text-decoration: none;
transition: color 0.2s ease;
}
.schoolLink:hover {
color: var(--brand-strong);
}
/* "Not published" is a fact about the school, not an error. */
.noData {
color: var(--text-muted);
font-size: 0.8125rem;
}
.neighbours {
margin-top: 2rem;
}
.neighbours h2 {
font-size: 1.125rem;
font-weight: 600;
color: var(--text-primary);
margin: 0 0 0.75rem;
font-family: var(--font-display);
}
.neighbours ul {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1rem;
list-style: none;
padding: 0;
margin: 0;
}
+258
View File
@@ -0,0 +1,258 @@
/**
* One place page, shared by all four families.
*
* They differ in what fills the registry, not in what the page shows, so a
* second component would be a second place to forget the same change.
*
* The local-versus-England comparison is the reason this page is not a list:
* it is the one number a parent cannot get by reading the schools one by one,
* and it is what keeps the page from reading as a name dropped into a
* template.
*/
import Link from 'next/link';
import type { PlaceDetail, PlaceSummary } from '@/lib/places';
import { placeUrl, authoritySlug } from '@/lib/places';
import type { School } from '@/lib/types';
import { schoolUrl } from '@/lib/utils';
import { absoluteUrl } from '@/lib/site';
import styles from './PlaceView.module.css';
interface Props {
detail: PlaceDetail;
phase?: 'primary' | 'secondary';
englandAverage: number | null;
/** Nearby places, so the page links onward instead of dead-ending. */
neighbours: PlaceSummary[];
}
// Ofsted grades in the order they are reported.
const OFSTED_LABELS: Array<[number, string]> = [
[1, 'Outstanding'], [2, 'Good'],
[3, 'Requires improvement'], [4, 'Inadequate'],
];
/**
* Column headings, taken from the site's own metric dictionary rather than
* invented here — see METRIC_DEFINITIONS in backend/schemas.py, surfaced at
* /api/metrics. The first cut said "RWM expected", which is jargon that
* appears nowhere else on the site.
*/
const METRICS = {
primary: {
key: 'rwm_expected_pct' as const,
heading: 'Reading, writing & maths',
hint: '% meeting the expected standard in reading, writing and maths',
unit: '%',
},
secondary: {
key: 'attainment_8_score' as const,
heading: 'Attainment 8',
hint: "Average grade across a pupil's best 8 GCSEs, including English and maths",
unit: '',
},
};
type PhaseKey = keyof typeof METRICS;
/** All-through schools sit in both phases, matching the search filters. */
function isPhase(school: School, phase: PhaseKey): boolean {
const p = (school.phase ?? '').toLowerCase();
if (p === 'all-through') return true;
return phase === 'secondary'
? p.includes('secondary') || p === '16 plus'
: p.includes('primary') || p.includes('middle');
}
function SchoolTable({ schools, phase }: { schools: School[]; phase: PhaseKey }) {
const metric = METRICS[phase];
return (
<div className={styles.tableWrap}>
<table className={styles.table}>
<thead>
<tr>
<th scope="col">School</th>
{/* Same class as the value cell below: one rule aligns both, so
they cannot drift apart. */}
<th scope="col" className={styles.num}>
<abbr className={styles.metricHead} title={metric.hint}>
{metric.heading}
</abbr>
</th>
</tr>
</thead>
<tbody>
{schools.map((s) => {
const value = s[metric.key];
return (
<tr key={s.urn}>
<td>
<Link href={schoolUrl(s.urn, s.school_name)} className={styles.schoolLink}>
{s.school_name}
</Link>
</td>
<td className={styles.num}>
{value == null
? <span className={styles.noData}>Not published</span>
: `${Math.round(Number(value))}${metric.unit}`}
</td>
</tr>
);
})}
</tbody>
</table>
</div>
);
}
export function PlaceView({ detail, phase, englandAverage, neighbours }: Props) {
const { place, schools, averages } = detail;
// Fall back to the single parent when the API predates the authorities
// field, so a stale cache never blanks the line entirely.
const authorities = place.authorities?.length
? place.authorities
: place.parent_authority
? [{ name: place.parent_authority, slug: authoritySlug(place.parent_authority), count: 0 }]
: [];
const local = averages[METRICS[phase ?? 'primary'].key];
const phaseWord = phase === 'secondary' ? 'Secondary schools'
: phase === 'primary' ? 'Primary schools' : 'Schools';
const graded = OFSTED_LABELS
.map(([grade, label]) => [label, schools.filter((s) => s.ofsted_grade === grade).length] as const)
.filter(([, n]) => n > 0);
/*
* An unphased page holds both primaries and secondaries, and they are
* scored on different measures — a percentage and a 0-90 score. Showing one
* column for both left 30% of rows blank on /schools/brentwood and put two
* incomparable scales in one column when it did not.
*
* So the phases get a table each. A blank cell inside one now means the
* school genuinely has no published result, which is worth saying.
*/
const groups: Array<[PhaseKey, School[]]> = phase
? [[phase, schools]]
: (['primary', 'secondary'] as PhaseKey[])
.map((p) => [p, schools.filter((s) => isPhase(s, p))] as [PhaseKey, School[]])
.filter(([, list]) => list.length > 0);
const jsonLd = {
'@context': 'https://schema.org',
'@graph': [
{
'@type': 'ItemList',
name: `${phaseWord} in ${place.name}`,
numberOfItems: schools.length,
// Alphabetical, and said so. Without this an ItemList carrying
// `position` reads as a ranking, which would be a claim the page
// stopped making when the table became A-Z.
itemListOrder: 'https://schema.org/ItemListOrderAscending',
itemListElement: schools.slice(0, 20).map((s, i) => ({
'@type': 'ListItem',
position: i + 1,
url: absoluteUrl(schoolUrl(s.urn, s.school_name)),
name: s.school_name,
})),
},
{
'@type': 'BreadcrumbList',
itemListElement: [
{ '@type': 'ListItem', position: 1, name: 'Schools', item: absoluteUrl('/') },
{ '@type': 'ListItem', position: 2, name: place.name },
],
},
],
};
return (
<div className={styles.container}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<header className={styles.header}>
<h1>{phaseWord} in {place.name}</h1>
<p className={styles.summary}>
{place.count} schools
{authorities.length > 0 && (
<>
{' · '}
{/* Every authority, not just the largest. A quarter of outcodes
and a third of towns cross a boundary: SW19 is mostly Merton
but partly Wandsworth, and naming one asserts otherwise. */}
{authorities.map((a, i) => (
<span key={a.name}>
{i > 0 && (i === authorities.length - 1 ? ' and ' : ', ')}
{/* No slug means no page: City of London and the Isles of
Scilly hold too few schools for one. Saying where the
place is stays right; linking there would 404. */}
{a.slug
? (
<Link href={`/schools/authority/${a.slug}`} className={styles.inlineLink}>
{a.name}
</Link>
)
: a.name}
</span>
))}
</>
)}
</p>
</header>
{!phase && (place.phases ?? []).length > 0 && (
<nav className={styles.phaseLinks} aria-label="By phase">
{(place.phases ?? []).map((ph) => (
/* placeUrl, not a template: the bare `/schools/[slug]/[phase]`
shape belongs to towns alone, and using it everywhere sent
every authority page into the town namespace. */
<Link key={ph} href={placeUrl(place.kind, place.slug, ph)}
className={styles.phaseLink}>
{ph === 'secondary' ? 'Secondary schools' : 'Primary schools'} in {place.name}
</Link>
))}
</nav>
)}
{local != null && englandAverage != null && (
<p className={styles.compare} data-testid="local-vs-england">
{place.name} averages <strong>{Math.round(local)}</strong> against{' '}
<strong>{Math.round(englandAverage)}</strong> across England.
</p>
)}
{graded.length > 0 && (
<ul className={styles.ofsted} data-testid="ofsted-distribution">
{graded.map(([label, n]) => (
<li key={label}>{label}: <strong>{n}</strong></li>
))}
</ul>
)}
{groups.map(([p, list]) => (
<section key={p} className={styles.group}>
{groups.length > 1 && (
<h2 className={styles.groupHeading}>
{p === 'secondary' ? 'Secondary schools' : 'Primary schools'}
<span className={styles.groupCount}>{list.length}</span>
</h2>
)}
<SchoolTable schools={list} phase={p} />
</section>
))}
{neighbours.length > 0 && (
<nav className={styles.neighbours} aria-label="Nearby places">
<h2>Nearby</h2>
<ul>
{neighbours.map((n) => (
<li key={n.kind + n.slug}>
<Link href={placeUrl(n.kind, n.slug)} className={styles.inlineLink}>{n.name}</Link>
</li>
))}
</ul>
</nav>
)}
</div>
);
}
+8
View File
@@ -12,6 +12,12 @@ jest.mock('next/navigation', () => ({
useSearchParams: () => new URLSearchParams(),
}));
// Everything below this line is browser furniture, and this file runs for
// every suite — including the ones that declare `@jest-environment node` to
// test route handlers, where NextRequest needs Fetch API globals jsdom does
// not provide. There is no `window` there, so guard rather than assume one.
if (typeof window !== 'undefined') {
// Mock window.matchMedia
Object.defineProperty(window, 'matchMedia', {
writable: true,
@@ -52,3 +58,5 @@ const localStorageMock = {
clear: jest.fn(),
};
global.localStorage = localStorageMock;
} // end: browser-only globals
+40
View File
@@ -0,0 +1,40 @@
/**
* Reading feature flags.
*
* Server-side only. No flag value reaches the browser bundle, and there is no
* Unleash dependency in package.json — the SDK lives in FastAPI, which already
* owns every other piece of data this app renders.
*
* Flags are declared in backend/flags.py. A purely front-end flag still has to
* be declared there; it is a flat data edit, and the return is that one list
* answers "what flags exist" for the whole system.
*/
export type Flags = Record<string, boolean>;
/*
* Reading flags pins the calling route to this ISR floor: Next uses the LOWEST
* revalidate among a route's fetches to set the whole route's revalidation
* frequency. 300s matches what /school/[slug] already sits at, so a page that
* reads flags is no more dynamic than a school page already is.
*
* It is also what makes a flip propagate without a webhook: five minutes on
* school pages, an hour on place pages, against flags that flip monthly.
*/
export const FLAGS_REVALIDATE = 300;
const API = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL
|| 'http://localhost:8000/api';
/** Every flag and its value. Never throws: an unreadable flag is a dark one. */
export async function getFlags(): Promise<Flags> {
try {
const res = await fetch(`${API}/flags`, {
next: { revalidate: FLAGS_REVALIDATE },
});
if (!res.ok) return {};
return await res.json();
} catch {
return {};
}
}
+84
View File
@@ -0,0 +1,84 @@
/**
* Client for the places API.
*
* Two namespaces, matching the backend: towns and localities share
* /schools/[place]; authorities take /schools/authority/[la]. 67 town names
* collide with an authority name and neither set contains the other, so one
* namespace would publish near-duplicate pages.
*/
import type { School } from '@/lib/types';
export interface PlaceSummary {
kind: string;
slug: string;
name: string;
count: number;
/** Phases that clear the threshold on their own, so the page links
* variants that exist rather than 404s. Absent on the registry listing. */
phases?: string[];
}
export interface PlaceAuthority {
name: string;
/** null when that authority has no page of its own — two English
* authorities hold fewer schools than the threshold. */
slug: string | null;
count: number;
}
export interface PlaceDetail {
place: PlaceSummary & {
parent_authority: string | null;
/** Every authority the place meaningfully sits in, largest first. SW19 is
* mostly Merton but partly Wandsworth. */
authorities?: PlaceAuthority[];
};
schools: School[];
averages: {
rwm_expected_pct: number | null;
attainment_8_score: number | null;
};
}
export function placeUrl(kind: string, slug: string, phase?: string): string {
const base =
kind === 'authority' ? `/schools/authority/${slug}`
: kind === 'outcode' ? `/schools/near/${slug}`
: `/schools/${slug}`;
return phase ? `${base}/${phase}` : base;
}
/**
* An authority name as it appears in a URL.
*
* Only a fallback: the API sends the slug it built, and that is what should
* be used. This mirrors `_slugify` in backend/app.py, collapsed runs and
* trimmed hyphens included, so the two cannot disagree about a name like
* "Bristol, City of".
*/
export function authoritySlug(name: string): string {
return name.toLowerCase().trim()
.replace(/[^\w\s-]/g, '')
.replace(/\s+/g, '-')
.replace(/-+/g, '-')
.replace(/^-|-$/g, '');
}
const API = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL
|| 'http://localhost:8000/api';
export async function fetchPlaces(): Promise<PlaceSummary[]> {
const res = await fetch(`${API}/places`, { next: { revalidate: 604800 } });
if (!res.ok) return [];
return (await res.json()).places ?? [];
}
export async function fetchPlace(
kind: string, slug: string, phase?: string,
): Promise<PlaceDetail | null> {
const q = phase ? `?phase=${encodeURIComponent(phase)}` : '';
const res = await fetch(`${API}/places/${kind}/${slug}${q}`,
{ next: { revalidate: 604800 } });
if (!res.ok) return null;
return res.json();
}
+19
View File
@@ -0,0 +1,19 @@
/**
* The one place the site's absolute origin is written down.
*
* The apex domain 301s to www at Cloudflare, so www is the host that actually
* serves a 200. Canonicals, og:url, sitemap <loc> entries and the robots.txt
* Sitemap: line must all agree with it — a canonical pointing at a redirect
* makes Google resolve the hop before it can consolidate the signal.
*
* backend/app.py holds the same value as BASE_URL for the sitemap. The two are
* asserted against each other by the e2e journeys rather than shared at build
* time, because the backend and frontend ship as separate images.
*/
export const SITE_URL = 'https://www.schoolcompare.co.uk';
/** Absolute URL for a site-relative path. Tolerates a missing leading slash. */
export function absoluteUrl(path: string): string {
const rooted = path.startsWith('/') ? path : `/${path}`;
return `${SITE_URL}${rooted}`;
}
+29
View File
@@ -0,0 +1,29 @@
/**
* Runtime proxy for the sitemap family → the FastAPI backend.
*
* Like the /api/* proxy, this reads FASTAPI_URL at request time rather than
* baking the backend host into the build, so one image works in every
* environment. robots.ts points crawlers at /sitemap.xml, which is the index;
* the index names children under /sitemaps/, which land on the same proxy.
*/
import { NextResponse } from 'next/server';
function backendOrigin(): string {
const base = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL || 'http://localhost:8000/api';
return base.replace(/\/api$/, '');
}
export async function proxySitemap(path: string): Promise<NextResponse> {
let upstream: Response;
try {
upstream = await fetch(`${backendOrigin()}${path}`, { cache: 'no-store' });
} catch {
return new NextResponse('Sitemap temporarily unavailable', { status: 502 });
}
const body = await upstream.text();
return new NextResponse(body, {
status: upstream.status,
headers: { 'content-type': upstream.headers.get('content-type') || 'application/xml' },
});
}
+6 -1
View File
@@ -361,7 +361,12 @@ export interface SchoolDetailsResponse {
* held back as a paid feature and are not part of this public payload — see
* data_loader._admission_distance.
*/
admission_distance: SchoolAdmissionDistance | null;
/**
* Absent — not null — when the admission_distance flag is off. Null means
* "this school has no published cut-off"; absent means "cut-offs are not
* being published at all". They are different claims and the type says so.
*/
admission_distance?: SchoolAdmissionDistance | null;
deprivation: SchoolDeprivation | null;
finance: SchoolFinance | null;
}
@@ -0,0 +1,16 @@
locality_slug,locality_name,outcodes,region
battersea,Battersea,SW11,London
canary-wharf,Canary Wharf,E14,London
clapham,Clapham,SW4,London
shoreditch,Shoreditch,EC2A|E1,London
peckham,Peckham,SE15,London
brixton,Brixton,SW2|SW9,London
camden-town,Camden Town,NW1,London
wimbledon,Wimbledon,SW19,London
putney,Putney,SW15,London
fulham,Fulham,SW6,London
chiswick,Chiswick,W4,London
stratford,Stratford,E15,London
walthamstow,Walthamstow,E17,London
tooting,Tooting,SW17,London
dulwich,Dulwich,SE21|SE22,London
1 locality_slug locality_name outcodes region
2 battersea Battersea SW11 London
3 canary-wharf Canary Wharf E14 London
4 clapham Clapham SW4 London
5 shoreditch Shoreditch EC2A|E1 London
6 peckham Peckham SE15 London
7 brixton Brixton SW2|SW9 London
8 camden-town Camden Town NW1 London
9 wimbledon Wimbledon SW19 London
10 putney Putney SW15 London
11 fulham Fulham SW6 London
12 chiswick Chiswick W4 London
13 stratford Stratford E15 London
14 walthamstow Walthamstow E17 London
15 tooting Tooting SW17 London
16 dulwich Dulwich SE21|SE22 London
+1 -1
View File
@@ -12,4 +12,4 @@ slowapi==0.1.9
secure==0.3.0
typesense==0.21.0
numpy==1.26.4
UnleashClient==6.0.1