Compare commits

...
Author SHA1 Message Date
TudorandClaude Opus 5 9cc87c41bb fix(places): a phase page needs results, not merely publishable schools
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m4s
Asked where schools with no results should sit in an alphabetical list, and
found that some pages were almost entirely made of them.

The per-phase threshold counted schools that were publishable — a result OR
an Ofsted grade — while a phase page exists for its results column.
/schools/kent/primary published with none of its five rows carrying a result;
Minehead had one of seven, Buntingford one of five. Forty-four phase pages
were majority-blank.

It is the same rule as "no page without a local average", which was written
into the spec as a thin-page control and never extended per phase.

The threshold now counts schools with a result for that phase. It gates
whether the page exists; it does not filter rows — a page that publishes still
lists every school of the phase, because someone looking up a school by name
has to find it whether or not it published results.

126 of 1,012 variant pages stop publishing: 62 primary, 64 secondary. Every
one of them was a table with too little in it to be worth a page.

The ordering itself is unchanged: pure A-Z, blanks interleaved. A school sits
where its name says it does, and at roughly a tenth of rows that reads fine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:14:31 +01:00
TudorandClaude Opus 5 8967966eef feat(places): list schools alphabetically on place pages
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 11s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 1m21s
Someone on a place page is usually looking for a school they can name, so the
order should serve scanning for it rather than ranking. /api/rankings keeps
its league-table ordering; this is a place-page decision, not a site-wide one.
Sorted case-insensitively, or a capitalised name would sort ahead of every
lowercase one.

The change made five pieces of copy untrue, so they go with it. The phase
variant titled itself "— Ranked", and all four route families described
themselves as "ranked by SATs and GCSE results". A page that opens by claiming
an order it does not keep is worse than one that claims nothing.

The ItemList markup carried `position` with no declared order, which reads as
a ranking. It now declares ItemListOrderAscending, so the structured data says
what the table does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-22 00:09:58 +01:00
tudor 4a9a5c734b Merge pull request 'fix(e2e): three assertions that were wrong about correct behaviour' (#122) from fix/e2e-canonical-and-robots into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Successful in 1m26s
Reviewed-on: #122
2026-08-21 23:07:57 +00:00
TudorandClaude Opus 5 4e82e6c916 fix(e2e): three assertions that were wrong about correct behaviour
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 41s
The staging gate was red on three journeys. All three were faults in the
tests; the site was behaving correctly in each case.

Next normalises canonical URLs against trailingSlash:false, so the homepage
ships "https://www.schoolcompare.co.uk" with no slash while every other route
keeps its path. Both address the same document. The test hardcoded the slash
and so failed only on the root — /rankings and /admissions passed throughout,
which is what made it look like a homepage bug rather than a test bug.
Compared with trailing slashes stripped from both sides.

The robots.txt assertion matched "Disallow: /" anywhere in the file and
tripped over the AI-crawler groups Cloudflare injects — ClaudeBot, GPTBot,
Amazonbot and six others all carry a blanket disallow, deliberately, and none
of them is Googlebot. It now parses the file into user-agent groups and checks
only the "*" group, which is also the thing the test was always trying to say:
Google may crawl the page, so it can see the noindex header.

Both were the same mistake as the doubled brand: asserting a naive string
rather than the semantics, and asserting against what the code assembles
rather than what the page renders.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 23:51:34 +01:00
tudor d4340a8fdd Merge pull request 'feat(places): name every authority a place sits in' (#121) from feat/place-multiple-authorities into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #121
2026-08-21 21:56:50 +00:00
TudorandClaude Opus 5 bb2f7a5841 fix(places): address review, and merge places GIAS spells more than one way
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m34s
Two findings from review on #121, plus a third the review prompted.

The cap at three authorities silently dropped the fourth in exactly the case
where the information matters most — a genuinely fragmented place — and
contradicted the stated goal of naming every authority a place sits in. It is
gone. The share rule was always the real limit and already bounds the list at
ten. Measured against the live corpus, one town would have been truncated
today: LONDON, split evenly between Hackney, Lambeth, Westminster and
Lewisham.

parent_authority used mode() while authorities used value_counts(), and on an
exact tie pandas does not guarantee the two pick the same name, so the 301
could have pointed somewhere other than the authority named first on the page.
The parent is now derived from authorities[0]: one computation, one answer.
It also inherits the sentinel filter, so a place can no longer redirect to
/schools/authority/does-not-apply.

Chasing the truncation case surfaced a worse bug. Places were grouped by raw
town value, but the registry is keyed by slug, and GIAS spells the same place
several ways. Five town slugs come from more than one spelling: "London"
(1,819 schools) and "LONDON" (12) both slugify to `london`, so the later group
simply overwrote the earlier one — /schools/london could have shown twelve
schools, silently, depending on row order. Weston-super-Mare was split 14/19
across two spellings and Newcastle-under-Lyme across three. Grouping is now by
slug, and the display name is the most common spelling.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:42:18 +01:00
TudorandClaude Opus 5 1cb5314c53 feat(places): name every authority a place sits in
SW19 is mostly Merton but partly Wandsworth, and the page said only Merton.
The cause was one field doing two jobs: _parent_authority takes the modal
authority, which is right for a 301 target and wrong as a statement about
where a place is.

This is not a corner case. A quarter of viable outcodes (425 of 1,760) and a
third of viable towns (263 of 783) cross an authority boundary — Bedford the
town spans Bedford and Central Bedfordshire.

Place now carries `authorities`, every authority holding at least a tenth of
the schools and at least two of them, largest first. parent_authority stays
single and unchanged, because a redirect still needs one target.

The share threshold exists because GIAS carries postcode errors: EN6 lists two
Shropshire schools among fourteen in Hertfordshire, and a bare "any authority
present" rule would print those as though they were real. A place too small or
too fragmented to clear the threshold still names its largest, so the page
never goes silent about where it is.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:40:58 +01:00
tudor 4cea26b813 Merge pull request 'fix(places): align the measure column's heading with its values' (#120) from fix/place-table-alignment into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #120
2026-08-21 21:21:03 +00:00
TudorandClaude Opus 5 dbb74d9b60 fix(places): align the measure column's heading with its values
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 14s
The heading sat on the right edge of the column and every value on the left.
A specificity collision, not a layout problem: the two were aligned by
different selectors and only one of them won.

  .table td            (0,1,1)  text-align: left    <- won for the value
  .num                 (0,1,0)  text-align: right   <- lost
  .table th:last-child (0,2,1)  text-align: right   <- won for the heading

The heading and the value cell now share one class and one rule, so they
cannot drift apart again whatever else changes around them.

The column also stretched to half the table. It now hugs its content with
width:1% and nowrap, so the school name takes the remaining width — which is
what made the gap read as misalignment on a wide screen, and what crowded the
name column on a narrow one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 22:15:02 +01:00
tudor 9545aec7f4 Merge pull request 'fix(places): phase-grouped tables, plain-English measures, styled links' (#119) from fix/place-presentation-to-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 56s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #119
2026-08-21 20:57:13 +00:00
Tudor 3365ebcb3a fix(places): phase-grouped tables, plain-English measures, styled links
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 10s
Three presentation faults on the place pages, all found by looking at a
rendered page rather than at a test.

An unphased place page showed one primary-only measure for a list holding both
phases: 8 of 27 rows on /schools/brentwood were blank, because secondaries
have no reading-writing-maths score. Picking the other measure would only have
inverted which rows were empty, and putting both in one column would have
mixed a percentage with a 0-90 score. Each phase now gets its own table, so a
blank cell means the school genuinely has no published result — which is worth
saying, and now says "Not published" rather than a bare dash.

"RWM expected" was invented here. The site already names the measure in
METRIC_DEFINITIONS, surfaced at /api/metrics: "Reading, Writing & Maths
Combined %". The heading now reads "Reading, writing & maths" with the full
definition in the tooltip.

Links carried no class at all, so they rendered as default blue underlined
browser links beside a site that styles table links as body colour with a
brand hover. They now follow RankingsView's convention, and running-copy links
take the brand colour.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 21:56:06 +01:00
tudor 6d79bd3331 Merge pull request 'fix(places): stop the place titles doubling the brand' (#117) from fix/place-title-brand-doubling into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m30s
Reviewed-on: #117
2026-08-21 20:53:14 +00:00
TudorandClaude Opus 5 24e114dee7 fix(places): stop the place titles doubling the brand
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 27s
Every place page shipped as 'Schools in Brentwood - Compare 27 Schools |
schoolcompare | schoolcompare'. The root layout's title template appends
'| schoolcompare' to any plain-string title, and all four place routes already
carried the brand. W8 opted the other routes out with an absolute title; the
place routes were written afterwards and did not inherit the lesson.

~2,600 titles affected, and the repetition pushed them past Google's
truncation point, so the doubled brand displaced real words in the result.

An e2e journey now asserts no title repeats the brand, across the static
routes and a place page, so this cannot come back on a route added later.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 21:43:32 +01:00
tudor 6c5db0c266 Merge pull request 'fix(places): submit and link the phase variants' (#116) from fix/place-phase-variants into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m29s
Reviewed-on: #116
2026-08-21 19:46:55 +00:00
TudorandClaude Opus 5 6f749ed21f fix(places): submit and link the phase variants
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 2m39s
/schools/[place]/[phase] shipped as routes but reached nothing. The sitemap
emitted one URL per registry entry and the registry had no phase dimension, so
~950 pages were absent from every sitemap — and PlaceView did not link them
either, leaving them reachable by nothing at all.

That is the query shape the baseline actually showed: 'primary schools in
beccles', 'secondary schools in brentwood'. Publishing the routes without a
path in meant building for the demand and then hiding from it.

Place now carries phase_urns so the per-phase threshold can be applied without
re-querying, the sitemap emits a variant wherever a phase clears the threshold
on its own, and the API exposes the qualifying phases so the place page links
only variants that exist. Outcodes are excluded: nobody searches 'primary
schools in SW11' and those routes do not exist.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 20:43:07 +01:00
tudor d423826840 Merge pull request 'fix(places): a locality collision must not break the sitemap' (#115) from fix/locality-collision-skip into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m13s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m38s
Reviewed-on: #115
2026-08-21 19:26:02 +00:00
TudorandClaude Opus 5 d3c63ccc6d fix(places): a locality collision must not break the sitemap
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 34s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 34s
Sitemap regeneration failed on staging. 'richmond' in the curated locality
list collides with the GIAS town Richmond in North Yorkshire (37 schools), the
registry raised, and the admin endpoint 500d — taking down sitemap generation
for all 25,000 school pages over one bad row of curated data.

The guard now skips the colliding locality and logs an error. Skipping still
achieves what the guard was for — a locality never silently shadows a town —
without letting curated data break the site. That matters beyond this bug:
GIAS town names change with no code change here, so a raise could fire
spontaneously in production later.

Also removes four localities that were London boroughs rather than districts.
Hackney, Islington, Greenwich and Ealing are local authorities with 104, 72,
108 and 115 schools and already have authority pages; a locality defined by
two or three outcodes would have been a partial near-duplicate of one — the
thin-content failure the two-namespace design exists to avoid. A test now
guards the whole borough list.

Validated against the live corpus: 15 localities, no town collisions, no
authority duplicates, all 15 clear the threshold.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 20:20:52 +01:00
tudor b93eb3a691 Merge pull request 'feat(seo): the location layer — town, locality, authority and outcode pages (W2)' (#114) from feat/w2-location-layer into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 19s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m18s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 4s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 11m8s
Reviewed-on: #114
2026-08-21 17:56:05 +00:00
TudorandClaude Opus 5 6b871ce1e9 feat(places): ItemList and BreadcrumbList, and the e2e gate
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 4m33s
ItemList tells Google the page is a ranked set rather than prose;
BreadcrumbList puts the place in a hierarchy. School URLs in the markup are
absolute on the canonical host, since a relative URL in JSON-LD is ambiguous.

Eight journeys covering all four families, the two-namespace guarantee, the
threshold, the canonical, the sitemap and the local-versus-England line — the
last because that comparison is the reason these pages are not a name dropped
into a template.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:19:08 +01:00
TudorandClaude Opus 5 c981d89137 feat(places): town, locality, authority and outcode routes
Every generateStaticParams is gated behind PRERENDER_PLACES and wrapped in the
same try/catch the school route uses. The plan claimed authority pages were
'few enough to always prebuild' — but few enough still means the API must be
reachable at build time, and in CI it is not: the build failed with
ECONNREFUSED rather than degrading to ISR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:17:57 +01:00
TudorandClaude Opus 5 de5e790112 feat(places): place page client and view component
One component for all four families: they differ in what fills the registry,
not in what the page shows, so a second would be a second place to forget the
same change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:15:10 +01:00
TudorandClaude Opus 5 42138fc402 feat(places): submit place and outcode sitemaps
Separate children per family so Search Console reports the location layer's
indexation apart from the school pages' — which is the point of the index
built in W1, and the number the stop condition watches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:13:55 +01:00
TudorandClaude Opus 5 c5af476213 feat(places): /api/places registry and place detail endpoints
The registry is cached for the process and reset by the same admin endpoint
that rebuilds the sitemaps, so places and sitemap always describe the same
corpus rather than drifting apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:13:05 +01:00
TudorandClaude Opus 5 de853b90b3 feat(places): London localities and postcode districts
The GIAS town field puts 1,819 London schools under the single value
'London', so it cannot answer 'schools in Battersea' — a query that appears in
the baseline. No single field can: parliamentary constituency gives Battersea
but not Canary Wharf, admin_ward gives Canary Wharf but not Battersea, and
neither gives Clapham or Shoreditch. So a locality is curated, defined by the
postcode districts it covers, which needs no new ingestion.

A locality may not shadow a published town: the registry raises rather than
silently costing a page that carries real demand. One below the threshold is
logged rather than raising, because a locality can legitimately be too small.

The pipeline seed mirrors the module, with a test guarding the drift — the
same arrangement gias_codes has, and for the same reason: the backend image
does not contain pipeline/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:11:56 +01:00
TudorandClaude Opus 5 759d9f5cea feat(places): registry of towns and authorities
Two namespaces because 67 town names collide with an authority name and
neither set contains the other — postal towns cross authority boundaries, so
Bedford the town holds 104 schools against the authority's 86.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:10:37 +01:00
TudorandClaude Opus 5 555d3f0a7d docs(seo): implementation plan for the W2 location layer
Seven tasks: the place registry, London localities and outcodes, the places
API, per-family sitemaps, the shared place view, the four route families, and
structured data plus the e2e gate.

Two things the plan corrects against the spec. The backend image does not
contain pipeline/, so the curated locality list cannot live only in a dbt
seed — it follows the gias_codes.py precedent instead, canonical in backend
with the seed as a mirror. And NationalAverages is nested by phase rather than
flat, which the first draft read wrongly and would have rendered every page
without its England comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:09:35 +01:00
TudorandClaude Opus 5 ecc847091c docs(seo): design for the W2 location layer
Supersedes the original spec's W2. The Search Console baseline inverted its
ordering: every measured location query is town or district level, none is an
administrative area, and phase is part of the query rather than a filter.

Two problems the original design did not anticipate. 67 viable towns share a
name with a local authority, and the authority is the larger set in only 43 of
them — postal towns cross authority boundaries, so neither can absorb the
other. Two namespaces resolve it by construction. And the GIAS town field
collapses 1,819 London schools into one value, which a curated
locality-to-outcode seed solves without new ingestion.

Sizing is measured against the live 25,185-school corpus rather than
estimated: 783 viable towns, 1,760 outcodes, 154 authorities.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 18:09:35 +01:00
tudor b187a478c9 Merge pull request 'feat(seo): rewrite the C1 snippets to earn the click (W8)' (#113) from feat/seo-metadata-c1 into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #113
2026-08-20 23:26:12 +00:00
TudorandClaude Opus 5 c0547c45e5 feat(seo): rewrite the C1 snippets to earn the click (W8)
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 11s
The baseline says these pages already rank and are not clicked. 'compare
school performance' sits at position 6.1 with 0.43% CTR; 'compare schools' at
7.2 with 0.87%. The brand query 'school compare' draws 9.16% from the same
neighbourhood of the same results page, which rules out a ranking explanation
— when the snippet gives a reason to click, it gets clicked.

These SERPs are owned by the DfE's own 'Compare school performance' service.
The old title put a lowercase brand nobody searches for in the most valuable
pixels, then a near-paraphrase of that service's name. Beside the government's
own result it read as a lookalike.

Intent in the title, differentiator in the description. Titles now match what
people type, and the descriptions carry the one fact gov.uk does not publish:
how close you had to live to get a place.

/compare deliberately takes the tool phrasing rather than the homepage's, so
the two pages stop competing for one phrase. The root layout's default and
Open Graph copy were saying something different again; they now agree.

No hard school counts in any of it. The corpus moves with every data refresh
and this repo has already shipped one copy bug of that kind.

Tests guard the mechanics — SERP length, intent keyword, the differentiator,
no brand-first title — and leave the wording free to iterate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:21:08 +01:00
tudor 4a3928df9f Merge pull request 'fix(seo): a school is publishable on any year's results, not the latest' (#112) from fix/sitemap-any-year-data into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #112
2026-08-20 23:15:19 +00:00
26 changed files with 4694 additions and 20 deletions

No files matched your search

+141 -1
View File
@@ -35,6 +35,7 @@ from .data_loader import (
search_schools_typesense,
)
from .data_loader import get_data_info as get_db_info
from .places import build_place_registry
from .schemas import METRIC_DEFINITIONS, RANKING_COLUMNS, SCHOOL_COLUMNS
from .utils import clean_for_json, convert_to_native
@@ -58,6 +59,12 @@ MAX_SLUG_LENGTH = 60
# regenerate endpoint after a pipeline run.
_sitemaps: dict[str, str] | None = None
# Built from the same DataFrame the sitemap uses, so places and sitemap can
# never describe different corpora. Reset by the same admin endpoint.
_place_registry: dict | None = None
VALID_PLACE_KINDS = ("town", "locality", "authority", "outcode")
def _slugify(text: str) -> str:
text = text.lower()
@@ -170,6 +177,14 @@ SITEMAP_CHUNK_SIZE = 10_000
SITEMAP_CHILD_PREFIX = "/sitemaps"
def get_place_registry() -> dict:
"""The place registry, built once and cached for the process."""
global _place_registry
if _place_registry is None:
_place_registry = build_place_registry(load_school_data())
return _place_registry
def _urlset(rows: list[str]) -> str:
return "\n".join([
'<?xml version="1.0" encoding="UTF-8"?>',
@@ -179,6 +194,44 @@ def _urlset(rows: list[str]) -> str:
])
def _place_url(place) -> str:
"""The canonical path for a place. Two namespaces, per the spec.
Towns and localities share /schools/[place]; authorities take their own
prefix because 67 town names collide with an authority name and neither
set contains the other.
"""
if place.kind == "authority":
return f"/schools/authority/{place.slug}"
if place.kind == "outcode":
return f"/schools/near/{place.slug}"
return f"/schools/{place.slug}"
def _place_sitemap_rows(kinds: tuple[str, ...]) -> list[str]:
"""A <url> per place, plus a phase variant wherever that phase clears the
threshold on its own.
Phase is part of the query — "primary schools in beccles" — so each
variant is its own indexable page. Submitting only the bare place URL left
~950 of them reachable by nothing: absent from every sitemap, and not
linked from the place page either.
"""
rows: list[str] = []
for p in sorted(get_place_registry().values(), key=lambda p: (p.kind, p.slug)):
if p.kind not in kinds:
continue
rows.append(_url_element(BASE_URL + _place_url(p)))
# Outcodes carry no phase variants: nobody searches "primary schools
# in SW11", so the routes do not exist to submit.
if p.kind == "outcode":
continue
for phase in ("primary", "secondary"):
if p.publishes_phase(phase):
rows.append(_url_element(f"{BASE_URL}{_place_url(p)}/{phase}"))
return rows
def build_sitemaps() -> dict[str, str]:
"""Build the sitemap index and every child, keyed by name."""
df = load_school_data()
@@ -196,6 +249,17 @@ def build_sitemaps() -> dict[str, str]:
for n, chunk in enumerate(chunks, start=1):
children[f"schools-{n}.xml"] = _urlset(chunk)
# Separate children per family: Search Console reports coverage per
# submitted sitemap, which is how the location layer's indexation is
# measured apart from the school pages'.
for label, kinds in (("places", ("town", "locality", "authority")),
("outcodes", ("outcode",))):
rows = _place_sitemap_rows(kinds)
chunks = [rows[i:i + SITEMAP_CHUNK_SIZE]
for i in range(0, len(rows), SITEMAP_CHUNK_SIZE)] or [[]]
for n, chunk in enumerate(chunks, start=1):
children[f"{label}-{n}.xml"] = _urlset(chunk)
# On a sitemap index, lastmod means "when this sitemap file last changed",
# so generation time is the correct value here — unlike on a <url>, where
# it would be a claim about content we cannot support.
@@ -1117,6 +1181,78 @@ async def get_rankings(
}
@app.get("/api/places")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def list_places(request: Request):
"""Every published place. The sitemap and the link modules read this."""
registry = get_place_registry()
return {"places": [
{"kind": p.kind, "slug": p.slug, "name": p.name, "count": len(p.urns)}
for p in sorted(registry.values(), key=lambda p: (p.kind, p.slug))
]}
@app.get("/api/places/{kind}/{slug}")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_place(request: Request, kind: str, slug: str,
phase: Optional[str] = None):
"""One place: its schools ranked, and its local averages."""
if kind not in VALID_PLACE_KINDS:
raise HTTPException(status_code=404, detail="No such place")
place = get_place_registry().get(f"{kind}:{slug}")
if place is None:
raise HTTPException(status_code=404, detail="No such place")
df = load_latest_school_data()
rows = df[df["urn"].isin(place.urns)]
if phase:
wanted = PHASE_GROUPS.get(phase.lower())
if wanted and "phase" in rows.columns:
rows = rows[rows["phase"].fillna("").str.lower().isin(wanted)]
# The metric the page shows, and averages.
metric = "attainment_8_score" if phase == "secondary" else "rwm_expected_pct"
# Alphabetical, not by score. A place page is read by someone looking for
# a school they can name, and scanning for it is what the order should
# serve. /rankings is where the league-table ordering lives, and it keeps
# sorting by metric.
if "school_name" in rows.columns:
rows = rows.sort_values("school_name", key=lambda c: c.str.lower())
averages = {
m: (None if m not in rows.columns or rows[m].dropna().empty
else float(rows[m].dropna().mean()))
for m in ("rwm_expected_pct", "attainment_8_score")
}
cols = [c for c in SCHOOL_COLUMNS + ["latitude", "longitude", "phase",
"rwm_expected_pct", "attainment_8_score",
"total_pupils"]
if c in rows.columns]
return {
"place": {"kind": place.kind, "slug": place.slug, "name": place.name,
"count": len(place.urns),
"parent_authority": place.parent_authority,
# Every authority the place meaningfully sits in. SW19 is
# mostly Merton but partly Wandsworth; naming one asserts
# something false.
"authorities": [
{"name": name, "slug": _slugify(name), "count": n}
for name, n in place.authorities
],
# Only phases that clear the threshold, so the page links
# variants that exist rather than 404s.
"phases": [ph for ph in ("primary", "secondary")
if place.publishes_phase(ph)]},
"schools": clean_for_json(rows[cols]),
"averages": averages,
}
@app.get("/api/data-info")
@limiter.limit(f"{settings.rate_limit_per_minute}/minute")
async def get_data_info(request: Request):
@@ -1228,7 +1364,11 @@ async def regenerate_sitemap(
_: bool = Depends(verify_admin_api_key),
):
"""Rebuild and cache the sitemap from current school data. Called by Airflow after data updates."""
global _sitemaps
global _sitemaps, _place_registry
# Places and sitemap are rebuilt together — they read the same marts, and
# letting them drift apart would submit URLs for places that no longer
# exist.
_place_registry = None
_sitemaps = build_sitemaps()
n = sum(x.count("<url>") for x in _sitemaps.values())
return {"status": "ok", "urls": n, "sitemaps": len(_sitemaps)}
+54
View File
@@ -0,0 +1,54 @@
"""Curated London localities, defined by the postcode districts they cover.
The GIAS `town` field puts 1,819 London schools under the single value
"London", so it cannot answer "schools in Battersea" — a query that appears in
the Search Console baseline. No single field can: parliamentary constituency
gives Battersea but not Canary Wharf; postcodes.io's admin_ward gives Canary
Wharf but not Battersea; neither gives Clapham or Shoreditch, which are postal
and colloquial rather than administrative.
So this is curated. Where a locality ends is a judgement, not a fact, and a
reviewable file is the honest place for a judgement. No new ingestion is
needed — the corpus already carries postcodes.
This is the canonical copy. `pipeline/transform/seeds/locality_outcodes.csv`
mirrors it for anyone querying the warehouse directly; the backend image does
not contain `pipeline/`, which is why the module rather than the seed is
canonical. Same arrangement as `backend/gias_codes.py`.
A locality whose outcodes hold fewer than MIN_SCHOOLS schools is not
published, so a typo produces no page rather than an empty one. Places that
fail that check are logged at startup, because a locality you meant to publish
quietly not appearing is the failure worth hearing about.
Two rules for anything added here.
**Sub-borough districts only.** A London borough is a local authority and
already has a page at /schools/authority/[la] covering all of its schools; a
locality defined by two or three outcodes would be a partial, near-duplicate
subset of it. Hackney, Islington, Greenwich and Ealing were all in the first
draft for that reason and have been removed.
**The slug must not match a GIAS town.** "Richmond" did — GIAS has a Richmond
in North Yorkshire with 37 schools — so the London one could never publish.
The registry skips any locality that collides and logs it.
"""
# slug -> (display name, outcodes)
LOCALITY_OUTCODES: dict[str, tuple[str, tuple[str, ...]]] = {
"battersea": ("Battersea", ("SW11",)),
"canary-wharf": ("Canary Wharf", ("E14",)),
"clapham": ("Clapham", ("SW4",)),
"shoreditch": ("Shoreditch", ("EC2A", "E1")),
"peckham": ("Peckham", ("SE15",)),
"brixton": ("Brixton", ("SW2", "SW9")),
"camden-town": ("Camden Town", ("NW1",)),
"wimbledon": ("Wimbledon", ("SW19",)),
"putney": ("Putney", ("SW15",)),
"fulham": ("Fulham", ("SW6",)),
"chiswick": ("Chiswick", ("W4",)),
"stratford": ("Stratford", ("E15",)),
"walthamstow": ("Walthamstow", ("E17",)),
"tooting": ("Tooting", ("SW17",)),
"dulwich": ("Dulwich", ("SE21", "SE22")),
}
+308
View File
@@ -0,0 +1,308 @@
"""The place registry: what places the site publishes, and what is in each.
One module owns this question. The pages, the sitemap and the internal-link
modules all read from here, so the threshold and the collision rules exist in
exactly one place and are testable without a browser or a database.
Two namespaces, never one. 67 viable town names collide with a local
authority name, and the authority is the larger set in only 43 of them —
postal towns cross authority boundaries, so neither can absorb the other.
Keys are "<kind>:<slug>" so the collision cannot reappear in the dict.
"""
from __future__ import annotations
import logging
import re
from dataclasses import dataclass, field
logger = logging.getLogger(__name__)
# Five schools with publishable data. Below this a place has nothing to say
# that a list of schools does not, and publishing it is index bloat.
MIN_SCHOOLS = 5
@dataclass(frozen=True)
class Place:
kind: str # "town" | "locality" | "authority" | "outcode"
slug: str
name: str
urns: tuple[int, ...]
parent_authority: str | None # authority NAME, for the 301 target
# Every authority the place meaningfully sits in, largest first. A quarter
# of outcodes and a third of towns straddle a boundary — SW19 is mostly
# Merton but partly Wandsworth — so naming only one asserts something
# false. parent_authority stays single because a redirect needs one
# target; this is what the page shows.
authorities: tuple[tuple[str, int], ...] = ()
# URNs per phase, so the per-phase threshold can be applied without
# re-querying. A place with 30 primaries and 2 secondaries publishes a
# primary variant and no secondary one.
phase_urns: dict[str, tuple[int, ...]] = field(default_factory=dict)
def publishes_phase(self, phase: str) -> bool:
return len(self.phase_urns.get(phase, ())) >= MIN_SCHOOLS
@property
def key(self) -> str:
return f"{self.kind}:{self.slug}"
def _publishable_urns(df) -> set[int]:
"""URNs with something a page could state, deduplicated across years."""
from backend.app import _PUBLISHABLE_FIELDS
cols = [c for c in _PUBLISHABLE_FIELDS if c in df.columns]
if not cols:
return set()
return set(df.loc[df[cols].notna().any(axis=1), "urn"].astype(int))
# The measure a phase page is built around. A page with no results in this
# column has nothing a list of school names does not already give.
_PHASE_METRIC = {
"primary": "rwm_expected_pct",
"secondary": "attainment_8_score",
}
def _phase_urns(group, publishable: set[int]) -> dict[str, tuple[int, ...]]:
"""URNs per phase, counting only schools with a result for that phase.
Not merely "publishable". A school with an Ofsted grade and no results is
worth a page of its own and belongs in the place list, but it cannot
populate a phase page's results column — and the threshold is there to ask
whether that column will have anything in it.
Counting publishable schools instead let /schools/kent/primary publish
with none of its five rows carrying a result, and left 44 phase pages
majority-blank. It is the same rule as "no page without a local average",
which was never extended per phase.
All-through schools count toward both phases, matching the PHASE_GROUPS
mapping the search filters already use.
"""
from backend.app import PHASE_GROUPS
if "phase" not in group.columns:
return {}
lowered = group["phase"].fillna("").str.lower()
out: dict[str, tuple[int, ...]] = {}
for phase in ("primary", "secondary"):
wanted = PHASE_GROUPS.get(phase, set())
subset = group[lowered.isin(wanted)]
# The page lists every school of the phase; the threshold counts only
# those carrying a result, so a mostly-empty table never publishes.
metric = _PHASE_METRIC[phase]
with_result = (
{int(u) for u in subset.loc[subset[metric].notna(), "urn"]}
if metric in subset.columns else set()
)
if len(with_result & publishable) < MIN_SCHOOLS:
continue
urns = tuple(sorted({int(u) for u in subset["urn"]} & publishable))
if urns:
out[phase] = urns
return out
# A place is described by an authority when it holds at least a tenth of the
# schools, and at least two. GIAS carries occasional postcode errors — EN6
# lists two Shropshire schools among fourteen in Hertfordshire — and a bare
# "any authority present" rule would print those as though they were real.
# There is deliberately no cap on how many are named. An earlier cut stopped
# at three, which silently dropped the fourth in exactly the case where the
# information matters most — a genuinely fragmented place. The share rule is
# the only limit, and it already bounds the list at ten.
_AUTHORITY_MIN_SHARE = 0.10
_AUTHORITY_MIN_SCHOOLS = 2
def _authorities(group) -> tuple[tuple[str, int], ...]:
"""Authorities this place meaningfully sits in, largest first."""
from backend.app import EXCLUDED_FILTER_VALUES
if "local_authority" not in group.columns:
return ()
counts = group["local_authority"].dropna().value_counts()
total = int(counts.sum())
if not total:
return ()
kept = [
(str(name), int(n)) for name, n in counts.items()
if str(name) not in EXCLUDED_FILTER_VALUES
and n >= _AUTHORITY_MIN_SCHOOLS
and n / total >= _AUTHORITY_MIN_SHARE
]
# A place too small or too fragmented for the share rule still names its
# largest authority, or the page would say nothing about where it is.
if not kept:
for name, n in counts.items():
if str(name) not in EXCLUDED_FILTER_VALUES:
return ((str(name), int(n)),)
return ()
return tuple(kept)
def _parent_authority(authorities: tuple[tuple[str, int], ...]) -> str | None:
"""The 301 target: the largest authority a place sits in.
Derived from `authorities` rather than computed separately. The first cut
used `mode()` here while `authorities` used `value_counts()`, and on an
exact tie pandas does not guarantee the two pick the same name — so the
redirect could have pointed somewhere other than the authority the page
named first. One computation, one answer.
Deriving it also inherits the sentinel filter, so a place can no longer
redirect to /schools/authority/does-not-apply.
"""
return authorities[0][0] if authorities else None
def _group(df, column: str, kind: str, publishable: set[int]) -> dict[str, Place]:
"""One Place per distinct SLUG in `column` that clears the threshold.
Grouped by slug, not by raw value, because GIAS spells the same place
several ways and they all resolve to one URL. Five town slugs come from
more than one spelling: "London" (1,819 schools) and "LONDON" (12) both
slugify to `london`; Weston-super-Mare is split 14/19 across two
spellings; Newcastle-under-Lyme across three.
Grouping by raw value meant the later group simply overwrote the earlier
one in this dict — so /schools/london could have shown twelve schools
instead of 1,819, silently and depending on row order.
The display name is the most common spelling, which is the one a reader
expects to see.
"""
from backend.app import _slugify
if column not in df.columns:
return {}
working = df.assign(_slug=df[column].map(
lambda v: _slugify(str(v).strip()) if isinstance(v, str) and v.strip() else None))
working = working[working["_slug"].notna() & (working["_slug"] != "")]
out: dict[str, Place] = {}
for slug, group in working.groupby("_slug"):
slug = str(slug)
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
continue
spellings = group[column].dropna().value_counts()
if spellings.empty:
continue
name = str(spellings.index[0]).strip()
authorities = () if kind == "authority" else _authorities(group)
place = Place(
kind=kind, slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable),
)
out[place.key] = place
return out
# "SW11 2AA" -> "SW11". Two letters max, one or two digits, optional letter.
_OUTCODE_RE = re.compile(r"^([A-Z]{1,2}\d{1,2}[A-Z]?)\s")
def _outcode(postcode) -> str | None:
if not isinstance(postcode, str):
return None
m = _OUTCODE_RE.match(postcode.upper().strip())
return m.group(1) if m else None
def _outcode_places(df, publishable: set[int]) -> dict[str, Place]:
"""One Place per postcode district clearing the threshold.
These carry no phase variants: nobody searches "primary schools in SW11".
"""
if "postcode" not in df.columns:
return {}
working = df.assign(_oc=df["postcode"].map(_outcode))
working = working[working["_oc"].notna()]
out: dict[str, Place] = {}
for oc, group in working.groupby("_oc"):
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
continue
authorities = _authorities(group)
place = Place(kind="outcode", slug=str(oc).lower(), name=str(oc),
urns=urns, parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable))
out[place.key] = place
return out
def _locality_places(df, publishable: set[int],
town_slugs: set[str]) -> dict[str, Place]:
"""One Place per curated locality clearing the threshold."""
from backend.localities import LOCALITY_OUTCODES
if "postcode" not in df.columns:
return {}
working = df.assign(_oc=df["postcode"].map(_outcode))
out: dict[str, Place] = {}
for slug, (name, outcodes) in LOCALITY_OUTCODES.items():
if slug in town_slugs:
# Skip, do not raise. The guard exists so a locality never
# silently shadows a town — skipping achieves that, and the error
# log makes it loud.
#
# Raising here took down sitemap generation for all 25,000 school
# pages when "richmond" met the GIAS town Richmond in North
# Yorkshire. Worse, GIAS town names change without any code change,
# so a raise means curated data can break the site spontaneously.
# A curation mistake must cost one page, not the sitemap.
logger.error(
"locality %r collides with the published town of the same "
"slug and has been skipped; rename it or remove it", slug)
continue
group = working[working["_oc"].isin(outcodes)]
urns = tuple(sorted({int(u) for u in group["urn"]} & publishable))
if len(urns) < MIN_SCHOOLS:
# Not an error — a locality can legitimately be too small. Logged
# because one you meant to publish quietly vanishing is the
# failure worth hearing about.
logger.warning(
"locality %s (%s) has %d publishable schools, below the "
"threshold of %d - not published",
slug, ", ".join(outcodes), len(urns), MIN_SCHOOLS)
continue
authorities = _authorities(group)
place = Place(kind="locality", slug=slug, name=name, urns=urns,
parent_authority=_parent_authority(authorities),
authorities=authorities,
phase_urns=_phase_urns(group, publishable))
out[place.key] = place
return out
def build_place_registry(df) -> dict[str, Place]:
"""Every place the site publishes, keyed by "<kind>:<slug>"."""
if df.empty or "urn" not in df.columns:
return {}
publishable = _publishable_urns(df)
registry: dict[str, Place] = {}
registry.update(_group(df, "local_authority", "authority", publishable))
towns = _group(df, "town", "town", publishable)
registry.update(towns)
town_slugs = {p.slug for p in towns.values()}
registry.update(_locality_places(df, publishable, town_slugs))
registry.update(_outcode_places(df, publishable))
return registry
+391
View File
@@ -0,0 +1,391 @@
"""Tests for the place registry (spec 2026-08-21).
The registry is built from the in-memory school DataFrame, so these build a
small frame directly rather than touching a database.
"""
import numpy as np
import pandas as pd
import pytest
from backend.places import MIN_SCHOOLS, build_place_registry
def _df(rows: list[dict]) -> pd.DataFrame:
base = {
"year": 202425, "ofsted_grade": 2.0, "ofsted_date": None,
"rwm_expected_pct": 60.0, "attainment_8_score": np.nan,
"phase": "Primary", "postcode": "AA1 1AA",
}
return pd.DataFrame([{**base, **r} for r in rows])
def _town(n: int, town: str, la: str, start: int = 100000, **kw) -> list[dict]:
"""`start` offsets the URNs so two calls can describe different schools —
the Bedford case needs two authorities' worth of distinct URNs in one
town."""
return [
{"urn": start + i, "school_name": f"{town} School {i}",
"town": town, "local_authority": la, **kw}
for i in range(n)
]
def test_town_clearing_the_threshold_is_published():
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "Brentwood", "Essex")))
assert "town:brentwood" in reg
assert reg["town:brentwood"].name == "Brentwood"
assert len(reg["town:brentwood"].urns) == MIN_SCHOOLS
def test_town_below_the_threshold_is_not_published():
reg = build_place_registry(_df(_town(MIN_SCHOOLS - 1, "Crosby", "Sefton")))
assert "town:crosby" not in reg
def test_a_town_below_threshold_still_names_its_authority():
# The route layer needs somewhere to 301 to.
reg = build_place_registry(_df(
_town(MIN_SCHOOLS - 1, "Crosby", "Sefton") + _town(MIN_SCHOOLS, "Bootle", "Sefton")))
assert "authority:sefton" in reg
def test_town_and_authority_of_the_same_name_are_separate_places():
# 67 real collisions. Neither set contains the other: Bedford the town has
# 104 schools, Bedford the authority 86, because postal towns cross
# authority boundaries.
rows = (_town(MIN_SCHOOLS, "Bedford", "Bedford")
+ _town(MIN_SCHOOLS, "Bedford", "Central Bedfordshire", start=200000))
reg = build_place_registry(_df(rows))
town, authority = reg["town:bedford"], reg["authority:bedford"]
assert set(town.urns) != set(authority.urns)
assert len(town.urns) == MIN_SCHOOLS * 2 # both authorities' schools
assert len(authority.urns) == MIN_SCHOOLS # only this authority's
def test_schools_without_publishable_data_do_not_count_toward_the_threshold():
rows = _town(MIN_SCHOOLS, "Ghosttown", "Nowhere")
for r in rows:
r["rwm_expected_pct"] = np.nan
r["ofsted_grade"] = np.nan
reg = build_place_registry(_df(rows))
assert "town:ghosttown" not in reg
def test_blank_town_is_ignored():
rows = _town(MIN_SCHOOLS, "", "Essex")
reg = build_place_registry(_df(rows))
assert not any(k.startswith("town:") for k in reg)
def test_a_school_is_counted_once_even_with_several_years_of_rows():
rows = []
for year in (202324, 202425):
rows += [{**r, "year": year} for r in _town(MIN_SCHOOLS, "Beccles", "Suffolk")]
reg = build_place_registry(_df(rows))
assert len(reg["town:beccles"].urns) == MIN_SCHOOLS
def test_locality_groups_schools_by_outcode(monkeypatch):
# The GIAS town field collapses 1,819 London schools into "London", so a
# locality is defined by its postcode districts instead.
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"battersea": ("Battersea", ("SW11",))})
rows = _town(MIN_SCHOOLS, "London", "Wandsworth")
for r in rows:
r["postcode"] = "SW11 2AA"
reg = build_place_registry(_df(rows))
assert reg["locality:battersea"].name == "Battersea"
assert len(reg["locality:battersea"].urns) == MIN_SCHOOLS
def test_locality_below_the_threshold_is_not_published(monkeypatch):
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"nowhere": ("Nowhere", ("ZZ99",))})
reg = build_place_registry(_df(_town(MIN_SCHOOLS, "London", "Wandsworth")))
assert "locality:nowhere" not in reg
def test_a_locality_may_not_shadow_a_viable_town(monkeypatch, caplog):
"""A colliding locality is skipped loudly, and the town survives.
This used to raise, which took down sitemap generation for all 25,000
school pages the first time a curated slug met a real GIAS town. Curated
data must not be able to break the site — and GIAS town names change with
no code change at all, so the raise could fire spontaneously.
"""
import logging
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"brentwood": ("Brentwood", ("CM13",))})
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
with caplog.at_level(logging.ERROR):
reg = build_place_registry(_df(rows))
assert "locality:brentwood" not in reg # skipped
assert "town:brentwood" in reg # the town is untouched
assert "brentwood" in caplog.text # and it was loud about it
def test_a_locality_collision_does_not_break_the_rest_of_the_registry(monkeypatch):
# The whole point of skipping rather than raising.
from backend import localities
monkeypatch.setattr(localities, "LOCALITY_OUTCODES",
{"brentwood": ("Brentwood", ("CM13",))})
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
reg = build_place_registry(_df(rows))
assert "authority:essex" in reg
assert "outcode:cm13" in reg
def test_outcode_places_are_built_from_postcodes():
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "CM13 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:cm13"].name == "CM13"
assert len(reg["outcode:cm13"].urns) == MIN_SCHOOLS
def test_malformed_postcodes_do_not_create_places():
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
for r in rows:
r["postcode"] = "not a postcode"
reg = build_place_registry(_df(rows))
assert not any(k.startswith("outcode:") for k in reg)
def test_every_curated_locality_is_structurally_valid():
# Guards the hand-maintained file: real slug, real name, real outcodes.
import re
from backend.localities import LOCALITY_OUTCODES
assert LOCALITY_OUTCODES, "the curated locality list must not be empty"
for slug, (name, outcodes) in LOCALITY_OUTCODES.items():
assert re.fullmatch(r"[a-z0-9-]+", slug), slug
assert name.strip() == name and name, slug
assert outcodes, f"{slug} has no outcodes"
for oc in outcodes:
assert re.fullmatch(r"[A-Z]{1,2}\d{1,2}[A-Z]?", oc), (slug, oc)
def test_the_pipeline_seed_mirrors_the_canonical_module():
"""Two copies with no drift guard is worse than one copy.
backend/localities.py is canonical because the backend image does not
contain pipeline/. The seed exists so the warehouse can join on the same
definitions, and this is what stops the two diverging — the same
arrangement assert_gias_code_names_match_seed.sql gives gias_codes.
"""
import csv
from pathlib import Path
from backend.localities import LOCALITY_OUTCODES
seed_path = (Path(__file__).resolve().parents[2]
/ "pipeline/transform/seeds/locality_outcodes.csv")
assert seed_path.exists(), f"missing seed mirror at {seed_path}"
seed = {
row["locality_slug"]: (row["locality_name"],
tuple(row["outcodes"].split("|")))
for row in csv.DictReader(seed_path.open())
}
assert seed == LOCALITY_OUTCODES
def test_no_curated_locality_names_a_london_borough():
"""Boroughs are authorities and already have a page.
A locality defined by two or three outcodes inside a borough would be a
partial, near-duplicate subset of that authority page — the exact
thin-content failure the two-namespace design exists to avoid. Hackney,
Islington, Greenwich and Ealing were all in the first draft.
Hardcoded rather than read from the corpus because this must fail in CI,
where there is no database.
"""
from backend.localities import LOCALITY_OUTCODES
boroughs = {
"barking-and-dagenham", "barnet", "bexley", "brent", "bromley",
"camden", "croydon", "ealing", "enfield", "greenwich", "hackney",
"hammersmith-and-fulham", "haringey", "harrow", "havering",
"hillingdon", "hounslow", "islington", "kensington-and-chelsea",
"kingston-upon-thames", "lambeth", "lewisham", "merton", "newham",
"redbridge", "richmond-upon-thames", "southwark", "sutton",
"tower-hamlets", "waltham-forest", "wandsworth", "westminster",
}
named = boroughs & set(LOCALITY_OUTCODES)
assert not named, (
f"these are boroughs, not districts: {sorted(named)} - they already "
"have an authority page covering every school"
)
def test_a_place_names_every_authority_it_straddles():
"""SW19 is mostly Merton but partly Wandsworth.
A quarter of viable outcodes and a third of viable towns cross an
authority boundary, so naming only the largest asserts something false.
"""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
names = [n for n, _ in reg["outcode:sw19"].authorities]
assert names == ["Merton", "Wandsworth"] # largest first
assert dict(reg["outcode:sw19"].authorities)["Wandsworth"] == 7
def test_the_redirect_target_stays_a_single_authority():
# parent_authority and authorities do different jobs: a 301 needs one
# target, the page needs the truth.
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert reg["outcode:sw19"].parent_authority == "Merton"
def test_a_stray_authority_below_the_share_threshold_is_not_named():
# GIAS carries postcode errors — EN6 lists two Shropshire schools among
# fourteen in Hertfordshire. Printing those as though real would be worse
# than omitting them.
rows = (_town(30, "Barnet", "Hertfordshire", start=300000)
+ _town(1, "Barnet", "Shropshire", start=400000))
for r in rows:
r["postcode"] = "EN6 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:en6"].authorities] == ["Hertfordshire"]
def test_a_sentinel_authority_is_never_named():
rows = (_town(20, "London", "Merton", start=300000)
+ _town(6, "London", "Does not apply", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
reg = build_place_registry(_df(rows))
assert [n for n, _ in reg["outcode:sw19"].authorities] == ["Merton"]
def test_a_place_always_names_at_least_one_authority():
# Even when every authority is below the share threshold, the page has to
# say where the place is.
rows = []
for i, la in enumerate(["A", "B", "C", "D", "E", "F", "G"]):
rows += _town(1, "Fragmented", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
place = reg.get("town:fragmented")
assert place is not None
assert len(place.authorities) == 1
def test_every_qualifying_authority_is_named_with_no_cap():
"""An earlier cut stopped at three, dropping the fourth silently.
That truncation bit exactly where the information matters most — a
genuinely fragmented place — and nothing recorded it.
"""
rows = []
for i, la in enumerate(["Hackney", "Lambeth", "Westminster", "Lewisham"]):
rows += _town(3, "Fourway", la, start=300000 + i * 100)
reg = build_place_registry(_df(rows))
assert len(reg["town:fourway"].authorities) == 4
def test_the_redirect_target_is_the_authority_named_first():
"""They were computed separately — mode() against value_counts() — and on
an exact tie pandas does not guarantee the two agree."""
rows = (_town(26, "London", "Merton", start=300000)
+ _town(7, "London", "Wandsworth", start=400000))
for r in rows:
r["postcode"] = "SW19 1AA"
place = build_place_registry(_df(rows))["outcode:sw19"]
assert place.parent_authority == place.authorities[0][0]
def test_a_place_never_redirects_to_a_sentinel_authority():
# Deriving the parent from `authorities` inherits its sentinel filter.
rows = (_town(6, "Someplace", "Does not apply", start=300000)
+ _town(5, "Someplace", "Essex", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:someplace"].parent_authority == "Essex"
def test_spellings_of_one_place_are_merged_not_overwritten():
"""GIAS spells the same place several ways, and they share a URL.
"London" (1,819 schools) and "LONDON" (12) both slugify to `london`.
Grouping by raw value let the later group overwrite the earlier one, so
the page could have shown twelve schools instead of 1,819 — silently, and
depending on row order.
"""
rows = (_town(6, "Weston-super-Mare", "North Somerset", start=300000)
+ _town(5, "Weston-Super-Mare", "North Somerset", start=400000))
reg = build_place_registry(_df(rows))
assert len(reg["town:weston-super-mare"].urns) == 11
def test_the_merged_place_takes_its_most_common_spelling():
rows = (_town(9, "Newcastle-under-Lyme", "Staffordshire", start=300000)
+ _town(5, "NEWCASTLE-UNDER-LYME", "Staffordshire", start=400000))
reg = build_place_registry(_df(rows))
assert reg["town:newcastle-under-lyme"].name == "Newcastle-under-Lyme"
def test_a_phase_page_needs_results_not_merely_publishable_schools():
"""/schools/kent/primary published with none of its five rows scored.
The threshold counted schools that were publishable — a result OR an
Ofsted grade — while the page exists for its results column. Forty-four
phase pages were majority-blank; one had no results at all.
"""
rows = _town(MIN_SCHOOLS, "Kent", "Kent")
for r in rows:
r["rwm_expected_pct"] = np.nan # Ofsted only, no results
reg = build_place_registry(_df(rows))
assert "town:kent" in reg # the place still publishes
assert not reg["town:kent"].publishes_phase("primary")
def test_a_phase_page_publishes_once_enough_schools_carry_a_result():
rows = _town(MIN_SCHOOLS, "Beccles", "Suffolk")
reg = build_place_registry(_df(rows))
assert reg["town:beccles"].publishes_phase("primary")
def test_a_publishing_phase_page_still_lists_its_unscored_schools():
"""The threshold gates whether the page exists; it does not filter rows.
A parent looking up a school by name has to find it whether or not it
published results.
"""
scored = _town(MIN_SCHOOLS, "Beccles", "Suffolk", start=300000)
unscored = _town(2, "Beccles", "Suffolk", start=400000)
for r in unscored:
r["rwm_expected_pct"] = np.nan
reg = build_place_registry(_df(scored + unscored))
place = reg["town:beccles"]
assert place.publishes_phase("primary")
assert len(place.phase_urns["primary"]) == MIN_SCHOOLS + 2
def test_the_secondary_threshold_counts_its_own_metric():
# A town full of scored primaries must not thereby publish a secondary page.
rows = _town(MIN_SCHOOLS, "Brentwood", "Essex")
reg = build_place_registry(_df(rows))
assert not reg["town:brentwood"].publishes_phase("secondary")
+91
View File
@@ -0,0 +1,91 @@
"""Tests for the places API (spec 2026-08-21)."""
import numpy as np
import pandas as pd
import pytest
from fastapi.testclient import TestClient
def _schools_df() -> pd.DataFrame:
base = {
"local_authority": "Essex", "school_type": "Academy",
"phase": "Primary", "year": 202425, "ofsted_grade": 2.0,
"ofsted_date": None, "attainment_8_score": np.nan,
"town": "Brentwood", "postcode": "CM13 1AA", "status": "Open",
"address": "1 Test Street", "latitude": 51.6, "longitude": 0.3,
}
return pd.DataFrame([
{**base, "urn": 100000 + i, "school_name": f"Brentwood School {i}",
"rwm_expected_pct": 50.0 + i}
for i in range(6)
])
@pytest.fixture()
def client(monkeypatch):
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _schools_df)
monkeypatch.setattr(app_module, "load_latest_school_data", _schools_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return TestClient(app_module.app, raise_server_exceptions=False)
def test_registry_lists_each_published_place(client):
body = client.get("/api/places").json()
slugs = {(p["kind"], p["slug"]) for p in body["places"]}
assert ("town", "brentwood") in slugs
assert ("authority", "essex") in slugs
assert ("outcode", "cm13") in slugs
def test_registry_carries_a_count_per_place(client):
body = client.get("/api/places").json()
town = next(p for p in body["places"] if p["slug"] == "brentwood")
assert town["count"] == 6
def test_place_detail_returns_its_schools_alphabetically(client):
"""A place page is read by someone looking for a school they can name.
Scanning for it is what the order should serve, so the list is A-Z.
/api/rankings is where the league-table ordering lives.
"""
body = client.get("/api/places/town/brentwood").json()
assert body["place"]["name"] == "Brentwood"
names = [s["school_name"] for s in body["schools"]]
assert names == sorted(names, key=str.lower)
def test_place_ordering_ignores_case(client):
body = client.get("/api/places/town/brentwood").json()
names = [s["school_name"] for s in body["schools"]]
# A capitalised name must not sort ahead of every lowercase one.
assert names == sorted(names, key=str.lower)
def test_the_rankings_endpoint_still_ranks_by_metric(client):
# Alphabetical is a place-page decision, not a site-wide one.
body = client.get("/api/rankings?metric=rwm_expected_pct&phase=primary").json()
scores = [r["rwm_expected_pct"] for r in body.get("rankings", [])
if r.get("rwm_expected_pct") is not None]
assert scores == sorted(scores, reverse=True)
def test_place_detail_carries_the_local_average(client):
body = client.get("/api/places/town/brentwood").json()
# 50..55 inclusive
assert body["averages"]["rwm_expected_pct"] == pytest.approx(52.5)
def test_phase_filter_narrows_the_school_list(client):
body = client.get("/api/places/town/brentwood?phase=secondary").json()
assert body["schools"] == []
def test_unknown_place_404s(client):
assert client.get("/api/places/town/atlantis").status_code == 404
def test_unknown_kind_404s(client):
assert client.get("/api/places/planet/mars").status_code == 404
+72 -1
View File
@@ -59,9 +59,14 @@ def static_child(monkeypatch) -> str:
def test_every_loc_uses_the_www_host(sitemaps):
# The apex 301s to www. A <loc> that redirects burns a crawl per URL.
# Checked across every file, index included, not just one.
#
# A child can legitimately be empty — this fixture holds two schools and no
# town clearing the threshold — so the presence check applies only to files
# that carry URLs. The absence check applies to all of them.
for name, xml in sitemaps.items():
assert "https://www.schoolcompare.co.uk" in xml, name
assert "https://schoolcompare.co.uk" not in xml, name
if "<loc>" in xml:
assert "https://www.schoolcompare.co.uk" in xml, name
def test_school_with_results_is_listed(schools_child):
@@ -217,3 +222,69 @@ def test_school_with_no_results_in_any_year_is_still_omitted(monkeypatch):
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
assert "/school/100002" not in app_module.build_sitemaps()["schools-1.xml"]
def _places_df() -> pd.DataFrame:
base = {
"local_authority": "Essex", "school_type": "Academy",
"phase": "Primary", "year": 202425, "ofsted_grade": 2.0,
"ofsted_date": None, "attainment_8_score": np.nan,
"town": "Brentwood", "postcode": "CM13 1AA",
}
return pd.DataFrame([
{**base, "urn": 100000 + i, "school_name": f"Brentwood School {i}",
"rwm_expected_pct": 60.0}
for i in range(6)
])
@pytest.fixture()
def place_sitemaps(monkeypatch) -> dict:
from backend import app as app_module
monkeypatch.setattr(app_module, "load_school_data", _places_df)
monkeypatch.setattr(app_module, "_place_registry", None)
return app_module.build_sitemaps()
def test_place_children_are_listed_in_the_index(place_sitemaps):
index = place_sitemaps["sitemap.xml"]
assert "/sitemaps/places-1.xml" in index
assert "/sitemaps/outcodes-1.xml" in index
def test_town_and_authority_urls_use_their_own_namespaces(place_sitemaps):
xml = place_sitemaps["places-1.xml"]
assert "<loc>https://www.schoolcompare.co.uk/schools/brentwood</loc>" in xml
assert "<loc>https://www.schoolcompare.co.uk/schools/authority/essex</loc>" in xml
def test_outcode_urls_live_in_their_own_child(place_sitemaps):
assert "/schools/near/cm13" in place_sitemaps["outcodes-1.xml"]
assert "/schools/near/cm13" not in place_sitemaps["places-1.xml"]
def test_place_urls_carry_no_priority_or_changefreq(place_sitemaps):
for name in ("places-1.xml", "outcodes-1.xml"):
assert "<priority>" not in place_sitemaps[name]
assert "<changefreq>" not in place_sitemaps[name]
def test_phase_variants_are_submitted_where_the_phase_clears_the_threshold(place_sitemaps):
# "primary schools in beccles" is the query shape the baseline showed, so
# each variant is its own page and has to be submitted. Emitting only the
# bare place URL left ~950 of them reachable by nothing.
xml = place_sitemaps["places-1.xml"]
assert "<loc>https://www.schoolcompare.co.uk/schools/brentwood/primary</loc>" in xml
def test_a_phase_below_its_own_threshold_is_not_submitted(place_sitemaps):
# The fixture is six primaries and no secondaries.
xml = place_sitemaps["places-1.xml"]
assert "/schools/brentwood/secondary" not in xml
def test_outcodes_get_no_phase_variants(place_sitemaps):
# Nobody searches "primary schools in CM13"; the routes do not exist.
xml = place_sitemaps["outcodes-1.xml"]
assert "/primary" not in xml and "/secondary" not in xml
File diff suppressed because it is too large. Load diff
@@ -0,0 +1,278 @@
# W2: The Location Layer — Design
Date: 2026-08-21
Status: awaiting review
Supersedes: workstream W2 in `2026-08-20-seo-programme-design.md`
Scope note: this covers four page families in one spec. Splitting them — towns
and authorities first, outcodes and localities after — was proposed and
declined in favour of building the layer in one pass. The decomposition
argument was that the curated locality seed needs human review and would hold
up 783 pages of measured demand behind it; that risk is accepted here, and the
implementation plan should sequence the seed early enough that review time
does not become the critical path.
## Problem
Location intent is the largest unserved demand the site has. In the 16-month
Search Console baseline it draws **874 impressions, one click, average
position 49.5**. The site does not compete.
Unlike named-school queries — which the same baseline showed to be
navigational and unwinnable, since a parent typing "audley junior school"
wants that school's own website — location queries have no incumbent owner.
Nobody owns "primary schools in Brentwood" the way a school owns its name.
The cause is structural: the site has no page about a place. Every competitor
ranking above it does.
## What the demand actually looks like
Every location query in the baseline is **town or district level**. Not one is
an administrative area:
| Query | Impressions | Position |
|-------|-------------|----------|
| colleges in solihull | 112 | 51.2 |
| schools in ramsey | 64 | 42.5 |
| schools in crosby | 57 | 47.7 |
| primary schools in beccles | 44 | 40.9 |
| private schools in battersea | 41 | 71.9 |
| secondary schools in brentwood | 37 | 56.1 |
| secondary schools in canary wharf | 30 | 35.9 |
Three patterns follow directly, and they drive the whole design.
**Towns, not authorities.** The superseded W2 put `/schools/[la]` first and
towns second. The data inverts that. Brentwood appears four times in different
phrasings; Beccles twice. Both are towns, not authorities.
**Phase is part of the query**, not a filter applied afterwards: "primary
schools in beccles", "secondary schools in brentwood", "colleges in solihull".
**London is searched by district** — Battersea, Canary Wharf — and the GIAS
`town` field cannot serve it at all.
## Measured sizing
Counted against the live corpus of 25,185 schools, not estimated.
| Family | Viable (≥5 schools) | Below threshold |
|--------|--------------------|-----------------|
| Towns | **783** | 907 → redirect to authority |
| Outcodes | **1,760** | 305 |
| Local authorities | 154 | — |
| London localities | ~100–150 (curated) | — |
With phase variants — 783 town pages plus roughly 700 primary and 250
secondary variants, 154 authorities across three variants, 1,760 outcodes and
the curated localities — the total lands near **4,000 pages**. Phase variants
need their own threshold: there are 17,426 primaries but only 4,456 secondaries nationally,
so most towns will support a primary page and not a secondary one.
## Two design problems this spec exists to solve
### 1. Town and authority names collide, and neither contains the other
67 viable towns share a name with a local authority. The obvious fix — let the
authority absorb the town, since it sounds like a superset — **does not work**:
| Place | Schools in the town | Schools in the authority |
|-------|--------------------|-----------------------|
| Bedford | 104 | 86 |
| Birmingham | 520 | 518 |
| Derby | 157 | 119 |
| Doncaster | 152 | 145 |
The authority is the larger set in only 43 of the 67. Postal towns cross
authority boundaries, so these are overlapping sets that happen to share a
name. Publishing both into one namespace produces near-duplicate pages, which
is the specific failure that sinks programmatic SEO.
**Resolution: two namespaces.**
```
/schools/[place] towns and London localities
/schools/[place]/primary
/schools/[place]/secondary
/schools/authority/[la] local authorities
/schools/authority/[la]/primary
/schools/authority/[la]/secondary
/schools/near/[outcode]
```
Outcodes carry no phase variants: nobody searches "primary schools in SW11",
so the variants would be pages without demand.
Every collision disappears by construction. `/schools/[place]` keeps the clean
URL for the pattern that carries the demand; authorities get a namespace whose
purpose is genuinely different — admissions are authority-run, and the
authority page is the one that can speak to catchment policy and LA averages.
A place page and an authority page of the same name must each say plainly
which set of schools they cover, or they read as duplicates to a reader even
when they differ in fact.
### 2. London has no locality field
`town` collapses **1,819 London schools into the single value "London"**. A
page listing all of them is useless, and borough pages do not help because
people search "Battersea", not "Wandsworth".
No single field solves it:
| Search term | `parliamentary_constituency` | postcodes.io `admin_ward` |
|-------------|------------------------------|---------------------------|
| Battersea | **Battersea** ✓ | Northcote / Wandsworth Town ✗ |
| Canary Wharf | Poplar and Limehouse ✗ | **Canary Wharf** ✓ |
| Vauxhall | Vauxhall and Camberwell Green ✗ | **Vauxhall** ✓ |
And neither covers Clapham, Shoreditch or Peckham, which are postal and
colloquial rather than administrative.
**Resolution: a curated seed mapping locality to outcodes.**
```
pipeline/transform/seeds/locality_outcodes.csv
locality_slug,locality_name,outcodes,region
battersea,Battersea,"SW11|SW8",London
canary-wharf,Canary Wharf,"E14",London
clapham,Clapham,"SW4|SW9",London
```
This needs **no new ingestion** — the corpus already has postcodes. It puts
the fuzzy, contested part of the problem in a reviewable file rather than in
derived logic, which suits it: locality boundaries are a judgement, not a
fact. The repo already uses dbt seeds for curated reference data
(`la_code_names.csv`, `gias_code_names.csv`), so this follows an established
pattern.
The seed generalises past London. Any colloquial place — Jesmond, Chorlton,
Clifton — can be defined by its outcodes without a schema change.
**Constraint:** a locality slug may not collide with a viable town slug. The
place registry enforces this and fails the build rather than silently
shadowing a town.
## Architecture
### The place registry
One module owns the question "what places do we publish, and what is in each".
Everything else reads from it: the pages, the sitemap, the internal links.
```
backend/places.py
Place = { kind: "town"|"locality"|"authority"|"outcode",
slug, name, urn_list, parent_authority | None }
build_place_registry(df) -> dict[str, Place]
place_schools(slug, phase=None) -> list[School]
```
Built once at startup from the same DataFrame the sitemap uses, and rebuilt by
the existing `/api/admin/regenerate-sitemap` path after a pipeline run.
Registry construction is where the threshold, the collision rules and the
seed's uniqueness constraint are enforced — in one place, testable without a
browser or a database.
### API
```
GET /api/places the registry: slug, kind, name, count
GET /api/places/{slug}?phase= aggregate + ranked schools for one place
```
`/api/places` is what the sitemap and the internal-link modules enumerate.
### Routes
Next App Router, ISR with the same 7-day revalidate the school pages use.
`generateStaticParams` gated behind an env flag, matching
`PRERENDER_SCHOOLS`, because 3,900 more routes cannot be statically built in
CI on every deploy.
## What each page must contain
A place page that is a name substituted into a template is the thing Google's
helpful-content stance exists to demote. Each page carries computed local
facts that exist nowhere else on the site:
- **H1** matching the query: "Primary schools in Brentwood"
- **Counts framed usefully**: "29 schools, 4 rated Outstanding"
- **A ranked table** of the top 20 on the phase's headline metric —
`rwm_expected_pct` for primary, `attainment_8_score` for secondary, and for
an unphased place page the metric matching whichever phase it holds more of
- **The local average against the England average** — the one number a parent
cannot get from a list
- **Ofsted grade distribution** for the place
- **A map**
- **Links to neighbouring places** and to the parent authority
- **An FAQ block**, feeding `FAQPage` structured data
- **A link to every school page in scope** — this is what finally de-orphans
the 23,000 school pages the original spec identified as near-orphans
## Thin-page controls
Three, and they are the difference between a location layer and index bloat:
1. **Five schools with current data minimum.** Below it, 301 to the parent
authority. This drops 907 towns and 305 outcodes.
2. **Per-phase thresholds.** A town with 30 primaries and 2 secondaries
publishes a primary page and no secondary page.
3. **No page without a local average.** If a place has too few schools with
results to compute one, it has nothing to say that a list does not, and it
falls back to the authority.
## Sitemap
Two new children in the existing index: `/sitemaps/places-{n}.xml` and
`/sitemaps/outcodes-{n}.xml`. Per-family children are why the index was built
in W1 — Search Console reports coverage per submitted sitemap, so indexation
of the location layer is measurable separately from the school pages.
## Testing
Per `CLAUDE.md`, user-facing behaviour extends `e2e/tests/journeys.spec.ts` in
the same PR.
**Unit (registry, no DB):** threshold enforcement; a sub-threshold town
resolves to its authority; a locality slug colliding with a town fails the
build; Bedford's town and authority pages hold different URN sets; per-phase
thresholds.
**Backend:** `/api/places` shape; `/api/places/{slug}` aggregate correctness
against a fixture; unknown slug 404s.
**e2e:** a known town, authority, locality and outcode page each render with
the expected count; a below-threshold town 301s; every place page declares a
canonical and appears in the sitemap; `/schools/bedford` and
`/schools/authority/bedford` both resolve and state which set they cover.
## Risks
**Index bloat** is the failure mode of every programmatic SEO programme. The
three controls above are the answer, and the per-family sitemap is how we find
out early if they were not enough.
**Helpful-content exposure.** Templated location pages are exactly what
Google's stance targets. The mitigation is that every page carries real
computed local data — counts, distributions, local-versus-national comparison
— rather than a name dropped into boilerplate. If indexation of the places
sitemap stalls below roughly half, that is the signal to stop and rethink
rather than to add more pages.
**Build cost.** ~4,000 additional ISR routes on top of 23,000 school pages.
The env-flag gate on `generateStaticParams` keeps CI viable.
**Curation drift.** The locality seed is hand-maintained and will go stale as
places change. It is small and reviewable, and a dbt test asserts every seed
outcode matches at least one school so a typo fails the pipeline rather than
publishing an empty page.
## Out of scope
Catchment-area estimation. It is a strong driver for this cluster and
`fact_admissions` carries the distances, but it is a modelling problem with
real accuracy risk and deserves its own design.
+256 -3
View File
@@ -1649,13 +1649,29 @@ const CANONICAL_ROUTES: Array<[string, string]> = [
['/admissions', 'https://www.schoolcompare.co.uk/admissions'],
];
/**
* Next normalises canonical URLs against `trailingSlash: false`, so the root
* ships as `https://www.schoolcompare.co.uk` with no slash while every other
* route keeps its path. Both forms address the same document, and which one
* Next emits is its business, not something worth pinning a test to.
*
* The first cut hardcoded the slash and failed only on the homepage — the
* same gap as the doubled brand: it asserted the metadata object rather than
* what the page actually renders.
*/
function sameUrl(a: string | null, b: string): boolean {
const strip = (u: string) => u.replace(/\/+$/, '');
return strip(a ?? '') === strip(b);
}
for (const [path, expected] of CANONICAL_ROUTES) {
test(`${path} declares exactly one canonical, on the www host`, async ({ page }) => {
await page.goto(path);
const hrefs = await page.locator('link[rel="canonical"]').evaluateAll(
(els) => els.map((e) => e.getAttribute('href')));
expect(hrefs, `${path} should declare one canonical`).toHaveLength(1);
expect(hrefs[0]).toBe(expected);
expect(sameUrl(hrefs[0], expected),
`${path} canonical was ${hrefs[0]}, expected ${expected}`).toBe(true);
});
}
@@ -1663,7 +1679,8 @@ test('a filtered homepage still canonicalises to the bare root', async ({ page }
await page.goto('/?search=primary&phase=primary&sort=name&page=2');
const href = await page.locator('link[rel="canonical"]').first()
.getAttribute('href');
expect(href).toBe('https://www.schoolcompare.co.uk/');
expect(sameUrl(href, 'https://www.schoolcompare.co.uk/'),
`filtered homepage canonical was ${href}`).toBe(true);
});
test('a school page canonicalises to its own slug on the www host', async ({ page }) => {
@@ -1714,10 +1731,33 @@ test('staging answers noindex, and stays crawlable so the noindex is seen', asyn
// The other half, and the reason this is one test rather than two: a
// Disallow would stop Google fetching the page at all, so it would never
// see the noindex above. The two only work together.
//
// Scoped to the `*` group. The first cut matched `Disallow: /` anywhere in
// the file and tripped over the AI-crawler groups Cloudflare injects —
// ClaudeBot, GPTBot, Amazonbot and friends all carry a blanket disallow,
// deliberately, and none of them is Googlebot.
const robots = await (await page.request.get('/robots.txt')).text();
expect(robots).not.toMatch(/^\s*Disallow:\s*\/\s*$/mi);
expect(blocksEverything(robots, '*'),
'the * group must not disallow the whole site, or the noindex is never seen')
.toBe(false);
});
/** True when `agent`'s group in a robots.txt disallows the entire site. */
function blocksEverything(robots: string, agent: string): boolean {
let current: string | null = null;
let blocked = false;
for (const raw of robots.split('\n')) {
const line = raw.split('#')[0].trim();
if (!line) continue;
const [key, ...rest] = line.split(':');
const value = rest.join(':').trim();
const k = key.trim().toLowerCase();
if (k === 'user-agent') current = value;
else if (current === agent && k === 'disallow' && value === '/') blocked = true;
}
return blocked;
}
test('a school page on staging is noindexed too, not just the homepage', async ({ page }) => {
const list = await page.request.get('/api/schools?search=primary&per_page=1');
const [first] = (await list.json()).schools ?? [];
@@ -1726,3 +1766,216 @@ test('a school page on staging is noindexed too, not just the homepage', async (
const res = await page.request.get(`/school/${first.urn}-x`);
expect(res.headers()['x-robots-tag']).toContain('noindex');
});
/*
* W8 — the C1 pages must ship a description, and it must differentiate.
*
* Baseline was 0.43% CTR at position 6.1 on "compare school performance",
* against 9.16% for the brand query from the same neighbourhood. The SERP is
* owned by the DfE's own service, so a description that paraphrases it earns
* nothing. Google may rewrite a snippet, but it cannot use one we never sent.
*/
test('every C1 page ships a description, and none opens its title with the brand', async ({ page }) => {
for (const path of ['/', '/compare', '/rankings', '/admissions']) {
await page.goto(path);
const desc = await page.locator('meta[name="description"]').first()
.getAttribute('content');
expect(desc, `${path} must ship a description`).toBeTruthy();
expect(desc!.length, `${path} description too short to be worth reading`)
.toBeGreaterThan(100);
const title = await page.title();
expect(title.toLowerCase().startsWith('schoolcompare'),
`${path} spends its most valuable pixels on the brand`).toBe(false);
}
});
test('the homepage snippet names what gov.uk does not publish', async ({ page }) => {
await page.goto('/');
const desc = await page.locator('meta[name="description"]').first()
.getAttribute('content');
// Admissions distance is the one fact the DfE service has no equivalent for.
expect(desc).toMatch(/close you had to live|distance/i);
});
/*
* The location layer (spec 2026-08-21, W2).
*
* Location intent sat at position 49.5 with one click across the whole 16-month
* baseline — the site published no page about a place. These assert the four
* families render, stay in their own namespaces, and reach the sitemap.
*/
async function firstPlaceOfKind(page: Page, kind: string) {
const res = await page.request.get('/api/places');
expect(res.ok()).toBeTruthy();
const { places } = await res.json();
const hit = places.find((p: { kind: string }) => p.kind === kind);
expect(hit, `no ${kind} in the registry`).toBeTruthy();
return hit as { kind: string; slug: string; name: string; count: number };
}
for (const [kind, prefix, article] of [
['town', '/schools/', 'a'],
['authority', '/schools/authority/', 'an'],
['outcode', '/schools/near/', 'an'],
] as const) {
test(`${article} ${kind} page renders with its school count`, async ({ page }) => {
const place = await firstPlaceOfKind(page, kind);
await page.goto(`${prefix}${place.slug}`);
await expect(page.locator('h1')).toContainText(place.name, { ignoreCase: true });
await expect(page.locator('a[href^="/school/"]').first()).toBeVisible();
});
}
test('a town and an authority sharing a name are different pages', async ({ page }) => {
// 67 real collisions, and the authority is the larger set in only 43 — so
// one namespace would have published near-duplicates.
const { places } = await (await page.request.get('/api/places')).json();
const townSlugs = new Set(
places.filter((p: { kind: string }) => p.kind === 'town')
.map((p: { slug: string }) => p.slug));
const clash = places.find((p: { kind: string; slug: string }) =>
p.kind === 'authority' && townSlugs.has(p.slug));
test.skip(!clash, 'no town/authority name collision in this environment');
const townRes = await page.request.get(`/api/places/town/${clash.slug}`);
const laRes = await page.request.get(`/api/places/authority/${clash.slug}`);
expect(townRes.ok() && laRes.ok()).toBeTruthy();
const townUrns = (await townRes.json()).schools.map((s: { urn: number }) => s.urn).sort();
const laUrns = (await laRes.json()).schools.map((s: { urn: number }) => s.urn).sort();
expect(townUrns).not.toEqual(laUrns);
});
test('a place below the threshold has no page', async ({ page }) => {
// Crosby holds one school; publishing it would be a page with nothing to say.
const res = await page.request.get('/api/places/town/crosby');
expect(res.status()).toBe(404);
});
test('place pages declare a canonical and reach the sitemap', async ({ page }) => {
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
const canonical = await page.locator('link[rel="canonical"]').first()
.getAttribute('href');
expect(canonical).toBe(`https://www.schoolcompare.co.uk/schools/${place.slug}`);
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toContain(`/schools/${place.slug}`);
});
test('the place page ships ItemList structured data that parses', async ({ page }) => {
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
const raw = await page.locator('script[type="application/ld+json"]').first()
.textContent();
const parsed = JSON.parse(raw!);
const types = (parsed['@graph'] ?? []).map((n: { '@type': string }) => n['@type']);
expect(types).toContain('ItemList');
expect(types).toContain('BreadcrumbList');
});
test('a place page states the local average against England', async ({ page }) => {
// The one number a list cannot give, and the reason these pages are not
// a name dropped into a template.
const place = await firstPlaceOfKind(page, 'town');
await page.goto(`/schools/${place.slug}`);
await expect(page.getByTestId('local-vs-england')).toContainText(/across England/i);
});
test('a place page links its phase variants, and they resolve', async ({ page }) => {
// "primary schools in beccles" is the query shape the baseline showed. The
// first cut submitted only the bare place URL and linked nothing, leaving
// ~950 variant pages reachable by nothing at all.
const res = await page.request.get('/api/places');
const { places } = await res.json();
const town = places.find((p: { kind: string }) => p.kind === 'town');
expect(town).toBeTruthy();
const detail = await (await page.request.get(`/api/places/town/${town.slug}`)).json();
test.skip(!(detail.place.phases ?? []).length, 'no phase clears the threshold here');
await page.goto(`/schools/${town.slug}`);
const phase = detail.place.phases[0];
const link = page.locator(`a[href="/schools/${town.slug}/${phase}"]`).first();
await expect(link).toBeVisible();
await link.click();
await expect(page.locator('h1')).toContainText(new RegExp(`${phase} schools in`, 'i'));
});
test('phase variants are submitted in the places sitemap', async ({ page }) => {
const xml = await (await page.request.get('/sitemaps/places-1.xml')).text();
expect(xml).toMatch(/\/schools\/[a-z0-9-]+\/primary</);
});
test('no page title repeats the brand', async ({ page }) => {
// The root layout appends '| schoolcompare' to a plain-string title. Any
// route whose title already carries the brand must opt out with
// `absolute`, or it ships '... | schoolcompare | schoolcompare' — which is
// how ~2,600 place pages first went out.
const res = await page.request.get('/api/places');
const { places } = await res.json();
const town = places.find((p: { kind: string }) => p.kind === 'town');
for (const path of ['/', '/rankings', '/admissions', `/schools/${town.slug}`]) {
await page.goto(path);
const title = await page.title();
const brands = (title.match(/schoolcompare/gi) ?? []).length;
expect(brands, `${path} repeats the brand: ${title}`).toBeLessThanOrEqual(1);
}
});
test('a place straddling a boundary names every authority it sits in', async ({ page }) => {
// A quarter of outcodes and a third of towns cross an authority boundary —
// SW19 is mostly Merton but partly Wandsworth. Naming only the largest
// asserts something false about the place.
const { places } = await (await page.request.get('/api/places')).json();
const outcode = places.find((p: { kind: string }) => p.kind === 'outcode');
expect(outcode).toBeTruthy();
// Find any place the registry reports as straddling.
let straddling: { kind: string; slug: string } | null = null;
for (const p of places.filter((p: { kind: string }) => p.kind === 'outcode').slice(0, 40)) {
const d = await (await page.request.get(`/api/places/outcode/${p.slug}`)).json();
if ((d.place.authorities ?? []).length > 1) { straddling = p; break; }
}
test.skip(!straddling, 'no straddling outcode found in the sample');
const detail = await (await page.request.get(
`/api/places/outcode/${straddling!.slug}`)).json();
await page.goto(`/schools/near/${straddling!.slug}`);
for (const a of detail.place.authorities) {
await expect(page.locator(`a[href="/schools/authority/${a.slug}"]`).first())
.toBeVisible();
}
});
test('a place page lists its schools alphabetically', async ({ page }) => {
// Someone on a place page is usually looking for a school they can name,
// so the order should serve scanning for it. /rankings is where the
// league-table ordering lives.
const { places } = await (await page.request.get('/api/places')).json();
const town = places.find((p: { kind: string; count: number }) =>
p.kind === 'town' && p.count >= 5);
expect(town).toBeTruthy();
await page.goto(`/schools/${town.slug}`);
const names = await page.locator('a[href^="/school/"]').allTextContents();
expect(names.length).toBeGreaterThan(1);
const sorted = [...names].sort((a, b) =>
a.toLowerCase().localeCompare(b.toLowerCase()));
expect(names).toEqual(sorted);
});
test('the rankings page still orders by score, not name', async ({ page }) => {
// Alphabetical is a place-page decision, not a site-wide one.
const res = await page.request.get('/api/rankings?metric=rwm_expected_pct&phase=primary');
expect(res.ok()).toBeTruthy();
const scores = ((await res.json()).rankings ?? [])
.map((r: { rwm_expected_pct: number | null }) => r.rwm_expected_pct)
.filter((v: number | null) => v != null);
expect(scores).toEqual([...scores].sort((a: number, b: number) => b - a));
});
+77
View File
@@ -51,3 +51,80 @@ describe('/compare indexability', () => {
.toBe('https://www.schoolcompare.co.uk/compare');
});
});
/*
* W8 — snippet copy for the C1 cluster.
*
* The baseline (GSC, 16 months to 2026-08-20) showed these pages ranking on
* page one and converting at a tenth of the normal rate: "compare school
* performance" at position 6.1 with 0.43% CTR, against 9.16% for the brand
* query from the same neighbourhood. The SERP is dominated by the DfE's own
* "Compare school performance" service, so the job of this copy is to say
* what that service does not offer, without losing intent match on the title.
*
* These tests guard the mechanics that make a snippet work — length, intent
* keyword, differentiator, no brand-first — not the exact wording, which
* should stay free to iterate.
*/
// Google truncates titles near 60 characters and descriptions near 155.
const TITLE_MAX = 60;
const DESC_MIN = 110;
const DESC_MAX = 155;
type Meta = { title?: unknown; description?: unknown };
const titleOf = (m: Meta): string => {
const t = m.title as string | { absolute?: string } | undefined;
return typeof t === 'string' ? t : (t?.absolute ?? '');
};
describe('C1 snippet copy', () => {
const pages: Array<[string, Meta, RegExp]> = [
['home', homeMetadata as Meta, /compare schools/i],
['rankings', rankingsMetadata as Meta, /league table/i],
['admissions', admissionsMetadata as Meta, /admission/i],
];
for (const [name, meta, intent] of pages) {
it(`${name}: title carries the search intent and fits the SERP`, () => {
const t = titleOf(meta);
expect(t).toMatch(intent);
expect(t.length).toBeLessThanOrEqual(TITLE_MAX);
});
it(`${name}: title does not open with the brand`, () => {
// The measured 0.43% CTR came from a brand-first title. The most
// valuable pixels go to the thing the searcher typed.
expect(titleOf(meta).toLowerCase().startsWith('schoolcompare')).toBe(false);
});
it(`${name}: description is long enough to be worth reading, short enough to survive`, () => {
const d = meta.description as string;
expect(d.length).toBeGreaterThanOrEqual(DESC_MIN);
expect(d.length).toBeLessThanOrEqual(DESC_MAX);
});
}
it('the homepage description names what gov.uk does not publish', () => {
// Admissions distance is the one fact the DfE service has no equivalent
// for. If it ever leaves this description, the snippet is competing with
// gov.uk on gov.uk's own ground.
expect(homeMetadata.description).toMatch(/close you had to live|distance/i);
});
it('/compare targets the tool phrasing rather than repeating the homepage', () => {
// Two pages chasing one phrase is how a site competes with itself.
return compareMetadata({ searchParams: Promise.resolve({}) }).then((m) => {
expect(m.title).toMatch(/comparison tool/i);
expect(m.title).not.toBe(titleOf(homeMetadata as Meta));
});
});
it('no C1 page claims a school count that will drift', () => {
// The corpus moves with every data refresh; this repo has already shipped
// one copy bug of that kind ("three schools" against MAX_SCHOOLS = 5).
for (const [, meta] of pages) {
expect(meta.description as string).not.toMatch(/\b\d{2},\d{3}\b|\b\d{2},000\b/);
}
});
});
@@ -0,0 +1,42 @@
import { generateMetadata as placeMeta } from '@/app/schools/[place]/page';
jest.mock('@/lib/places', () => ({
...jest.requireActual('@/lib/places'),
fetchPlace: jest.fn(async (kind: string, slug: string) =>
slug === 'atlantis' ? null : ({
place: { kind, slug, name: 'Brentwood', count: 29,
parent_authority: 'Essex' },
schools: [], averages: { rwm_expected_pct: 63, attainment_8_score: null },
})),
fetchPlaces: jest.fn(async () => []),
}));
describe('place page metadata', () => {
it('titles the page the way the place is searched', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'brentwood' }) });
expect((m.title as { absolute: string }).absolute).toMatch(/schools in brentwood/i);
});
it('canonicalises to its own path on the www host', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'brentwood' }) });
expect(m.alternates?.canonical)
.toBe('https://www.schoolcompare.co.uk/schools/brentwood');
});
it('opts out of the layout template, which would double the brand', () => {
// The root layout appends '| schoolcompare' to a plain string title, and
// these titles already carry it — every place page shipped reading
// '... | schoolcompare | schoolcompare' until this was made absolute.
return placeMeta({ params: Promise.resolve({ place: 'brentwood' }) })
.then((m) => {
expect(typeof m.title).toBe('object');
expect((m.title as { absolute: string }).absolute)
.not.toMatch(/schoolcompare.*schoolcompare/);
});
});
it('an unknown place gets a not-found title rather than inventing one', async () => {
const m = await placeMeta({ params: Promise.resolve({ place: 'atlantis' }) });
expect(m.title).toMatch(/not found/i);
});
});
@@ -0,0 +1,278 @@
import { render, screen } from '@testing-library/react';
import { PlaceView } from '@/components/places/PlaceView';
import type { PlaceDetail } from '@/lib/places';
const detail: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 29,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', rwm_expected_pct: 82,
ofsted_grade: 1, phase: 'Primary' } as never,
{ urn: 2, school_name: 'Beta Primary', rwm_expected_pct: 44,
ofsted_grade: 3, phase: 'Primary' } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
describe('PlaceView', () => {
it('leads with an H1 that matches how the place is searched', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByRole('heading', { level: 1 }))
.toHaveTextContent(/primary schools in brentwood/i);
});
it('states the count so the page says something before the table', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByText(/29 schools/i)).toBeInTheDocument();
});
it('compares the local average against England, which a list cannot', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByTestId('local-vs-england')).toHaveTextContent('63');
expect(screen.getByTestId('local-vs-england')).toHaveTextContent('61');
});
it('links every school in scope, which is what de-orphans them', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getAllByRole('link', { name: /Primary$/ })).toHaveLength(2);
});
it('links to the parent authority so the place sits in a hierarchy', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByRole('link', { name: /Essex/i }))
.toHaveAttribute('href', '/schools/authority/essex');
});
it('shows the Ofsted distribution, not just a count of Outstanding', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.getByTestId('ofsted-distribution')).toBeInTheDocument();
});
it('links to neighbouring places so the page is not a dead end', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[{ kind: 'town', slug: 'romford', name: 'Romford', count: 40 }]} />);
expect(screen.getByRole('link', { name: /Romford/ }))
.toHaveAttribute('href', '/schools/romford');
});
it('says nothing about an average it does not have', () => {
render(<PlaceView detail={{ ...detail, averages:
{ rwm_expected_pct: null, attainment_8_score: null } }}
phase="primary" englandAverage={61} neighbours={[]} />);
expect(screen.queryByTestId('local-vs-england')).not.toBeInTheDocument();
});
});
describe('PlaceView structured data', () => {
function jsonLd() {
const { container } = render(<PlaceView detail={detail} phase="primary"
englandAverage={61} neighbours={[]} />);
const el = container.querySelector('script[type="application/ld+json"]');
return JSON.parse(el!.textContent!);
}
it('declares the page as a ranked list, not prose', () => {
const types = jsonLd()['@graph'].map((n: { '@type': string }) => n['@type']);
expect(types).toContain('ItemList');
expect(types).toContain('BreadcrumbList');
});
it('gives every listed school an absolute URL on the canonical host', () => {
const list = jsonLd()['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListElement).toHaveLength(2);
for (const item of list.itemListElement) {
expect(item.url).toMatch(/^https:\/\/www\.schoolcompare\.co\.uk\/school\//);
}
});
});
describe('PlaceView phase variants', () => {
it('links the phase variants that exist', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: /Primary schools in Brentwood/i }))
.toHaveAttribute('href', '/schools/brentwood/primary');
});
it('links no variant for a phase below its own threshold', () => {
render(<PlaceView detail={detail} englandAverage={61} neighbours={[]} />);
expect(screen.queryByRole('link', { name: /Secondary schools in Brentwood/i }))
.not.toBeInTheDocument();
});
it('does not link sideways from a variant page to itself', () => {
render(<PlaceView detail={detail} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.queryByRole('link', { name: /Primary schools in Brentwood/i }))
.not.toBeInTheDocument();
});
});
describe('PlaceView presentation', () => {
// /schools/brentwood shipped with 8 of 27 rows blank: an unphased page shows
// one primary-only measure for a list that also holds secondaries.
const mixed: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 4,
parent_authority: 'Essex', phases: ['primary', 'secondary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta High', phase: 'Secondary',
rwm_expected_pct: null, attainment_8_score: 47 } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: 45 },
};
it('gives each phase its own table rather than one column of blanks', () => {
render(<PlaceView detail={mixed} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('heading', { name: /^Primary schools/ })).toBeInTheDocument();
expect(screen.getByRole('heading', { name: /^Secondary schools/ })).toBeInTheDocument();
expect(screen.getByText('82%')).toBeInTheDocument();
expect(screen.getByText('47')).toBeInTheDocument();
});
it('names the measure in plain words, not jargon', () => {
// The first cut said "RWM expected", which appears nowhere else on the site.
render(<PlaceView detail={mixed} englandAverage={61} neighbours={[]} />);
expect(screen.getByText('Reading, writing & maths')).toBeInTheDocument();
expect(screen.getByText('Attainment 8')).toBeInTheDocument();
expect(screen.queryByText(/RWM expected/i)).not.toBeInTheDocument();
});
it('says a missing result is unpublished rather than showing a bare dash', () => {
const noResult: PlaceDetail = {
...mixed,
schools: [{ urn: 3, school_name: 'New Primary', phase: 'Primary',
rwm_expected_pct: null, attainment_8_score: null } as never],
};
render(<PlaceView detail={noResult} englandAverage={61} neighbours={[]} />);
expect(screen.getByText('Not published')).toBeInTheDocument();
});
it('styles school links to the site convention rather than browser default', () => {
const { container } = render(<PlaceView detail={mixed} englandAverage={61}
neighbours={[]} />);
const link = container.querySelector('a[href^="/school/"]');
expect(link?.className).toBeTruthy();
});
it('a phased page shows one table and no phase headings', () => {
render(<PlaceView detail={mixed} phase="primary" englandAverage={61}
neighbours={[]} />);
expect(screen.queryByRole('heading', { name: /^Secondary schools/ }))
.not.toBeInTheDocument();
});
});
describe('PlaceView table alignment', () => {
const aligned: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('aligns the measure heading and its values with the same class', () => {
// They were aligned by two different selectors whose specificity did not
// match: `.table th:last-child` (0,2,1) won and went right, while `.num`
// (0,1,0) lost to `.table td` (0,1,1) and stayed left. Sharing one class
// is what makes them impossible to drift apart.
const { container } = render(<PlaceView detail={aligned} englandAverage={61}
neighbours={[]} />);
const th = container.querySelectorAll('th')[1];
const td = container.querySelectorAll('tbody td')[1];
expect(th.className).toBeTruthy();
expect(td.className).toBe(th.className);
});
it('leaves the school-name column unclassed so it takes the spare width', () => {
const { container } = render(<PlaceView detail={aligned} englandAverage={61}
neighbours={[]} />);
expect(container.querySelectorAll('th')[0].className).toBe('');
});
});
describe('PlaceView authorities', () => {
const straddling: PlaceDetail = {
place: { kind: 'outcode', slug: 'sw19', name: 'SW19', count: 33,
parent_authority: 'Merton', phases: ['primary'],
authorities: [
{ name: 'Merton', slug: 'merton', count: 26 },
{ name: 'Wandsworth', slug: 'wandsworth', count: 7 },
] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 82, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 63, attainment_8_score: null },
};
it('names every authority the place straddles, not just the largest', () => {
// SW19 is mostly Merton but partly Wandsworth. Naming one asserts
// something false about a quarter of outcodes.
render(<PlaceView detail={straddling} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' }))
.toHaveAttribute('href', '/schools/authority/merton');
expect(screen.getByRole('link', { name: 'Wandsworth' }))
.toHaveAttribute('href', '/schools/authority/wandsworth');
});
it('joins them readably rather than as a bare list', () => {
// Asserted on the summary line's whole text: a loose /and/ matcher also
// hits "Wandsworth".
const { container } = render(<PlaceView detail={straddling}
englandAverage={61} neighbours={[]} />);
const summary = container.querySelector('header p');
expect(summary?.textContent).toContain('Merton and Wandsworth');
});
it('falls back to the single parent when the field is absent', () => {
// A cached API response predating the authorities field must not blank
// the line entirely.
const legacy = { ...straddling,
place: { ...straddling.place, authorities: undefined } };
render(<PlaceView detail={legacy} englandAverage={61} neighbours={[]} />);
expect(screen.getByRole('link', { name: 'Merton' })).toBeInTheDocument();
});
});
describe('PlaceView list ordering', () => {
const detail3: PlaceDetail = {
place: { kind: 'town', slug: 'brentwood', name: 'Brentwood', count: 2,
parent_authority: 'Essex', phases: ['primary'] },
schools: [
{ urn: 1, school_name: 'Alpha Primary', phase: 'Primary',
rwm_expected_pct: 40, attainment_8_score: null } as never,
{ urn: 2, school_name: 'Beta Primary', phase: 'Primary',
rwm_expected_pct: 90, attainment_8_score: null } as never,
],
averages: { rwm_expected_pct: 65, attainment_8_score: null },
};
it('renders schools in the order the API sent them, not by score', () => {
// The API sorts alphabetically now; the component must not re-sort.
render(<PlaceView detail={detail3} englandAverage={61} neighbours={[]} />);
const links = screen.getAllByRole('link', { name: /Primary$/ });
expect(links.map((l) => l.textContent))
.toEqual(['Alpha Primary', 'Beta Primary']);
});
it('declares the list as ascending rather than implying a ranking', () => {
// An ItemList carrying `position` reads as a ranking unless it says
// otherwise, and the table is A-Z.
const { container } = render(<PlaceView detail={detail3} englandAverage={61}
neighbours={[]} />);
const ld = JSON.parse(
container.querySelector('script[type="application/ld+json"]')!.textContent!);
const list = ld['@graph'].find((n: { '@type': string }) => n['@type'] === 'ItemList');
expect(list.itemListOrder).toBe('https://schema.org/ItemListOrderAscending');
});
});
+4 -2
View File
@@ -5,9 +5,11 @@ import { AdmissionsView } from '@/components/AdmissionsView';
export const dynamic = 'force-static';
export const metadata: Metadata = {
title: 'School Admissions Guide',
// Deadlines and offer days are what gets searched, and what this page is
// genuinely best at — the countdowns are live.
title: { absolute: 'School Admissions Deadlines & Offer Days | schoolcompare' },
description:
'Understand the Primary and Secondary school admissions process in England, with live countdowns to every key deadline and National Offer Day.',
'Every key date for primary and secondary school admissions in England, with live countdowns to the application deadline and National Offer Day.',
alternates: { canonical: absoluteUrl('/admissions') },
};
+5 -2
View File
@@ -30,9 +30,12 @@ export async function generateMetadata(
const { urns } = await searchParams;
const base: Metadata = {
title: 'Compare Schools',
// Deliberately not the homepage's phrase. Two pages chasing "compare
// schools" is how a site competes with itself; this one takes the tool
// phrasing instead.
title: 'School Comparison Tool — Up to Five at Once | schoolcompare',
description:
'Compare schools in England side by side — Ofsted inspections, KS2 and GCSE results against the England average, admissions odds and school community.',
'Put up to five English schools in one table: SATs and GCSE results against the England average, Ofsted grades, and the distance places were offered.',
keywords:
'school comparison, compare schools, Ofsted comparison, school admissions, KS2 comparison, primary school performance',
alternates: { canonical: absoluteUrl('/compare') },
+9 -6
View File
@@ -48,10 +48,11 @@ export const metadata: Metadata = {
statusBarStyle: 'default',
},
title: {
default: 'schoolcompare | Compare School Performance',
default: 'Compare Schools Side by Side | schoolcompare',
template: '%s | schoolcompare',
},
description: 'Compare primary and secondary school SATs and GCSE performance across England',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place. Free, no sign-up.',
keywords: 'school comparison, KS2 results, KS4 results, primary school, secondary school, England schools, SATs results, GCSE results',
authors: [{ name: 'schoolcompare' }],
manifest: '/manifest.json',
@@ -61,16 +62,18 @@ export const metadata: Metadata = {
metadataBase: new URL(SITE_URL),
openGraph: {
type: 'website',
title: 'schoolcompare | Compare School Performance',
description: 'Compare primary and secondary school SATs and GCSE performance across England',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
url: SITE_URL,
siteName: 'schoolcompare',
},
twitter: {
// summary_large_image now that there is an image worth showing.
card: 'summary_large_image',
title: 'schoolcompare | Compare School Performance',
description: 'Compare primary and secondary school SATs and GCSE performance across England',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
},
};
+15 -2
View File
@@ -34,8 +34,21 @@ interface HomePageProps {
* saying the brand twice.
*/
export const metadata: Metadata = {
title: { absolute: 'schoolcompare | Compare every school in England' },
description: 'Search and compare school performance across England',
/*
* Intent in the title, differentiator in the description.
*
* These queries are owned by the DfE's own "Compare school performance"
* service, and the old title — brand first, then a near-paraphrase of that
* service's name — gave a searcher no reason to pick us over it. It drew
* 0.43% CTR at position 6.1 while the brand query drew 9.16% from the same
* neighbourhood, so the ranking was never the problem.
*
* The title now matches what people type. The description carries the one
* fact gov.uk does not publish: how close you had to live to get a place.
*/
title: { absolute: 'Compare Schools Side by Side | schoolcompare' },
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place. Free, no sign-up.',
// This page reads eleven search params. They filter a result set; they do
// not make a new document. Collapsing every combination onto "/" stops the
// homepage competing with itself for its own head terms.
+5 -2
View File
@@ -18,8 +18,11 @@ interface RankingsPageProps {
}
export const metadata: Metadata = {
title: 'School Rankings',
description: 'Top-ranked schools by SATs and GCSE performance across England',
// 'School Rankings' matched nothing anyone types. League tables is the
// phrase parents actually search, and it spikes each results day.
title: { absolute: 'Primary & Secondary School League Tables | schoolcompare' },
description:
'Rank English schools by SATs results, GCSEs, Progress 8 or Attainment 8, and filter by local authority or year. Built from the DfE’s own figures.',
keywords: 'school rankings, top schools, best schools, KS2 rankings, KS4 rankings, school league tables',
// Param forms (?metric=&local_authority=&year=&phase=) collapse here for
// now. W3 replaces them with real indexable paths.
@@ -0,0 +1,63 @@
/**
* Phase variants of a place page.
*
* Phase is part of the query — "primary schools in beccles", "secondary
* schools in brentwood" — not a filter applied afterwards, so each gets its
* own indexable path. A place with no schools of the phase has no page: the
* per-phase threshold, not an error.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ place: string; phase: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
const PHASES = ['primary', 'secondary'] as const;
type Phase = (typeof PHASES)[number];
const isPhase = (v: string): v is Phase => (PHASES as readonly string[]).includes(v);
async function resolve(slug: string, phase: Phase) {
return (await fetchPlace('town', slug, phase))
?? (await fetchPlace('locality', slug, phase));
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { place: slug, phase } = await params;
if (!isPhase(phase)) return { title: 'Place Not Found' };
const detail = await resolve(slug, phase);
if (!detail || detail.schools.length === 0) return { title: 'Place Not Found' };
const word = phase === 'secondary' ? 'Secondary' : 'Primary';
const { name } = detail.place;
return {
// Not "Ranked": the table is alphabetical, so the word would be a claim
// the page does not keep.
title: { absolute: `${word} Schools in ${name} | schoolcompare` },
description:
`Every ${phase} school in ${name}, with results, Ofsted grades and the local `
+ `average against England.`,
alternates: { canonical: absoluteUrl(`/schools/${slug}/${phase}`) },
};
}
export default async function PlacePhasePage({ params }: Props) {
const { place: slug, phase } = await params;
if (!isPhase(phase)) notFound();
const detail = await resolve(slug, phase);
if (!detail || detail.schools.length === 0) notFound();
const national = await fetchNationalAverages().catch(() => null);
const englandAverage = phase === 'secondary'
? national?.secondary?.attainment_8_score ?? null
: national?.primary?.rwm_expected_pct ?? null;
return <PlaceView detail={detail} phase={phase}
englandAverage={englandAverage} neighbours={[]} />;
}
+94
View File
@@ -0,0 +1,94 @@
/**
* Town and locality pages.
*
* A place below the five-school threshold is not in the registry, so
* fetchPlace returns null and the request 404s rather than rendering a page
* with nothing to say.
*/
import { notFound, redirect } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces, authoritySlug } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ place: string }> }
// ISR: place aggregates change only when the pipeline runs.
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ place: string }>> {
// Off by default: ~2,000 place routes cannot be built in CI on every deploy.
// Matches the PRERENDER_SCHOOLS gate on the school route.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'town' || p.kind === 'locality')
.map((p) => ({ place: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
async function resolve(slug: string) {
return (await fetchPlace('town', slug)) ?? (await fetchPlace('locality', slug));
}
/** Other towns in the same authority — the cheapest honest definition of
* "nearby", and enough to stop each place page being a dead end. */
async function neighboursOf(detail: { place: { slug: string; parent_authority: string | null } }) {
if (!detail.place.parent_authority) return [];
const all = await fetchPlaces();
return all
.filter((p) => p.kind === 'town' && p.slug !== detail.place.slug)
.slice(0, 12);
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { place: slug } = await params;
const detail = await resolve(slug);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
// absolute: the root layout's template appends '| schoolcompare' to a
// plain string, and this title already carries it. Without this every
// place title read '... | schoolcompare | schoolcompare'.
title: { absolute: `Schools in ${name} — Compare ${count} Schools | schoolcompare` },
description:
`Every school in ${name}, with SATs and GCSE results, Ofsted grades, the local `
+ `average against England, and how close you had to live to get a place.`,
alternates: { canonical: absoluteUrl(`/schools/${slug}`) },
};
}
export default async function PlacePage({ params }: Props) {
const { place: slug } = await params;
const detail = await resolve(slug);
if (!detail) notFound();
// Global constraint: no page without a local average. A place with too few
// schools carrying results has nothing to say that a list does not, so it
// defers to its authority rather than publishing a thin page.
if (detail.averages.rwm_expected_pct == null
&& detail.averages.attainment_8_score == null) {
if (detail.place.parent_authority) {
redirect(`/schools/authority/${authoritySlug(detail.place.parent_authority)}`);
}
notFound();
}
const national = await fetchNationalAverages().catch(() => null);
// NationalAverages is nested by phase — { primary: {...}, secondary: {...} }
// — not flat. Reading it flat silently yields undefined and the page renders
// with no comparison, which is the one thing that makes it not a list.
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={await neighboursOf(detail)}
/>
);
}
@@ -0,0 +1,67 @@
/**
* Local authority pages.
*
* A separate namespace from /schools/[place] because 67 town names collide
* with an authority name and neither set contains the other — Bedford the
* town holds 104 schools, Bedford the authority 86, because postal towns
* cross authority boundaries. The title says "Local Authority" so a reader
* landing on both knows which set each covers.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ la: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ la: string }>> {
// Gated like every other prerender in this app. There are only ~154
// authorities, but "few enough to always build" still means the API must be
// reachable at build time, and in CI it is not — the build fails with
// ECONNREFUSED rather than degrading. The catch is the same fallback the
// school route uses.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'authority')
.map((p) => ({ la: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { la } = await params;
const detail = await fetchPlace('authority', la);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
title: { absolute: `Schools in ${name} — Local Authority | schoolcompare` },
description:
`All ${count} schools in the ${name} local authority, with SATs and GCSE results, `
+ `Ofsted grades and the authority average against England.`,
alternates: { canonical: absoluteUrl(`/schools/authority/${la}`) },
};
}
export default async function AuthorityPage({ params }: Props) {
const { la } = await params;
const detail = await fetchPlace('authority', la);
if (!detail) notFound();
const national = await fetchNationalAverages().catch(() => null);
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={[]}
/>
);
}
@@ -0,0 +1,62 @@
/**
* Postcode district pages.
*
* No phase variants: nobody searches "primary schools in SW11", so the
* variants would be pages without demand. These exist to catch
* "schools near <postcode>" and to give London districts a geographic page
* where the GIAS town field cannot.
*/
import { notFound } from 'next/navigation';
import type { Metadata } from 'next';
import { fetchPlace, fetchPlaces } from '@/lib/places';
import { fetchNationalAverages } from '@/lib/api';
import { PlaceView } from '@/components/places/PlaceView';
import { absoluteUrl } from '@/lib/site';
interface Props { params: Promise<{ outcode: string }> }
export const revalidate = 604800;
export const dynamicParams = true;
export async function generateStaticParams(): Promise<Array<{ outcode: string }>> {
// 1,760 of these; same CI budget argument as the town routes.
if (process.env.PRERENDER_PLACES !== '1') return [];
try {
return (await fetchPlaces())
.filter((p) => p.kind === 'outcode')
.map((p) => ({ outcode: p.slug }));
} catch (error) {
console.warn('generateStaticParams: API unreachable, falling back to on-demand ISR.', error);
return [];
}
}
export async function generateMetadata({ params }: Props): Promise<Metadata> {
const { outcode } = await params;
const detail = await fetchPlace('outcode', outcode);
if (!detail) return { title: 'Place Not Found' };
const { name, count } = detail.place;
return {
title: { absolute: `Schools near ${name} | schoolcompare` },
description:
`${count} schools in the ${name} postcode district, with results, Ofsted grades `
+ `and how close you had to live to get a place.`,
alternates: { canonical: absoluteUrl(`/schools/near/${outcode}`) },
};
}
export default async function OutcodePage({ params }: Props) {
const { outcode } = await params;
const detail = await fetchPlace('outcode', outcode);
if (!detail) notFound();
const national = await fetchNationalAverages().catch(() => null);
return (
<PlaceView
detail={detail}
englandAverage={national?.primary?.rwm_expected_pct ?? null}
neighbours={[]}
/>
);
}
+1 -1
View File
@@ -9,7 +9,7 @@ export const runtime = 'nodejs';
* validated here rather than passed through, so this route cannot be used to
* reach arbitrary backend paths.
*/
const CHILD = /^(static|schools-\d+)\.xml$/;
const CHILD = /^(static|schools-\d+|places-\d+|outcodes-\d+)\.xml$/;
export async function GET(
_request: Request,
@@ -0,0 +1,210 @@
/* Tokens only — see globals.css. Follows RankingsView's conventions, and in
particular its link treatment: table links take --text-primary with no
underline and a brand-coloured hover, not the browser default. The first
cut used bare <Link> with no class at all, which rendered as default blue
underlined links and read as unstyled beside the rest of the site. */
.container {
width: 100%;
min-width: 0;
}
.header {
margin-bottom: 1.5rem;
}
.header h1 {
font-size: 2.25rem;
font-weight: 700;
color: var(--text-primary);
margin-bottom: 0.5rem;
font-family: var(--font-display);
text-wrap: balance;
}
.summary {
font-size: 1rem;
color: var(--text-secondary);
margin: 0;
line-height: 1.6;
}
/* Links in running copy: brand colour, underline on hover only. */
.inlineLink {
color: var(--brand);
text-decoration: none;
transition: color 0.2s ease;
}
.inlineLink:hover {
color: var(--brand-strong);
text-decoration: underline;
}
/* Phase variants are separate indexable pages, so the bare place page has to
link them — a sitemap entry alone leaves them with no internal path in. */
.phaseLinks {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 0.75rem;
margin: 0 0 1.25rem;
}
.phaseLink {
display: inline-block;
padding: 0.4rem 0.875rem;
border: 1px solid var(--border-strong);
border-radius: 999px;
font-size: 0.875rem;
font-weight: 500;
color: var(--text-primary);
text-decoration: none;
transition: border-color 0.2s ease, color 0.2s ease;
}
.phaseLink:hover {
border-color: var(--brand);
color: var(--brand-strong);
}
/* The one number a list cannot give you, so it gets its own band. */
.compare {
background: var(--bg-secondary);
border: 1px solid var(--border);
border-radius: 8px;
padding: 0.875rem 1.125rem;
margin: 0 0 1.5rem;
color: var(--text-primary);
font-size: 1rem;
}
.ofsted {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1.25rem;
list-style: none;
padding: 0;
margin: 0 0 1.5rem;
font-size: 0.9375rem;
color: var(--text-secondary);
}
.group {
margin-bottom: 2rem;
}
.groupHeading {
display: flex;
align-items: baseline;
gap: 0.625rem;
font-size: 1.25rem;
font-weight: 600;
color: var(--text-primary);
font-family: var(--font-display);
margin: 0 0 0.75rem;
}
.groupCount {
font-size: 0.8125rem;
font-weight: 500;
color: var(--text-secondary);
background: var(--bg-secondary);
border-radius: 999px;
padding: 0.125rem 0.5rem;
}
/* Wide content scrolls in its own container so the page body never does. */
.tableWrap {
overflow-x: auto;
border: 1px solid var(--border);
border-radius: 8px;
background: var(--bg-card);
}
.table {
width: 100%;
border-collapse: collapse;
font-size: 0.9375rem;
}
.table th,
.table td {
padding: 0.75rem 1rem;
text-align: left;
border-bottom: 1px solid var(--border);
}
.table th {
background: var(--bg-secondary);
color: var(--text-secondary);
font-weight: 600;
font-size: 0.8125rem;
}
.table tbody tr:last-child td {
border-bottom: none;
}
/*
* Header and value share one class and one rule, so they cannot drift apart.
*
* The first cut aligned them with two different selectors: `.table th:last-child`
* at (0,2,1) beat the element rule and went right, while `.num` at (0,1,0) lost
* to `.table td` at (0,1,1) and stayed left. The heading and its numbers sat on
* opposite edges of the column.
*
* width:1% with nowrap makes the measure column hug its content so the school
* name takes the remaining width — without it the two columns split evenly and
* the gap between heading and value reads as misalignment on a wide screen.
*/
.table th.num,
.table td.num {
text-align: right;
font-variant-numeric: tabular-nums;
width: 1%;
white-space: nowrap;
}
/* The measure is spelled out; the tooltip carries the definition. */
.metricHead {
text-decoration: none;
cursor: help;
border-bottom: 1px dotted var(--border-strong);
}
/* Table links: site convention is body colour, brand on hover. */
.schoolLink {
color: var(--text-primary);
text-decoration: none;
transition: color 0.2s ease;
}
.schoolLink:hover {
color: var(--brand-strong);
}
/* "Not published" is a fact about the school, not an error. */
.noData {
color: var(--text-muted);
font-size: 0.8125rem;
}
.neighbours {
margin-top: 2rem;
}
.neighbours h2 {
font-size: 1.125rem;
font-weight: 600;
color: var(--text-primary);
margin: 0 0 0.75rem;
font-family: var(--font-display);
}
.neighbours ul {
display: flex;
flex-wrap: wrap;
gap: 0.5rem 1rem;
list-style: none;
padding: 0;
margin: 0;
}
+247
View File
@@ -0,0 +1,247 @@
/**
* One place page, shared by all four families.
*
* They differ in what fills the registry, not in what the page shows, so a
* second component would be a second place to forget the same change.
*
* The local-versus-England comparison is the reason this page is not a list:
* it is the one number a parent cannot get by reading the schools one by one,
* and it is what keeps the page from reading as a name dropped into a
* template.
*/
import Link from 'next/link';
import type { PlaceDetail, PlaceSummary } from '@/lib/places';
import { placeUrl, authoritySlug } from '@/lib/places';
import type { School } from '@/lib/types';
import { schoolUrl } from '@/lib/utils';
import { absoluteUrl } from '@/lib/site';
import styles from './PlaceView.module.css';
interface Props {
detail: PlaceDetail;
phase?: 'primary' | 'secondary';
englandAverage: number | null;
/** Nearby places, so the page links onward instead of dead-ending. */
neighbours: PlaceSummary[];
}
// Ofsted grades in the order they are reported.
const OFSTED_LABELS: Array<[number, string]> = [
[1, 'Outstanding'], [2, 'Good'],
[3, 'Requires improvement'], [4, 'Inadequate'],
];
/**
* Column headings, taken from the site's own metric dictionary rather than
* invented here — see METRIC_DEFINITIONS in backend/schemas.py, surfaced at
* /api/metrics. The first cut said "RWM expected", which is jargon that
* appears nowhere else on the site.
*/
const METRICS = {
primary: {
key: 'rwm_expected_pct' as const,
heading: 'Reading, writing & maths',
hint: '% meeting the expected standard in reading, writing and maths',
unit: '%',
},
secondary: {
key: 'attainment_8_score' as const,
heading: 'Attainment 8',
hint: "Average grade across a pupil's best 8 GCSEs, including English and maths",
unit: '',
},
};
type PhaseKey = keyof typeof METRICS;
/** All-through schools sit in both phases, matching the search filters. */
function isPhase(school: School, phase: PhaseKey): boolean {
const p = (school.phase ?? '').toLowerCase();
if (p === 'all-through') return true;
return phase === 'secondary'
? p.includes('secondary') || p === '16 plus'
: p.includes('primary') || p.includes('middle');
}
function SchoolTable({ schools, phase }: { schools: School[]; phase: PhaseKey }) {
const metric = METRICS[phase];
return (
<div className={styles.tableWrap}>
<table className={styles.table}>
<thead>
<tr>
<th scope="col">School</th>
{/* Same class as the value cell below: one rule aligns both, so
they cannot drift apart. */}
<th scope="col" className={styles.num}>
<abbr className={styles.metricHead} title={metric.hint}>
{metric.heading}
</abbr>
</th>
</tr>
</thead>
<tbody>
{schools.map((s) => {
const value = s[metric.key];
return (
<tr key={s.urn}>
<td>
<Link href={schoolUrl(s.urn, s.school_name)} className={styles.schoolLink}>
{s.school_name}
</Link>
</td>
<td className={styles.num}>
{value == null
? <span className={styles.noData}>Not published</span>
: `${Math.round(Number(value))}${metric.unit}`}
</td>
</tr>
);
})}
</tbody>
</table>
</div>
);
}
export function PlaceView({ detail, phase, englandAverage, neighbours }: Props) {
const { place, schools, averages } = detail;
// Fall back to the single parent when the API predates the authorities
// field, so a stale cache never blanks the line entirely.
const authorities = place.authorities?.length
? place.authorities
: place.parent_authority
? [{ name: place.parent_authority, slug: authoritySlug(place.parent_authority), count: 0 }]
: [];
const local = averages[METRICS[phase ?? 'primary'].key];
const phaseWord = phase === 'secondary' ? 'Secondary schools'
: phase === 'primary' ? 'Primary schools' : 'Schools';
const graded = OFSTED_LABELS
.map(([grade, label]) => [label, schools.filter((s) => s.ofsted_grade === grade).length] as const)
.filter(([, n]) => n > 0);
/*
* An unphased page holds both primaries and secondaries, and they are
* scored on different measures — a percentage and a 0-90 score. Showing one
* column for both left 30% of rows blank on /schools/brentwood and put two
* incomparable scales in one column when it did not.
*
* So the phases get a table each. A blank cell inside one now means the
* school genuinely has no published result, which is worth saying.
*/
const groups: Array<[PhaseKey, School[]]> = phase
? [[phase, schools]]
: (['primary', 'secondary'] as PhaseKey[])
.map((p) => [p, schools.filter((s) => isPhase(s, p))] as [PhaseKey, School[]])
.filter(([, list]) => list.length > 0);
const jsonLd = {
'@context': 'https://schema.org',
'@graph': [
{
'@type': 'ItemList',
name: `${phaseWord} in ${place.name}`,
numberOfItems: schools.length,
// Alphabetical, and said so. Without this an ItemList carrying
// `position` reads as a ranking, which would be a claim the page
// stopped making when the table became A-Z.
itemListOrder: 'https://schema.org/ItemListOrderAscending',
itemListElement: schools.slice(0, 20).map((s, i) => ({
'@type': 'ListItem',
position: i + 1,
url: absoluteUrl(schoolUrl(s.urn, s.school_name)),
name: s.school_name,
})),
},
{
'@type': 'BreadcrumbList',
itemListElement: [
{ '@type': 'ListItem', position: 1, name: 'Schools', item: absoluteUrl('/') },
{ '@type': 'ListItem', position: 2, name: place.name },
],
},
],
};
return (
<div className={styles.container}>
<script
type="application/ld+json"
dangerouslySetInnerHTML={{ __html: JSON.stringify(jsonLd) }}
/>
<header className={styles.header}>
<h1>{phaseWord} in {place.name}</h1>
<p className={styles.summary}>
{place.count} schools
{authorities.length > 0 && (
<>
{' · '}
{/* Every authority, not just the largest. A quarter of outcodes
and a third of towns cross a boundary: SW19 is mostly Merton
but partly Wandsworth, and naming one asserts otherwise. */}
{authorities.map((a, i) => (
<span key={a.slug}>
{i > 0 && (i === authorities.length - 1 ? ' and ' : ', ')}
<Link href={`/schools/authority/${a.slug}`} className={styles.inlineLink}>
{a.name}
</Link>
</span>
))}
</>
)}
</p>
</header>
{!phase && (place.phases ?? []).length > 0 && (
<nav className={styles.phaseLinks} aria-label="By phase">
{(place.phases ?? []).map((ph) => (
<Link key={ph} href={`/schools/${place.slug}/${ph}`} className={styles.phaseLink}>
{ph === 'secondary' ? 'Secondary schools' : 'Primary schools'} in {place.name}
</Link>
))}
</nav>
)}
{local != null && englandAverage != null && (
<p className={styles.compare} data-testid="local-vs-england">
{place.name} averages <strong>{Math.round(local)}</strong> against{' '}
<strong>{Math.round(englandAverage)}</strong> across England.
</p>
)}
{graded.length > 0 && (
<ul className={styles.ofsted} data-testid="ofsted-distribution">
{graded.map(([label, n]) => (
<li key={label}>{label}: <strong>{n}</strong></li>
))}
</ul>
)}
{groups.map(([p, list]) => (
<section key={p} className={styles.group}>
{groups.length > 1 && (
<h2 className={styles.groupHeading}>
{p === 'secondary' ? 'Secondary schools' : 'Primary schools'}
<span className={styles.groupCount}>{list.length}</span>
</h2>
)}
<SchoolTable schools={list} phase={p} />
</section>
))}
{neighbours.length > 0 && (
<nav className={styles.neighbours} aria-label="Nearby places">
<h2>Nearby</h2>
<ul>
{neighbours.map((n) => (
<li key={n.kind + n.slug}>
<Link href={placeUrl(n.kind, n.slug)} className={styles.inlineLink}>{n.name}</Link>
</li>
))}
</ul>
</nav>
)}
</div>
);
}
+71
View File
@@ -0,0 +1,71 @@
/**
* Client for the places API.
*
* Two namespaces, matching the backend: towns and localities share
* /schools/[place]; authorities take /schools/authority/[la]. 67 town names
* collide with an authority name and neither set contains the other, so one
* namespace would publish near-duplicate pages.
*/
import type { School } from '@/lib/types';
export interface PlaceSummary {
kind: string;
slug: string;
name: string;
count: number;
/** Phases that clear the threshold on their own, so the page links
* variants that exist rather than 404s. Absent on the registry listing. */
phases?: string[];
}
export interface PlaceAuthority {
name: string;
slug: string;
count: number;
}
export interface PlaceDetail {
place: PlaceSummary & {
parent_authority: string | null;
/** Every authority the place meaningfully sits in, largest first. SW19 is
* mostly Merton but partly Wandsworth. */
authorities?: PlaceAuthority[];
};
schools: School[];
averages: {
rwm_expected_pct: number | null;
attainment_8_score: number | null;
};
}
export function placeUrl(kind: string, slug: string, phase?: string): string {
const base =
kind === 'authority' ? `/schools/authority/${slug}`
: kind === 'outcode' ? `/schools/near/${slug}`
: `/schools/${slug}`;
return phase ? `${base}/${phase}` : base;
}
/** An authority name as it appears in a URL. */
export function authoritySlug(name: string): string {
return name.toLowerCase().trim().replace(/[^\w\s-]/g, '').replace(/\s+/g, '-');
}
const API = process.env.FASTAPI_URL || process.env.NEXT_PUBLIC_API_URL
|| 'http://localhost:8000/api';
export async function fetchPlaces(): Promise<PlaceSummary[]> {
const res = await fetch(`${API}/places`, { next: { revalidate: 604800 } });
if (!res.ok) return [];
return (await res.json()).places ?? [];
}
export async function fetchPlace(
kind: string, slug: string, phase?: string,
): Promise<PlaceDetail | null> {
const q = phase ? `?phase=${encodeURIComponent(phase)}` : '';
const res = await fetch(`${API}/places/${kind}/${slug}${q}`,
{ next: { revalidate: 604800 } });
if (!res.ok) return null;
return res.json();
}
@@ -0,0 +1,16 @@
locality_slug,locality_name,outcodes,region
battersea,Battersea,SW11,London
canary-wharf,Canary Wharf,E14,London
clapham,Clapham,SW4,London
shoreditch,Shoreditch,EC2A|E1,London
peckham,Peckham,SE15,London
brixton,Brixton,SW2|SW9,London
camden-town,Camden Town,NW1,London
wimbledon,Wimbledon,SW19,London
putney,Putney,SW15,London
fulham,Fulham,SW6,London
chiswick,Chiswick,W4,London
stratford,Stratford,E15,London
walthamstow,Walthamstow,E17,London
tooting,Tooting,SW17,London
dulwich,Dulwich,SE21|SE22,London
1 locality_slug locality_name outcodes region
2 battersea Battersea SW11 London
3 canary-wharf Canary Wharf E14 London
4 clapham Clapham SW4 London
5 shoreditch Shoreditch EC2A|E1 London
6 peckham Peckham SE15 London
7 brixton Brixton SW2|SW9 London
8 camden-town Camden Town NW1 London
9 wimbledon Wimbledon SW19 London
10 putney Putney SW15 London
11 fulham Fulham SW6 London
12 chiswick Chiswick W4 London
13 stratford Stratford E15 London
14 walthamstow Walthamstow E17 London
15 tooting Tooting SW17 London
16 dulwich Dulwich SE21|SE22 London