Supersedes the original spec's W2. The Search Console baseline inverted its ordering: every measured location query is town or district level, none is an administrative area, and phase is part of the query rather than a filter. Two problems the original design did not anticipate. 67 viable towns share a name with a local authority, and the authority is the larger set in only 43 of them — postal towns cross authority boundaries, so neither can absorb the other. Two namespaces resolve it by construction. And the GIAS town field collapses 1,819 London schools into one value, which a curated locality-to-outcode seed solves without new ingestion. Sizing is measured against the live 25,185-school corpus rather than estimated: 783 viable towns, 1,760 outcodes, 154 authorities. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
11 KiB
W2: The Location Layer — Design
Date: 2026-08-21
Status: awaiting review
Supersedes: workstream W2 in 2026-08-20-seo-programme-design.md
Scope note: this covers four page families in one spec. Splitting them — towns and authorities first, outcodes and localities after — was proposed and declined in favour of building the layer in one pass. The decomposition argument was that the curated locality seed needs human review and would hold up 783 pages of measured demand behind it; that risk is accepted here, and the implementation plan should sequence the seed early enough that review time does not become the critical path.
Problem
Location intent is the largest unserved demand the site has. In the 16-month Search Console baseline it draws 874 impressions, one click, average position 49.5. The site does not compete.
Unlike named-school queries — which the same baseline showed to be navigational and unwinnable, since a parent typing "audley junior school" wants that school's own website — location queries have no incumbent owner. Nobody owns "primary schools in Brentwood" the way a school owns its name.
The cause is structural: the site has no page about a place. Every competitor ranking above it does.
What the demand actually looks like
Every location query in the baseline is town or district level. Not one is an administrative area:
| Query | Impressions | Position |
|---|---|---|
| colleges in solihull | 112 | 51.2 |
| schools in ramsey | 64 | 42.5 |
| schools in crosby | 57 | 47.7 |
| primary schools in beccles | 44 | 40.9 |
| private schools in battersea | 41 | 71.9 |
| secondary schools in brentwood | 37 | 56.1 |
| secondary schools in canary wharf | 30 | 35.9 |
Three patterns follow directly, and they drive the whole design.
Towns, not authorities. The superseded W2 put /schools/[la] first and
towns second. The data inverts that. Brentwood appears four times in different
phrasings; Beccles twice. Both are towns, not authorities.
Phase is part of the query, not a filter applied afterwards: "primary schools in beccles", "secondary schools in brentwood", "colleges in solihull".
London is searched by district — Battersea, Canary Wharf — and the GIAS
town field cannot serve it at all.
Measured sizing
Counted against the live corpus of 25,185 schools, not estimated.
| Family | Viable (≥5 schools) | Below threshold |
|---|---|---|
| Towns | 783 | 907 → redirect to authority |
| Outcodes | 1,760 | 305 |
| Local authorities | 154 | — |
| London localities | ~100–150 (curated) | — |
With phase variants — 783 town pages plus roughly 700 primary and 250 secondary variants, 154 authorities across three variants, 1,760 outcodes and the curated localities — the total lands near 4,000 pages. Phase variants need their own threshold: there are 17,426 primaries but only 4,456 secondaries nationally, so most towns will support a primary page and not a secondary one.
Two design problems this spec exists to solve
1. Town and authority names collide, and neither contains the other
67 viable towns share a name with a local authority. The obvious fix — let the authority absorb the town, since it sounds like a superset — does not work:
| Place | Schools in the town | Schools in the authority |
|---|---|---|
| Bedford | 104 | 86 |
| Birmingham | 520 | 518 |
| Derby | 157 | 119 |
| Doncaster | 152 | 145 |
The authority is the larger set in only 43 of the 67. Postal towns cross authority boundaries, so these are overlapping sets that happen to share a name. Publishing both into one namespace produces near-duplicate pages, which is the specific failure that sinks programmatic SEO.
Resolution: two namespaces.
/schools/[place] towns and London localities
/schools/[place]/primary
/schools/[place]/secondary
/schools/authority/[la] local authorities
/schools/authority/[la]/primary
/schools/authority/[la]/secondary
/schools/near/[outcode]
Outcodes carry no phase variants: nobody searches "primary schools in SW11", so the variants would be pages without demand.
Every collision disappears by construction. /schools/[place] keeps the clean
URL for the pattern that carries the demand; authorities get a namespace whose
purpose is genuinely different — admissions are authority-run, and the
authority page is the one that can speak to catchment policy and LA averages.
A place page and an authority page of the same name must each say plainly which set of schools they cover, or they read as duplicates to a reader even when they differ in fact.
2. London has no locality field
town collapses 1,819 London schools into the single value "London". A
page listing all of them is useless, and borough pages do not help because
people search "Battersea", not "Wandsworth".
No single field solves it:
| Search term | parliamentary_constituency |
postcodes.io admin_ward |
|---|---|---|
| Battersea | Battersea ✓ | Northcote / Wandsworth Town ✗ |
| Canary Wharf | Poplar and Limehouse ✗ | Canary Wharf ✓ |
| Vauxhall | Vauxhall and Camberwell Green ✗ | Vauxhall ✓ |
And neither covers Clapham, Shoreditch or Peckham, which are postal and colloquial rather than administrative.
Resolution: a curated seed mapping locality to outcodes.
pipeline/transform/seeds/locality_outcodes.csv
locality_slug,locality_name,outcodes,region
battersea,Battersea,"SW11|SW8",London
canary-wharf,Canary Wharf,"E14",London
clapham,Clapham,"SW4|SW9",London
This needs no new ingestion — the corpus already has postcodes. It puts
the fuzzy, contested part of the problem in a reviewable file rather than in
derived logic, which suits it: locality boundaries are a judgement, not a
fact. The repo already uses dbt seeds for curated reference data
(la_code_names.csv, gias_code_names.csv), so this follows an established
pattern.
The seed generalises past London. Any colloquial place — Jesmond, Chorlton, Clifton — can be defined by its outcodes without a schema change.
Constraint: a locality slug may not collide with a viable town slug. The place registry enforces this and fails the build rather than silently shadowing a town.
Architecture
The place registry
One module owns the question "what places do we publish, and what is in each". Everything else reads from it: the pages, the sitemap, the internal links.
backend/places.py
Place = { kind: "town"|"locality"|"authority"|"outcode",
slug, name, urn_list, parent_authority | None }
build_place_registry(df) -> dict[str, Place]
place_schools(slug, phase=None) -> list[School]
Built once at startup from the same DataFrame the sitemap uses, and rebuilt by
the existing /api/admin/regenerate-sitemap path after a pipeline run.
Registry construction is where the threshold, the collision rules and the
seed's uniqueness constraint are enforced — in one place, testable without a
browser or a database.
API
GET /api/places the registry: slug, kind, name, count
GET /api/places/{slug}?phase= aggregate + ranked schools for one place
/api/places is what the sitemap and the internal-link modules enumerate.
Routes
Next App Router, ISR with the same 7-day revalidate the school pages use.
generateStaticParams gated behind an env flag, matching
PRERENDER_SCHOOLS, because 3,900 more routes cannot be statically built in
CI on every deploy.
What each page must contain
A place page that is a name substituted into a template is the thing Google's helpful-content stance exists to demote. Each page carries computed local facts that exist nowhere else on the site:
- H1 matching the query: "Primary schools in Brentwood"
- Counts framed usefully: "29 schools, 4 rated Outstanding"
- A ranked table of the top 20 on the phase's headline metric —
rwm_expected_pctfor primary,attainment_8_scorefor secondary, and for an unphased place page the metric matching whichever phase it holds more of - The local average against the England average — the one number a parent cannot get from a list
- Ofsted grade distribution for the place
- A map
- Links to neighbouring places and to the parent authority
- An FAQ block, feeding
FAQPagestructured data - A link to every school page in scope — this is what finally de-orphans the 23,000 school pages the original spec identified as near-orphans
Thin-page controls
Three, and they are the difference between a location layer and index bloat:
- Five schools with current data minimum. Below it, 301 to the parent authority. This drops 907 towns and 305 outcodes.
- Per-phase thresholds. A town with 30 primaries and 2 secondaries publishes a primary page and no secondary page.
- No page without a local average. If a place has too few schools with results to compute one, it has nothing to say that a list does not, and it falls back to the authority.
Sitemap
Two new children in the existing index: /sitemaps/places-{n}.xml and
/sitemaps/outcodes-{n}.xml. Per-family children are why the index was built
in W1 — Search Console reports coverage per submitted sitemap, so indexation
of the location layer is measurable separately from the school pages.
Testing
Per CLAUDE.md, user-facing behaviour extends e2e/tests/journeys.spec.ts in
the same PR.
Unit (registry, no DB): threshold enforcement; a sub-threshold town resolves to its authority; a locality slug colliding with a town fails the build; Bedford's town and authority pages hold different URN sets; per-phase thresholds.
Backend: /api/places shape; /api/places/{slug} aggregate correctness
against a fixture; unknown slug 404s.
e2e: a known town, authority, locality and outcode page each render with
the expected count; a below-threshold town 301s; every place page declares a
canonical and appears in the sitemap; /schools/bedford and
/schools/authority/bedford both resolve and state which set they cover.
Risks
Index bloat is the failure mode of every programmatic SEO programme. The three controls above are the answer, and the per-family sitemap is how we find out early if they were not enough.
Helpful-content exposure. Templated location pages are exactly what Google's stance targets. The mitigation is that every page carries real computed local data — counts, distributions, local-versus-national comparison — rather than a name dropped into boilerplate. If indexation of the places sitemap stalls below roughly half, that is the signal to stop and rethink rather than to add more pages.
Build cost. ~4,000 additional ISR routes on top of 23,000 school pages.
The env-flag gate on generateStaticParams keeps CI viable.
Curation drift. The locality seed is hand-maintained and will go stale as places change. It is small and reviewable, and a dbt test asserts every seed outcode matches at least one school so a typo fails the pipeline rather than publishing an empty page.
Out of scope
Catchment-area estimation. It is a strong driver for this cluster and
fact_admissions carries the distances, but it is a modelling problem with
real accuracy risk and deserves its own design.