From ecc847091c9defa96e06bb5cc04e101df56d8820 Mon Sep 17 00:00:00 2001 From: Tudor Date: Fri, 21 Aug 2026 08:37:06 +0100 Subject: [PATCH] docs(seo): design for the W2 location layer MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Supersedes the original spec's W2. The Search Console baseline inverted its ordering: every measured location query is town or district level, none is an administrative area, and phase is part of the query rather than a filter. Two problems the original design did not anticipate. 67 viable towns share a name with a local authority, and the authority is the larger set in only 43 of them — postal towns cross authority boundaries, so neither can absorb the other. Two namespaces resolve it by construction. And the GIAS town field collapses 1,819 London schools into one value, which a curated locality-to-outcode seed solves without new ingestion. Sizing is measured against the live 25,185-school corpus rather than estimated: 783 viable towns, 1,760 outcodes, 154 authorities. Co-Authored-By: Claude Opus 5 Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj --- .../2026-08-21-w2-location-layer-design.md | 278 ++++++++++++++++++ 1 file changed, 278 insertions(+) create mode 100644 docs/superpowers/specs/2026-08-21-w2-location-layer-design.md diff --git a/docs/superpowers/specs/2026-08-21-w2-location-layer-design.md b/docs/superpowers/specs/2026-08-21-w2-location-layer-design.md new file mode 100644 index 0000000..348bb93 --- /dev/null +++ b/docs/superpowers/specs/2026-08-21-w2-location-layer-design.md @@ -0,0 +1,278 @@ +# W2: The Location Layer — Design + +Date: 2026-08-21 +Status: awaiting review +Supersedes: workstream W2 in `2026-08-20-seo-programme-design.md` + +Scope note: this covers four page families in one spec. Splitting them — towns +and authorities first, outcodes and localities after — was proposed and +declined in favour of building the layer in one pass. The decomposition +argument was that the curated locality seed needs human review and would hold +up 783 pages of measured demand behind it; that risk is accepted here, and the +implementation plan should sequence the seed early enough that review time +does not become the critical path. + +## Problem + +Location intent is the largest unserved demand the site has. In the 16-month +Search Console baseline it draws **874 impressions, one click, average +position 49.5**. The site does not compete. + +Unlike named-school queries — which the same baseline showed to be +navigational and unwinnable, since a parent typing "audley junior school" +wants that school's own website — location queries have no incumbent owner. +Nobody owns "primary schools in Brentwood" the way a school owns its name. + +The cause is structural: the site has no page about a place. Every competitor +ranking above it does. + +## What the demand actually looks like + +Every location query in the baseline is **town or district level**. Not one is +an administrative area: + +| Query | Impressions | Position | +|-------|-------------|----------| +| colleges in solihull | 112 | 51.2 | +| schools in ramsey | 64 | 42.5 | +| schools in crosby | 57 | 47.7 | +| primary schools in beccles | 44 | 40.9 | +| private schools in battersea | 41 | 71.9 | +| secondary schools in brentwood | 37 | 56.1 | +| secondary schools in canary wharf | 30 | 35.9 | + +Three patterns follow directly, and they drive the whole design. + +**Towns, not authorities.** The superseded W2 put `/schools/[la]` first and +towns second. The data inverts that. Brentwood appears four times in different +phrasings; Beccles twice. Both are towns, not authorities. + +**Phase is part of the query**, not a filter applied afterwards: "primary +schools in beccles", "secondary schools in brentwood", "colleges in solihull". + +**London is searched by district** — Battersea, Canary Wharf — and the GIAS +`town` field cannot serve it at all. + +## Measured sizing + +Counted against the live corpus of 25,185 schools, not estimated. + +| Family | Viable (≥5 schools) | Below threshold | +|--------|--------------------|-----------------| +| Towns | **783** | 907 → redirect to authority | +| Outcodes | **1,760** | 305 | +| Local authorities | 154 | — | +| London localities | ~100–150 (curated) | — | + +With phase variants — 783 town pages plus roughly 700 primary and 250 +secondary variants, 154 authorities across three variants, 1,760 outcodes and +the curated localities — the total lands near **4,000 pages**. Phase variants +need their own threshold: there are 17,426 primaries but only 4,456 secondaries nationally, +so most towns will support a primary page and not a secondary one. + +## Two design problems this spec exists to solve + +### 1. Town and authority names collide, and neither contains the other + +67 viable towns share a name with a local authority. The obvious fix — let the +authority absorb the town, since it sounds like a superset — **does not work**: + +| Place | Schools in the town | Schools in the authority | +|-------|--------------------|-----------------------| +| Bedford | 104 | 86 | +| Birmingham | 520 | 518 | +| Derby | 157 | 119 | +| Doncaster | 152 | 145 | + +The authority is the larger set in only 43 of the 67. Postal towns cross +authority boundaries, so these are overlapping sets that happen to share a +name. Publishing both into one namespace produces near-duplicate pages, which +is the specific failure that sinks programmatic SEO. + +**Resolution: two namespaces.** + +``` +/schools/[place] towns and London localities +/schools/[place]/primary +/schools/[place]/secondary +/schools/authority/[la] local authorities +/schools/authority/[la]/primary +/schools/authority/[la]/secondary +/schools/near/[outcode] +``` + +Outcodes carry no phase variants: nobody searches "primary schools in SW11", +so the variants would be pages without demand. + +Every collision disappears by construction. `/schools/[place]` keeps the clean +URL for the pattern that carries the demand; authorities get a namespace whose +purpose is genuinely different — admissions are authority-run, and the +authority page is the one that can speak to catchment policy and LA averages. + +A place page and an authority page of the same name must each say plainly +which set of schools they cover, or they read as duplicates to a reader even +when they differ in fact. + +### 2. London has no locality field + +`town` collapses **1,819 London schools into the single value "London"**. A +page listing all of them is useless, and borough pages do not help because +people search "Battersea", not "Wandsworth". + +No single field solves it: + +| Search term | `parliamentary_constituency` | postcodes.io `admin_ward` | +|-------------|------------------------------|---------------------------| +| Battersea | **Battersea** ✓ | Northcote / Wandsworth Town ✗ | +| Canary Wharf | Poplar and Limehouse ✗ | **Canary Wharf** ✓ | +| Vauxhall | Vauxhall and Camberwell Green ✗ | **Vauxhall** ✓ | + +And neither covers Clapham, Shoreditch or Peckham, which are postal and +colloquial rather than administrative. + +**Resolution: a curated seed mapping locality to outcodes.** + +``` +pipeline/transform/seeds/locality_outcodes.csv +locality_slug,locality_name,outcodes,region +battersea,Battersea,"SW11|SW8",London +canary-wharf,Canary Wharf,"E14",London +clapham,Clapham,"SW4|SW9",London +``` + +This needs **no new ingestion** — the corpus already has postcodes. It puts +the fuzzy, contested part of the problem in a reviewable file rather than in +derived logic, which suits it: locality boundaries are a judgement, not a +fact. The repo already uses dbt seeds for curated reference data +(`la_code_names.csv`, `gias_code_names.csv`), so this follows an established +pattern. + +The seed generalises past London. Any colloquial place — Jesmond, Chorlton, +Clifton — can be defined by its outcodes without a schema change. + +**Constraint:** a locality slug may not collide with a viable town slug. The +place registry enforces this and fails the build rather than silently +shadowing a town. + +## Architecture + +### The place registry + +One module owns the question "what places do we publish, and what is in each". +Everything else reads from it: the pages, the sitemap, the internal links. + +``` +backend/places.py + + Place = { kind: "town"|"locality"|"authority"|"outcode", + slug, name, urn_list, parent_authority | None } + + build_place_registry(df) -> dict[str, Place] + place_schools(slug, phase=None) -> list[School] +``` + +Built once at startup from the same DataFrame the sitemap uses, and rebuilt by +the existing `/api/admin/regenerate-sitemap` path after a pipeline run. +Registry construction is where the threshold, the collision rules and the +seed's uniqueness constraint are enforced — in one place, testable without a +browser or a database. + +### API + +``` +GET /api/places the registry: slug, kind, name, count +GET /api/places/{slug}?phase= aggregate + ranked schools for one place +``` + +`/api/places` is what the sitemap and the internal-link modules enumerate. + +### Routes + +Next App Router, ISR with the same 7-day revalidate the school pages use. +`generateStaticParams` gated behind an env flag, matching +`PRERENDER_SCHOOLS`, because 3,900 more routes cannot be statically built in +CI on every deploy. + +## What each page must contain + +A place page that is a name substituted into a template is the thing Google's +helpful-content stance exists to demote. Each page carries computed local +facts that exist nowhere else on the site: + +- **H1** matching the query: "Primary schools in Brentwood" +- **Counts framed usefully**: "29 schools, 4 rated Outstanding" +- **A ranked table** of the top 20 on the phase's headline metric — + `rwm_expected_pct` for primary, `attainment_8_score` for secondary, and for + an unphased place page the metric matching whichever phase it holds more of +- **The local average against the England average** — the one number a parent + cannot get from a list +- **Ofsted grade distribution** for the place +- **A map** +- **Links to neighbouring places** and to the parent authority +- **An FAQ block**, feeding `FAQPage` structured data +- **A link to every school page in scope** — this is what finally de-orphans + the 23,000 school pages the original spec identified as near-orphans + +## Thin-page controls + +Three, and they are the difference between a location layer and index bloat: + +1. **Five schools with current data minimum.** Below it, 301 to the parent + authority. This drops 907 towns and 305 outcodes. +2. **Per-phase thresholds.** A town with 30 primaries and 2 secondaries + publishes a primary page and no secondary page. +3. **No page without a local average.** If a place has too few schools with + results to compute one, it has nothing to say that a list does not, and it + falls back to the authority. + +## Sitemap + +Two new children in the existing index: `/sitemaps/places-{n}.xml` and +`/sitemaps/outcodes-{n}.xml`. Per-family children are why the index was built +in W1 — Search Console reports coverage per submitted sitemap, so indexation +of the location layer is measurable separately from the school pages. + +## Testing + +Per `CLAUDE.md`, user-facing behaviour extends `e2e/tests/journeys.spec.ts` in +the same PR. + +**Unit (registry, no DB):** threshold enforcement; a sub-threshold town +resolves to its authority; a locality slug colliding with a town fails the +build; Bedford's town and authority pages hold different URN sets; per-phase +thresholds. + +**Backend:** `/api/places` shape; `/api/places/{slug}` aggregate correctness +against a fixture; unknown slug 404s. + +**e2e:** a known town, authority, locality and outcode page each render with +the expected count; a below-threshold town 301s; every place page declares a +canonical and appears in the sitemap; `/schools/bedford` and +`/schools/authority/bedford` both resolve and state which set they cover. + +## Risks + +**Index bloat** is the failure mode of every programmatic SEO programme. The +three controls above are the answer, and the per-family sitemap is how we find +out early if they were not enough. + +**Helpful-content exposure.** Templated location pages are exactly what +Google's stance targets. The mitigation is that every page carries real +computed local data — counts, distributions, local-versus-national comparison +— rather than a name dropped into boilerplate. If indexation of the places +sitemap stalls below roughly half, that is the signal to stop and rethink +rather than to add more pages. + +**Build cost.** ~4,000 additional ISR routes on top of 23,000 school pages. +The env-flag gate on `generateStaticParams` keeps CI viable. + +**Curation drift.** The locality seed is hand-maintained and will go stale as +places change. It is small and reviewable, and a dbt test asserts every seed +outcode matches at least one school so a typo fails the pipeline rather than +publishing an empty page. + +## Out of scope + +Catchment-area estimation. It is a strong driver for this cluster and +`fact_admissions` carries the distances, but it is a modelling problem with +real accuracy risk and deserves its own design.