# W2: The Location Layer — Design Date: 2026-08-21 Status: awaiting review Supersedes: workstream W2 in `2026-08-20-seo-programme-design.md` Scope note: this covers four page families in one spec. Splitting them — towns and authorities first, outcodes and localities after — was proposed and declined in favour of building the layer in one pass. The decomposition argument was that the curated locality seed needs human review and would hold up 783 pages of measured demand behind it; that risk is accepted here, and the implementation plan should sequence the seed early enough that review time does not become the critical path. ## Problem Location intent is the largest unserved demand the site has. In the 16-month Search Console baseline it draws **874 impressions, one click, average position 49.5**. The site does not compete. Unlike named-school queries — which the same baseline showed to be navigational and unwinnable, since a parent typing "audley junior school" wants that school's own website — location queries have no incumbent owner. Nobody owns "primary schools in Brentwood" the way a school owns its name. The cause is structural: the site has no page about a place. Every competitor ranking above it does. ## What the demand actually looks like Every location query in the baseline is **town or district level**. Not one is an administrative area: | Query | Impressions | Position | |-------|-------------|----------| | colleges in solihull | 112 | 51.2 | | schools in ramsey | 64 | 42.5 | | schools in crosby | 57 | 47.7 | | primary schools in beccles | 44 | 40.9 | | private schools in battersea | 41 | 71.9 | | secondary schools in brentwood | 37 | 56.1 | | secondary schools in canary wharf | 30 | 35.9 | Three patterns follow directly, and they drive the whole design. **Towns, not authorities.** The superseded W2 put `/schools/[la]` first and towns second. The data inverts that. Brentwood appears four times in different phrasings; Beccles twice. Both are towns, not authorities. **Phase is part of the query**, not a filter applied afterwards: "primary schools in beccles", "secondary schools in brentwood", "colleges in solihull". **London is searched by district** — Battersea, Canary Wharf — and the GIAS `town` field cannot serve it at all. ## Measured sizing Counted against the live corpus of 25,185 schools, not estimated. | Family | Viable (≥5 schools) | Below threshold | |--------|--------------------|-----------------| | Towns | **783** | 907 → redirect to authority | | Outcodes | **1,760** | 305 | | Local authorities | 154 | — | | London localities | ~100–150 (curated) | — | With phase variants — 783 town pages plus roughly 700 primary and 250 secondary variants, 154 authorities across three variants, 1,760 outcodes and the curated localities — the total lands near **4,000 pages**. Phase variants need their own threshold: there are 17,426 primaries but only 4,456 secondaries nationally, so most towns will support a primary page and not a secondary one. ## Two design problems this spec exists to solve ### 1. Town and authority names collide, and neither contains the other 67 viable towns share a name with a local authority. The obvious fix — let the authority absorb the town, since it sounds like a superset — **does not work**: | Place | Schools in the town | Schools in the authority | |-------|--------------------|-----------------------| | Bedford | 104 | 86 | | Birmingham | 520 | 518 | | Derby | 157 | 119 | | Doncaster | 152 | 145 | The authority is the larger set in only 43 of the 67. Postal towns cross authority boundaries, so these are overlapping sets that happen to share a name. Publishing both into one namespace produces near-duplicate pages, which is the specific failure that sinks programmatic SEO. **Resolution: two namespaces.** ``` /schools/[place] towns and London localities /schools/[place]/primary /schools/[place]/secondary /schools/authority/[la] local authorities /schools/authority/[la]/primary /schools/authority/[la]/secondary /schools/near/[outcode] ``` Outcodes carry no phase variants: nobody searches "primary schools in SW11", so the variants would be pages without demand. Every collision disappears by construction. `/schools/[place]` keeps the clean URL for the pattern that carries the demand; authorities get a namespace whose purpose is genuinely different — admissions are authority-run, and the authority page is the one that can speak to catchment policy and LA averages. A place page and an authority page of the same name must each say plainly which set of schools they cover, or they read as duplicates to a reader even when they differ in fact. ### 2. London has no locality field `town` collapses **1,819 London schools into the single value "London"**. A page listing all of them is useless, and borough pages do not help because people search "Battersea", not "Wandsworth". No single field solves it: | Search term | `parliamentary_constituency` | postcodes.io `admin_ward` | |-------------|------------------------------|---------------------------| | Battersea | **Battersea** ✓ | Northcote / Wandsworth Town ✗ | | Canary Wharf | Poplar and Limehouse ✗ | **Canary Wharf** ✓ | | Vauxhall | Vauxhall and Camberwell Green ✗ | **Vauxhall** ✓ | And neither covers Clapham, Shoreditch or Peckham, which are postal and colloquial rather than administrative. **Resolution: a curated seed mapping locality to outcodes.** ``` pipeline/transform/seeds/locality_outcodes.csv locality_slug,locality_name,outcodes,region battersea,Battersea,"SW11|SW8",London canary-wharf,Canary Wharf,"E14",London clapham,Clapham,"SW4|SW9",London ``` This needs **no new ingestion** — the corpus already has postcodes. It puts the fuzzy, contested part of the problem in a reviewable file rather than in derived logic, which suits it: locality boundaries are a judgement, not a fact. The repo already uses dbt seeds for curated reference data (`la_code_names.csv`, `gias_code_names.csv`), so this follows an established pattern. The seed generalises past London. Any colloquial place — Jesmond, Chorlton, Clifton — can be defined by its outcodes without a schema change. **Constraint:** a locality slug may not collide with a viable town slug. The place registry enforces this and fails the build rather than silently shadowing a town. ## Architecture ### The place registry One module owns the question "what places do we publish, and what is in each". Everything else reads from it: the pages, the sitemap, the internal links. ``` backend/places.py Place = { kind: "town"|"locality"|"authority"|"outcode", slug, name, urn_list, parent_authority | None } build_place_registry(df) -> dict[str, Place] place_schools(slug, phase=None) -> list[School] ``` Built once at startup from the same DataFrame the sitemap uses, and rebuilt by the existing `/api/admin/regenerate-sitemap` path after a pipeline run. Registry construction is where the threshold, the collision rules and the seed's uniqueness constraint are enforced — in one place, testable without a browser or a database. ### API ``` GET /api/places the registry: slug, kind, name, count GET /api/places/{slug}?phase= aggregate + ranked schools for one place ``` `/api/places` is what the sitemap and the internal-link modules enumerate. ### Routes Next App Router, ISR with the same 7-day revalidate the school pages use. `generateStaticParams` gated behind an env flag, matching `PRERENDER_SCHOOLS`, because 3,900 more routes cannot be statically built in CI on every deploy. ## What each page must contain A place page that is a name substituted into a template is the thing Google's helpful-content stance exists to demote. Each page carries computed local facts that exist nowhere else on the site: - **H1** matching the query: "Primary schools in Brentwood" - **Counts framed usefully**: "29 schools, 4 rated Outstanding" - **A ranked table** of the top 20 on the phase's headline metric — `rwm_expected_pct` for primary, `attainment_8_score` for secondary, and for an unphased place page the metric matching whichever phase it holds more of - **The local average against the England average** — the one number a parent cannot get from a list - **Ofsted grade distribution** for the place - **A map** - **Links to neighbouring places** and to the parent authority - **An FAQ block**, feeding `FAQPage` structured data - **A link to every school page in scope** — this is what finally de-orphans the 23,000 school pages the original spec identified as near-orphans ## Thin-page controls Three, and they are the difference between a location layer and index bloat: 1. **Five schools with current data minimum.** Below it, 301 to the parent authority. This drops 907 towns and 305 outcodes. 2. **Per-phase thresholds.** A town with 30 primaries and 2 secondaries publishes a primary page and no secondary page. 3. **No page without a local average.** If a place has too few schools with results to compute one, it has nothing to say that a list does not, and it falls back to the authority. ## Sitemap Two new children in the existing index: `/sitemaps/places-{n}.xml` and `/sitemaps/outcodes-{n}.xml`. Per-family children are why the index was built in W1 — Search Console reports coverage per submitted sitemap, so indexation of the location layer is measurable separately from the school pages. ## Testing Per `CLAUDE.md`, user-facing behaviour extends `e2e/tests/journeys.spec.ts` in the same PR. **Unit (registry, no DB):** threshold enforcement; a sub-threshold town resolves to its authority; a locality slug colliding with a town fails the build; Bedford's town and authority pages hold different URN sets; per-phase thresholds. **Backend:** `/api/places` shape; `/api/places/{slug}` aggregate correctness against a fixture; unknown slug 404s. **e2e:** a known town, authority, locality and outcode page each render with the expected count; a below-threshold town 301s; every place page declares a canonical and appears in the sitemap; `/schools/bedford` and `/schools/authority/bedford` both resolve and state which set they cover. ## Risks **Index bloat** is the failure mode of every programmatic SEO programme. The three controls above are the answer, and the per-family sitemap is how we find out early if they were not enough. **Helpful-content exposure.** Templated location pages are exactly what Google's stance targets. The mitigation is that every page carries real computed local data — counts, distributions, local-versus-national comparison — rather than a name dropped into boilerplate. If indexation of the places sitemap stalls below roughly half, that is the signal to stop and rethink rather than to add more pages. **Build cost.** ~4,000 additional ISR routes on top of 23,000 school pages. The env-flag gate on `generateStaticParams` keeps CI viable. **Curation drift.** The locality seed is hand-maintained and will go stale as places change. It is small and reviewable, and a dbt test asserts every seed outcode matches at least one school so a typo fails the pipeline rather than publishing an empty page. ## Out of scope Catchment-area estimation. It is a strong driver for this cluster and `fact_admissions` carries the distances, but it is a modelling problem with real accuracy risk and deserves its own design.