docs(seo): design for the W2 location layer
Supersedes the original spec's W2. The Search Console baseline inverted its ordering: every measured location query is town or district level, none is an administrative area, and phase is part of the query rather than a filter. Two problems the original design did not anticipate. 67 viable towns share a name with a local authority, and the authority is the larger set in only 43 of them — postal towns cross authority boundaries, so neither can absorb the other. Two namespaces resolve it by construction. And the GIAS town field collapses 1,819 London schools into one value, which a curated locality-to-outcode seed solves without new ingestion. Sizing is measured against the live 25,185-school corpus rather than estimated: 783 viable towns, 1,760 outcodes, 154 authorities. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
This commit is contained in:
1 parent
b187a478c9
commit
ecc847091c
1 file changed
+278
@@ -0,0 +1,278 @@
|
||||
# W2: The Location Layer — Design
|
||||
|
||||
Date: 2026-08-21
|
||||
Status: awaiting review
|
||||
Supersedes: workstream W2 in `2026-08-20-seo-programme-design.md`
|
||||
|
||||
Scope note: this covers four page families in one spec. Splitting them — towns
|
||||
and authorities first, outcodes and localities after — was proposed and
|
||||
declined in favour of building the layer in one pass. The decomposition
|
||||
argument was that the curated locality seed needs human review and would hold
|
||||
up 783 pages of measured demand behind it; that risk is accepted here, and the
|
||||
implementation plan should sequence the seed early enough that review time
|
||||
does not become the critical path.
|
||||
|
||||
## Problem
|
||||
|
||||
Location intent is the largest unserved demand the site has. In the 16-month
|
||||
Search Console baseline it draws **874 impressions, one click, average
|
||||
position 49.5**. The site does not compete.
|
||||
|
||||
Unlike named-school queries — which the same baseline showed to be
|
||||
navigational and unwinnable, since a parent typing "audley junior school"
|
||||
wants that school's own website — location queries have no incumbent owner.
|
||||
Nobody owns "primary schools in Brentwood" the way a school owns its name.
|
||||
|
||||
The cause is structural: the site has no page about a place. Every competitor
|
||||
ranking above it does.
|
||||
|
||||
## What the demand actually looks like
|
||||
|
||||
Every location query in the baseline is **town or district level**. Not one is
|
||||
an administrative area:
|
||||
|
||||
| Query | Impressions | Position |
|
||||
|-------|-------------|----------|
|
||||
| colleges in solihull | 112 | 51.2 |
|
||||
| schools in ramsey | 64 | 42.5 |
|
||||
| schools in crosby | 57 | 47.7 |
|
||||
| primary schools in beccles | 44 | 40.9 |
|
||||
| private schools in battersea | 41 | 71.9 |
|
||||
| secondary schools in brentwood | 37 | 56.1 |
|
||||
| secondary schools in canary wharf | 30 | 35.9 |
|
||||
|
||||
Three patterns follow directly, and they drive the whole design.
|
||||
|
||||
**Towns, not authorities.** The superseded W2 put `/schools/[la]` first and
|
||||
towns second. The data inverts that. Brentwood appears four times in different
|
||||
phrasings; Beccles twice. Both are towns, not authorities.
|
||||
|
||||
**Phase is part of the query**, not a filter applied afterwards: "primary
|
||||
schools in beccles", "secondary schools in brentwood", "colleges in solihull".
|
||||
|
||||
**London is searched by district** — Battersea, Canary Wharf — and the GIAS
|
||||
`town` field cannot serve it at all.
|
||||
|
||||
## Measured sizing
|
||||
|
||||
Counted against the live corpus of 25,185 schools, not estimated.
|
||||
|
||||
| Family | Viable (≥5 schools) | Below threshold |
|
||||
|--------|--------------------|-----------------|
|
||||
| Towns | **783** | 907 → redirect to authority |
|
||||
| Outcodes | **1,760** | 305 |
|
||||
| Local authorities | 154 | — |
|
||||
| London localities | ~100–150 (curated) | — |
|
||||
|
||||
With phase variants — 783 town pages plus roughly 700 primary and 250
|
||||
secondary variants, 154 authorities across three variants, 1,760 outcodes and
|
||||
the curated localities — the total lands near **4,000 pages**. Phase variants
|
||||
need their own threshold: there are 17,426 primaries but only 4,456 secondaries nationally,
|
||||
so most towns will support a primary page and not a secondary one.
|
||||
|
||||
## Two design problems this spec exists to solve
|
||||
|
||||
### 1. Town and authority names collide, and neither contains the other
|
||||
|
||||
67 viable towns share a name with a local authority. The obvious fix — let the
|
||||
authority absorb the town, since it sounds like a superset — **does not work**:
|
||||
|
||||
| Place | Schools in the town | Schools in the authority |
|
||||
|-------|--------------------|-----------------------|
|
||||
| Bedford | 104 | 86 |
|
||||
| Birmingham | 520 | 518 |
|
||||
| Derby | 157 | 119 |
|
||||
| Doncaster | 152 | 145 |
|
||||
|
||||
The authority is the larger set in only 43 of the 67. Postal towns cross
|
||||
authority boundaries, so these are overlapping sets that happen to share a
|
||||
name. Publishing both into one namespace produces near-duplicate pages, which
|
||||
is the specific failure that sinks programmatic SEO.
|
||||
|
||||
**Resolution: two namespaces.**
|
||||
|
||||
```
|
||||
/schools/[place] towns and London localities
|
||||
/schools/[place]/primary
|
||||
/schools/[place]/secondary
|
||||
/schools/authority/[la] local authorities
|
||||
/schools/authority/[la]/primary
|
||||
/schools/authority/[la]/secondary
|
||||
/schools/near/[outcode]
|
||||
```
|
||||
|
||||
Outcodes carry no phase variants: nobody searches "primary schools in SW11",
|
||||
so the variants would be pages without demand.
|
||||
|
||||
Every collision disappears by construction. `/schools/[place]` keeps the clean
|
||||
URL for the pattern that carries the demand; authorities get a namespace whose
|
||||
purpose is genuinely different — admissions are authority-run, and the
|
||||
authority page is the one that can speak to catchment policy and LA averages.
|
||||
|
||||
A place page and an authority page of the same name must each say plainly
|
||||
which set of schools they cover, or they read as duplicates to a reader even
|
||||
when they differ in fact.
|
||||
|
||||
### 2. London has no locality field
|
||||
|
||||
`town` collapses **1,819 London schools into the single value "London"**. A
|
||||
page listing all of them is useless, and borough pages do not help because
|
||||
people search "Battersea", not "Wandsworth".
|
||||
|
||||
No single field solves it:
|
||||
|
||||
| Search term | `parliamentary_constituency` | postcodes.io `admin_ward` |
|
||||
|-------------|------------------------------|---------------------------|
|
||||
| Battersea | **Battersea** ✓ | Northcote / Wandsworth Town ✗ |
|
||||
| Canary Wharf | Poplar and Limehouse ✗ | **Canary Wharf** ✓ |
|
||||
| Vauxhall | Vauxhall and Camberwell Green ✗ | **Vauxhall** ✓ |
|
||||
|
||||
And neither covers Clapham, Shoreditch or Peckham, which are postal and
|
||||
colloquial rather than administrative.
|
||||
|
||||
**Resolution: a curated seed mapping locality to outcodes.**
|
||||
|
||||
```
|
||||
pipeline/transform/seeds/locality_outcodes.csv
|
||||
locality_slug,locality_name,outcodes,region
|
||||
battersea,Battersea,"SW11|SW8",London
|
||||
canary-wharf,Canary Wharf,"E14",London
|
||||
clapham,Clapham,"SW4|SW9",London
|
||||
```
|
||||
|
||||
This needs **no new ingestion** — the corpus already has postcodes. It puts
|
||||
the fuzzy, contested part of the problem in a reviewable file rather than in
|
||||
derived logic, which suits it: locality boundaries are a judgement, not a
|
||||
fact. The repo already uses dbt seeds for curated reference data
|
||||
(`la_code_names.csv`, `gias_code_names.csv`), so this follows an established
|
||||
pattern.
|
||||
|
||||
The seed generalises past London. Any colloquial place — Jesmond, Chorlton,
|
||||
Clifton — can be defined by its outcodes without a schema change.
|
||||
|
||||
**Constraint:** a locality slug may not collide with a viable town slug. The
|
||||
place registry enforces this and fails the build rather than silently
|
||||
shadowing a town.
|
||||
|
||||
## Architecture
|
||||
|
||||
### The place registry
|
||||
|
||||
One module owns the question "what places do we publish, and what is in each".
|
||||
Everything else reads from it: the pages, the sitemap, the internal links.
|
||||
|
||||
```
|
||||
backend/places.py
|
||||
|
||||
Place = { kind: "town"|"locality"|"authority"|"outcode",
|
||||
slug, name, urn_list, parent_authority | None }
|
||||
|
||||
build_place_registry(df) -> dict[str, Place]
|
||||
place_schools(slug, phase=None) -> list[School]
|
||||
```
|
||||
|
||||
Built once at startup from the same DataFrame the sitemap uses, and rebuilt by
|
||||
the existing `/api/admin/regenerate-sitemap` path after a pipeline run.
|
||||
Registry construction is where the threshold, the collision rules and the
|
||||
seed's uniqueness constraint are enforced — in one place, testable without a
|
||||
browser or a database.
|
||||
|
||||
### API
|
||||
|
||||
```
|
||||
GET /api/places the registry: slug, kind, name, count
|
||||
GET /api/places/{slug}?phase= aggregate + ranked schools for one place
|
||||
```
|
||||
|
||||
`/api/places` is what the sitemap and the internal-link modules enumerate.
|
||||
|
||||
### Routes
|
||||
|
||||
Next App Router, ISR with the same 7-day revalidate the school pages use.
|
||||
`generateStaticParams` gated behind an env flag, matching
|
||||
`PRERENDER_SCHOOLS`, because 3,900 more routes cannot be statically built in
|
||||
CI on every deploy.
|
||||
|
||||
## What each page must contain
|
||||
|
||||
A place page that is a name substituted into a template is the thing Google's
|
||||
helpful-content stance exists to demote. Each page carries computed local
|
||||
facts that exist nowhere else on the site:
|
||||
|
||||
- **H1** matching the query: "Primary schools in Brentwood"
|
||||
- **Counts framed usefully**: "29 schools, 4 rated Outstanding"
|
||||
- **A ranked table** of the top 20 on the phase's headline metric —
|
||||
`rwm_expected_pct` for primary, `attainment_8_score` for secondary, and for
|
||||
an unphased place page the metric matching whichever phase it holds more of
|
||||
- **The local average against the England average** — the one number a parent
|
||||
cannot get from a list
|
||||
- **Ofsted grade distribution** for the place
|
||||
- **A map**
|
||||
- **Links to neighbouring places** and to the parent authority
|
||||
- **An FAQ block**, feeding `FAQPage` structured data
|
||||
- **A link to every school page in scope** — this is what finally de-orphans
|
||||
the 23,000 school pages the original spec identified as near-orphans
|
||||
|
||||
## Thin-page controls
|
||||
|
||||
Three, and they are the difference between a location layer and index bloat:
|
||||
|
||||
1. **Five schools with current data minimum.** Below it, 301 to the parent
|
||||
authority. This drops 907 towns and 305 outcodes.
|
||||
2. **Per-phase thresholds.** A town with 30 primaries and 2 secondaries
|
||||
publishes a primary page and no secondary page.
|
||||
3. **No page without a local average.** If a place has too few schools with
|
||||
results to compute one, it has nothing to say that a list does not, and it
|
||||
falls back to the authority.
|
||||
|
||||
## Sitemap
|
||||
|
||||
Two new children in the existing index: `/sitemaps/places-{n}.xml` and
|
||||
`/sitemaps/outcodes-{n}.xml`. Per-family children are why the index was built
|
||||
in W1 — Search Console reports coverage per submitted sitemap, so indexation
|
||||
of the location layer is measurable separately from the school pages.
|
||||
|
||||
## Testing
|
||||
|
||||
Per `CLAUDE.md`, user-facing behaviour extends `e2e/tests/journeys.spec.ts` in
|
||||
the same PR.
|
||||
|
||||
**Unit (registry, no DB):** threshold enforcement; a sub-threshold town
|
||||
resolves to its authority; a locality slug colliding with a town fails the
|
||||
build; Bedford's town and authority pages hold different URN sets; per-phase
|
||||
thresholds.
|
||||
|
||||
**Backend:** `/api/places` shape; `/api/places/{slug}` aggregate correctness
|
||||
against a fixture; unknown slug 404s.
|
||||
|
||||
**e2e:** a known town, authority, locality and outcode page each render with
|
||||
the expected count; a below-threshold town 301s; every place page declares a
|
||||
canonical and appears in the sitemap; `/schools/bedford` and
|
||||
`/schools/authority/bedford` both resolve and state which set they cover.
|
||||
|
||||
## Risks
|
||||
|
||||
**Index bloat** is the failure mode of every programmatic SEO programme. The
|
||||
three controls above are the answer, and the per-family sitemap is how we find
|
||||
out early if they were not enough.
|
||||
|
||||
**Helpful-content exposure.** Templated location pages are exactly what
|
||||
Google's stance targets. The mitigation is that every page carries real
|
||||
computed local data — counts, distributions, local-versus-national comparison
|
||||
— rather than a name dropped into boilerplate. If indexation of the places
|
||||
sitemap stalls below roughly half, that is the signal to stop and rethink
|
||||
rather than to add more pages.
|
||||
|
||||
**Build cost.** ~4,000 additional ISR routes on top of 23,000 school pages.
|
||||
The env-flag gate on `generateStaticParams` keeps CI viable.
|
||||
|
||||
**Curation drift.** The locality seed is hand-maintained and will go stale as
|
||||
places change. It is small and reviewable, and a dbt test asserts every seed
|
||||
outcode matches at least one school so a typo fails the pipeline rather than
|
||||
publishing an empty page.
|
||||
|
||||
## Out of scope
|
||||
|
||||
Catchment-area estimation. It is a strong driver for this cluster and
|
||||
`fact_admissions` carries the distances, but it is a modelling problem with
|
||||
real accuracy risk and deserves its own design.
|
||||
Reference in new issue
Block a user