feat(seo): the location layer — town, locality, authority and outcode pages (W2) #114

Merged
tudor merged 9 commits from feat/w2-location-layer into main 2026-08-21 17:56:06 +00:00
Owner

W2 of the SEO programme. Spec: docs/superpowers/specs/2026-08-21-w2-location-layer-design.md. Plan: docs/superpowers/plans/2026-08-21-w2-location-layer.md.

Why

Location intent is the largest winnable gap in the W0 baseline: 874 impressions, one click, average position 49.5 across 16 months. The site published no page about a place.

The contrast with C4 is the whole argument. Named-school queries draw far more impressions but are navigational — a parent typing audley junior school wants that school own website, and no SEO changes that. Nobody owns primary schools in Brentwood the way a school owns its name.

What ships

Four page families, roughly 4,000 pages, sized against the live 25,185-school corpus rather than estimated:

Family Route Viable
Towns /schools/[place] 783
Phase variants /schools/[place]/[phase] ~950
Authorities /schools/authority/[la] 154
Outcodes /schools/near/[outcode] 1,760
London localities /schools/[place] 20 curated

Two design problems worth reviewing

Town and authority names collide, and neither contains the other. 67 viable towns share a name with an authority. The obvious fix — let the authority absorb the town — is wrong in 24 cases: Bedford the town has 104 schools against the authority 86, Derby 157 against 119, because postal towns cross authority boundaries. Two namespaces eliminate all 67 by construction.

London has no locality field. The GIAS town field puts 1,819 schools under London. Neither candidate field solves it: parliamentary constituency gives Battersea but not Canary Wharf; admin_ward gives Canary Wharf but not Battersea; neither gives Clapham or Shoreditch. So localities are curated as locality-to-outcode mappings, needing no new ingestion since the corpus already has postcodes.

The curated list lives in backend/localities.py rather than a dbt seed, because the backend image does not contain pipeline/ — it copies only backend/ and scripts/. Same arrangement as gias_codes.py, with a test guarding drift against the seed mirror.

Three thin-page controls

Index bloat is how programmatic SEO fails, so these are the load-bearing parts:

  1. Five schools with publishable data minimum — drops 907 towns and 305 outcodes.
  2. Per-phase thresholds — 17,426 primaries against 4,456 secondaries nationally, so most towns get a primary page and no secondary one.
  3. No page without a local average — a place that cannot compute one has nothing a list does not, and defers to its authority.

Every page carries computed local facts — counts, Ofsted distribution, and the local average against England — rather than a name dropped into boilerplate. That last one is the reason these are not lists, and there is a test pinning it.

Two bugs found while executing

A test fixture that proved nothing. The Bedford collision test built both authorities schools from the same URN range, so the sets were identical and the assertion passed for the wrong reason.

The build failed with ECONNREFUSED. The plan claimed authority pages were few enough to always prebuild — but few enough still means the API must be reachable at build time, and in CI it is not. All three prerenders are now gated behind PRERENDER_PLACES and wrapped in the same try/catch fallback the school route uses.

Verification

Backend 95 passed (70 at baseline) · frontend 244 · tsc --noEmit clean · next build green with all four families resolving · 81 e2e journeys registering.

Before merging

Staging needs an Airflow run, then POST /api/admin/regenerate-sitemap. The registry is built from the marts and cached per process; against a stale staging database it publishes a different set of places than the tests expect.

The stop condition

Agreed in the spec, and the reason the sitemap splits per family: if indexation of the places sitemap stalls below roughly half, stop and rethink rather than adding more pages.

Deliberately deferred

A map per place page (no SEO weight, and 4,000 Leaflet instances need their own performance measurement), an FAQ block (a copy task that would hold 4,000 pages behind copywriting), and catchment estimation (a modelling problem where a wrong boundary is worse than none).

🤖 Generated with Claude Code

https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj

W2 of the SEO programme. Spec: `docs/superpowers/specs/2026-08-21-w2-location-layer-design.md`. Plan: `docs/superpowers/plans/2026-08-21-w2-location-layer.md`. ## Why Location intent is the largest **winnable** gap in the [W0 baseline](https://claude.ai/code/artifact/418c0e30-1782-464a-bb48-3feadefe829f): **874 impressions, one click, average position 49.5** across 16 months. The site published no page about a place. The contrast with C4 is the whole argument. Named-school queries draw far more impressions but are navigational — a parent typing `audley junior school` wants that school own website, and no SEO changes that. Nobody owns `primary schools in Brentwood` the way a school owns its name. ## What ships Four page families, roughly **4,000 pages**, sized against the live 25,185-school corpus rather than estimated: | Family | Route | Viable | |---|---|---| | Towns | `/schools/[place]` | 783 | | Phase variants | `/schools/[place]/[phase]` | ~950 | | Authorities | `/schools/authority/[la]` | 154 | | Outcodes | `/schools/near/[outcode]` | 1,760 | | London localities | `/schools/[place]` | 20 curated | ## Two design problems worth reviewing **Town and authority names collide, and neither contains the other.** 67 viable towns share a name with an authority. The obvious fix — let the authority absorb the town — is wrong in 24 cases: Bedford the town has 104 schools against the authority 86, Derby 157 against 119, because postal towns cross authority boundaries. Two namespaces eliminate all 67 by construction. **London has no locality field.** The GIAS `town` field puts 1,819 schools under `London`. Neither candidate field solves it: parliamentary constituency gives Battersea but not Canary Wharf; `admin_ward` gives Canary Wharf but not Battersea; neither gives Clapham or Shoreditch. So localities are curated as locality-to-outcode mappings, needing **no new ingestion** since the corpus already has postcodes. The curated list lives in `backend/localities.py` rather than a dbt seed, because **the backend image does not contain `pipeline/`** — it copies only `backend/` and `scripts/`. Same arrangement as `gias_codes.py`, with a test guarding drift against the seed mirror. ## Three thin-page controls Index bloat is how programmatic SEO fails, so these are the load-bearing parts: 1. **Five schools with publishable data minimum** — drops 907 towns and 305 outcodes. 2. **Per-phase thresholds** — 17,426 primaries against 4,456 secondaries nationally, so most towns get a primary page and no secondary one. 3. **No page without a local average** — a place that cannot compute one has nothing a list does not, and defers to its authority. Every page carries computed local facts — counts, Ofsted distribution, and the local average against England — rather than a name dropped into boilerplate. That last one is the reason these are not lists, and there is a test pinning it. ## Two bugs found while executing **A test fixture that proved nothing.** The Bedford collision test built both authorities schools from the same URN range, so the sets were identical and the assertion passed for the wrong reason. **The build failed with ECONNREFUSED.** The plan claimed authority pages were `few enough to always prebuild` — but few enough still means the API must be reachable at build time, and in CI it is not. All three prerenders are now gated behind `PRERENDER_PLACES` and wrapped in the same try/catch fallback the school route uses. ## Verification Backend 95 passed (70 at baseline) · frontend 244 · `tsc --noEmit` clean · `next build` green with all four families resolving · 81 e2e journeys registering. ## Before merging **Staging needs an Airflow run**, then `POST /api/admin/regenerate-sitemap`. The registry is built from the marts and cached per process; against a stale staging database it publishes a different set of places than the tests expect. ## The stop condition Agreed in the spec, and the reason the sitemap splits per family: **if indexation of the places sitemap stalls below roughly half, stop and rethink rather than adding more pages.** ## Deliberately deferred A **map** per place page (no SEO weight, and 4,000 Leaflet instances need their own performance measurement), an **FAQ block** (a copy task that would hold 4,000 pages behind copywriting), and **catchment estimation** (a modelling problem where a wrong boundary is worse than none). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
tudor added 9 commits 2026-08-21 17:28:26 +00:00
Supersedes the original spec's W2. The Search Console baseline inverted its
ordering: every measured location query is town or district level, none is an
administrative area, and phase is part of the query rather than a filter.

Two problems the original design did not anticipate. 67 viable towns share a
name with a local authority, and the authority is the larger set in only 43 of
them — postal towns cross authority boundaries, so neither can absorb the
other. Two namespaces resolve it by construction. And the GIAS town field
collapses 1,819 London schools into one value, which a curated
locality-to-outcode seed solves without new ingestion.

Sizing is measured against the live 25,185-school corpus rather than
estimated: 783 viable towns, 1,760 outcodes, 154 authorities.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Seven tasks: the place registry, London localities and outcodes, the places
API, per-family sitemaps, the shared place view, the four route families, and
structured data plus the e2e gate.

Two things the plan corrects against the spec. The backend image does not
contain pipeline/, so the curated locality list cannot live only in a dbt
seed — it follows the gias_codes.py precedent instead, canonical in backend
with the seed as a mirror. And NationalAverages is nested by phase rather than
flat, which the first draft read wrongly and would have rendered every page
without its England comparison.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Two namespaces because 67 town names collide with an authority name and
neither set contains the other — postal towns cross authority boundaries, so
Bedford the town holds 104 schools against the authority's 86.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The GIAS town field puts 1,819 London schools under the single value
'London', so it cannot answer 'schools in Battersea' — a query that appears in
the baseline. No single field can: parliamentary constituency gives Battersea
but not Canary Wharf, admin_ward gives Canary Wharf but not Battersea, and
neither gives Clapham or Shoreditch. So a locality is curated, defined by the
postcode districts it covers, which needs no new ingestion.

A locality may not shadow a published town: the registry raises rather than
silently costing a page that carries real demand. One below the threshold is
logged rather than raising, because a locality can legitimately be too small.

The pipeline seed mirrors the module, with a test guarding the drift — the
same arrangement gias_codes has, and for the same reason: the backend image
does not contain pipeline/.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The registry is cached for the process and reset by the same admin endpoint
that rebuilds the sitemaps, so places and sitemap always describe the same
corpus rather than drifting apart.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Separate children per family so Search Console reports the location layer's
indexation apart from the school pages' — which is the point of the index
built in W1, and the number the stop condition watches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
One component for all four families: they differ in what fills the registry,
not in what the page shows, so a second would be a second place to forget the
same change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Every generateStaticParams is gated behind PRERENDER_PLACES and wrapped in the
same try/catch the school route uses. The plan claimed authority pages were
'few enough to always prebuild' — but few enough still means the API must be
reachable at build time, and in CI it is not: the build failed with
ECONNREFUSED rather than degrading to ISR.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
feat(places): ItemList and BreadcrumbList, and the e2e gate
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 18s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 36s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 4m33s
6b871ce1e9
ItemList tells Google the page is a ranked set rather than prose;
BreadcrumbList puts the place in a hierarchy. School URLs in the markup are
absolute on the canonical host, since a relative URL in JSON-LD is ambiguous.

Eight journeys covering all four families, the two-namespace guarantee, the
threshold, the canonical, the sitemap and the local-versus-England line — the
last because that comparison is the reason these pages are not a name dropped
into a template.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj

🤖 AI Code Review (Claude Code)

This PR adds a place registry (towns, curated London localities, local authorities, postcode districts) with two new backend endpoints, extended sitemaps, and four new Next.js route families to serve ~4,000 location pages. The core registry logic is well-tested, but error handling around the new places code is inconsistent with the existing sitemap code it sits next to, and one frontend feature (nearby-places links) doesn't do what its own comment claims.

🔴 Severe (blocks merge)

  • backend/app.py: get_place_registry() (used by list_places, get_place, and build_sitemaps via _place_sitemap_rows) has no exception handling around build_place_registry(). places.py's _locality_places() deliberately raises ValueError when a curated locality slug (backend/localities.py) collides with a real GIAS town slug of the same name — several curated slugs (richmond, greenwich, wimbledon, ealing, fulham, putney, chiswick, etc.) are plausible standalone GIAS town values outside the 'London' bucket this feature exists to work around. Unlike the two pre-existing sitemap call sites (lifespan startup and _serve_sitemap), which already wrap the equivalent build call in try/except, the new /api/places*, /api/places/{kind}/{slug} and POST /api/admin/regenerate-sitemap endpoints call it unguarded. If the collision occurs on real data, the next Airflow-triggered regenerate-sitemap call: (1) raises before build_sitemaps() returns, so _sitemaps is never reassigned and the entire sitemap — schools included, ~23k URLs — silently stops updating on every future pipeline run, and (2) leaves _place_registry permanently None, so every subsequent call to /api/places and /api/places/{kind}/{slug} 500s with no self-healing until a code fix is deployed.

🟡 Minor

  • nextjs-app/app/schools/[place]/page.tsx: neighboursOf() is documented as 'Other towns in the same authority — the cheapest honest definition of nearby', but it only filters by kind === 'town' and excludes the current slug, then slices the first 12 — it never filters by parent_authority, and structurally can't from this call, since the PlaceSummary objects returned by fetchPlaces()/'/api/places' don't carry parent_authority at all (only the single-place detail endpoint does). Every place page's 'Nearby' section therefore shows an arbitrary alphabetically-first slice of all ~783 towns nationwide rather than places actually near the current one. PlaceView.test.tsx doesn't catch this because it passes a mocked neighbours prop directly instead of exercising neighboursOf().
  • backend/app.py: In get_place(), the phase filter uses phase.lower() (PHASE_GROUPS.get(phase.lower())) but the very next line selects the ranking/averaging metric with a case-sensitive comparison: metric = "attainment_8_score" if phase == "secondary" else "rwm_expected_pct". A request like ?phase=Secondary correctly restricts rows to secondary schools but still sorts and reports the average using rwm_expected_pct instead of attainment_8_score. Not reachable from the shipped frontend routes (which always pass lowercase 'secondary'), but reachable via direct API calls.
## 🤖 AI Code Review (Claude Code) This PR adds a place registry (towns, curated London localities, local authorities, postcode districts) with two new backend endpoints, extended sitemaps, and four new Next.js route families to serve ~4,000 location pages. The core registry logic is well-tested, but error handling around the new places code is inconsistent with the existing sitemap code it sits next to, and one frontend feature (nearby-places links) doesn't do what its own comment claims. ### 🔴 Severe (blocks merge) - **backend/app.py**: get_place_registry() (used by list_places, get_place, and build_sitemaps via _place_sitemap_rows) has no exception handling around build_place_registry(). places.py's _locality_places() deliberately raises ValueError when a curated locality slug (backend/localities.py) collides with a real GIAS town slug of the same name — several curated slugs (richmond, greenwich, wimbledon, ealing, fulham, putney, chiswick, etc.) are plausible standalone GIAS town values outside the 'London' bucket this feature exists to work around. Unlike the two pre-existing sitemap call sites (lifespan startup and _serve_sitemap), which already wrap the equivalent build call in try/except, the new /api/places*, /api/places/{kind}/{slug} and POST /api/admin/regenerate-sitemap endpoints call it unguarded. If the collision occurs on real data, the next Airflow-triggered regenerate-sitemap call: (1) raises before build_sitemaps() returns, so _sitemaps is never reassigned and the entire sitemap — schools included, ~23k URLs — silently stops updating on every future pipeline run, and (2) leaves _place_registry permanently None, so every subsequent call to /api/places and /api/places/{kind}/{slug} 500s with no self-healing until a code fix is deployed. ### 🟡 Minor - **nextjs-app/app/schools/[place]/page.tsx**: neighboursOf() is documented as 'Other towns in the same authority — the cheapest honest definition of nearby', but it only filters by kind === 'town' and excludes the current slug, then slices the first 12 — it never filters by parent_authority, and structurally can't from this call, since the PlaceSummary objects returned by fetchPlaces()/'/api/places' don't carry parent_authority at all (only the single-place detail endpoint does). Every place page's 'Nearby' section therefore shows an arbitrary alphabetically-first slice of all ~783 towns nationwide rather than places actually near the current one. PlaceView.test.tsx doesn't catch this because it passes a mocked neighbours prop directly instead of exercising neighboursOf(). - **backend/app.py**: In get_place(), the phase filter uses phase.lower() (PHASE_GROUPS.get(phase.lower())) but the very next line selects the ranking/averaging metric with a case-sensitive comparison: metric = "attainment_8_score" if phase == "secondary" else "rwm_expected_pct". A request like ?phase=Secondary correctly restricts rows to secondary schools but still sorts and reports the average using rwm_expected_pct instead of attainment_8_score. Not reachable from the shipped frontend routes (which always pass lowercase 'secondary'), but reachable via direct API calls.
tudor merged commit b93eb3a691 into main 2026-08-21 17:56:06 +00:00
Sign in to join this conversation.
No Reviewers
No labels
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: tudor/school_compare#114