fix(seo): crawl hygiene and a per-family sitemap index (W1) #108

Merged
tudor merged 5 commits from feat/seo-crawl-hygiene into feat/england-only-corpus 2026-08-20 22:12:05 +00:00
Owner

Stacked on #107. Retarget to main once that merges.

Workstream W1 of the SEO programme — spec in docs/superpowers/specs/2026-08-20-seo-programme-design.md, plan in docs/superpowers/plans/2026-08-20-w1-crawl-hygiene.md.

The fault this started from

The site ranks second for "school compare" because that query is an exact match for the brand and domain. That is a naming win, not an SEO win, and it transfers to nothing. Investigating why the terms that carry actual demand do not rank turned up four structural problems, all fixed here.

What changed

Every canonical pointed at a redirect. The apex 301s to www at Cloudflare, but metadataBase, the school-page canonical, robots.txt's Sitemap: line and every sitemap <loc> named the apex. ~25,000 URLs each costing a redirect hop. Now one shared SITE_URL in nextjs-app/lib/site.ts, matched by BASE_URL in the backend.

The homepage had no canonical at all, while accepting eleven search params — so every filter combination was a crawlable near-duplicate of the page we most want to rank for "compare schools". All combinations now collapse onto /. Rankings and admissions had no canonical either.

/compare?urns=… was indexable. 25,193 schools make ~317 million pairs before triples. Now noindex, follow with a canonical to the bare path, so its outbound links to each school page still count. Bare /compare stays indexable — it is the landing page for the "compare schools" head term.

The sitemap listed every URN with priority and changefreq (both ignored by Google) and no lastmod (which Google does read). Now it drops schools with neither results nor an Ofsted grade, adds /admissions which was never listed, and carries a real lastmod.

Split into a sitemap index with children under /sitemaps/, chunked at 10,000. This is for diagnostics, not size — 25,193 is well under the 50,000 limit, but Search Console reports coverage per submitted sitemap, so one file per page family is what will make the location pages measurable when W2 lands.

Two judgement calls worth reviewing

lastmod comes from each school's Ofsted date and is omitted when unknown. An always-now lastmod is a claim Google learns to distrust; absent honestly means unknown. On the index it legitimately means "when this file changed", so generation time is correct there — different semantics, different value.

Children sit under /sitemaps/ rather than /sitemap-*.xml. Next only treats a whole bracketed path segment as dynamic — verified in Next's router source, where UrlNode._insert only reads a segment as dynamic if it startsWith('[') && endsWith(']'). A folder named sitemap-[...parts] would have been inserted as a literal static segment and never matched. The build output listing /sitemaps/[...parts] as a dynamic route is the proof it resolves.

Verification

  • Backend: 68 passed (54 at baseline)
  • Frontend: 218 passed, tsc --noEmit clean
  • e2e: 67 journeys register, up 9
  • next build succeeds, both sitemap routes present

⚠️ Before the e2e gate runs

Same as #107 — staging's marts need an Airflow run after the deploy and before the Playwright journeys, or the sitemap and England-only tests fail. Manual trigger, agreed.

Not in this PR

A better template for the ~3,927 English schools with no data. This PR stops submitting them, which is the crawl-hygiene half; giving them something worth showing is a product change and belongs with its own design.

🤖 Generated with Claude Code

https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj

**Stacked on #107.** Retarget to `main` once that merges. Workstream W1 of the SEO programme — spec in `docs/superpowers/specs/2026-08-20-seo-programme-design.md`, plan in `docs/superpowers/plans/2026-08-20-w1-crawl-hygiene.md`. ## The fault this started from The site ranks second for "school compare" because that query is an exact match for the brand and domain. That is a naming win, not an SEO win, and it transfers to nothing. Investigating why the terms that carry actual demand do not rank turned up four structural problems, all fixed here. ## What changed **Every canonical pointed at a redirect.** The apex 301s to `www` at Cloudflare, but `metadataBase`, the school-page canonical, `robots.txt`'s `Sitemap:` line and every sitemap `<loc>` named the apex. ~25,000 URLs each costing a redirect hop. Now one shared `SITE_URL` in `nextjs-app/lib/site.ts`, matched by `BASE_URL` in the backend. **The homepage had no canonical at all**, while accepting eleven search params — so every filter combination was a crawlable near-duplicate of the page we most want to rank for "compare schools". All combinations now collapse onto `/`. Rankings and admissions had no canonical either. **`/compare?urns=…` was indexable.** 25,193 schools make ~317 million pairs before triples. Now `noindex, follow` with a canonical to the bare path, so its outbound links to each school page still count. Bare `/compare` stays indexable — it is the landing page for the "compare schools" head term. **The sitemap listed every URN** with `priority` and `changefreq` (both ignored by Google) and no `lastmod` (which Google does read). Now it drops schools with neither results nor an Ofsted grade, adds `/admissions` which was never listed, and carries a real `lastmod`. **Split into a sitemap index** with children under `/sitemaps/`, chunked at 10,000. This is for diagnostics, not size — 25,193 is well under the 50,000 limit, but Search Console reports coverage per submitted sitemap, so one file per page family is what will make the location pages measurable when W2 lands. ## Two judgement calls worth reviewing **`lastmod` comes from each school's Ofsted date and is omitted when unknown.** An always-`now` `lastmod` is a claim Google learns to distrust; absent honestly means unknown. On the *index* it legitimately means "when this file changed", so generation time is correct there — different semantics, different value. **Children sit under `/sitemaps/` rather than `/sitemap-*.xml`.** Next only treats a whole bracketed path segment as dynamic — verified in Next's router source, where `UrlNode._insert` only reads a segment as dynamic if it `startsWith('[') && endsWith(']')`. A folder named `sitemap-[...parts]` would have been inserted as a literal static segment and never matched. The build output listing `/sitemaps/[...parts]` as a dynamic route is the proof it resolves. ## Verification - Backend: 68 passed (54 at baseline) - Frontend: 218 passed, `tsc --noEmit` clean - e2e: 67 journeys register, up 9 - `next build` succeeds, both sitemap routes present ## ⚠️ Before the e2e gate runs Same as #107 — staging's marts need an Airflow run after the deploy and before the Playwright journeys, or the sitemap and England-only tests fail. Manual trigger, agreed. ## Not in this PR A better template for the ~3,927 English schools with no data. This PR stops submitting them, which is the crawl-hygiene half; giving them something worth showing is a product change and belongs with its own design. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
tudor added 5 commits 2026-08-20 21:17:49 +00:00
The apex 301s to www at Cloudflare, but metadataBase, the school-page
canonical, robots.txt's Sitemap: line and the sitemap's own <loc> entries all
named the apex. Every one of those pointed Google at a redirect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The homepage read eleven search params and declared no canonical, so every
filter combination was a crawlable near-duplicate of the page we most want to
rank. Rankings and admissions declared none either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
25,193 schools make ~317 million pairs. The bare page stays indexable as the
landing page for the head term; the parameter space goes noindex, follow so
its outbound links still count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Drops the schools with neither results nor an Ofsted grade, adds /admissions
which was never listed, replaces the invented priority and changefreq with a
lastmod taken from each school's Ofsted date.

lastmod is omitted where no date is known rather than defaulted to now. An
always-now lastmod is a claim Google learns to distrust; absent honestly
means unknown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Search Console reports coverage per submitted sitemap, so one file per page
family is what will make W2's location pages measurable when they land. The
index's lastmod is generation time, which is the correct semantic there —
unlike on a <url>, where it would be a claim we cannot support.

Children sit under /sitemaps/ because Next only treats a whole bracketed path
segment as dynamic; a route folder named sitemap-[...parts] would be read as a
literal static segment and never match. Confirmed by the build output, which
lists /sitemaps/[...parts] as a dynamic route.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
tudor merged commit a5ac0bcd1b into feat/england-only-corpus 2026-08-20 22:12:05 +00:00
Sign in to join this conversation.
No Reviewers
No labels
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: tudor/school_compare#108