fix(seo): crawl hygiene and a per-family sitemap index (W1) #110

Merged
tudor merged 5 commits from feat/seo-crawl-hygiene-main into main 2026-08-20 22:23:10 +00:00
Owner

Re-targets #108, which merged into feat/england-only-corpus instead of main and so never reached the trunk. Same five commits, cherry-picked onto current main (which already carries #107 and #109). No content differences.

Spec: docs/superpowers/specs/2026-08-20-seo-programme-design.md. Plan: docs/superpowers/plans/2026-08-20-w1-crawl-hygiene.md.

What changed

Every canonical pointed at a redirect. The apex 301s to www at Cloudflare, but metadataBase, the school-page canonical, robots.txt's Sitemap: line and every sitemap <loc> named the apex — ~25,000 URLs each costing a hop. Now one shared SITE_URL in nextjs-app/lib/site.ts, matched by BASE_URL in the backend.

The homepage had no canonical at all, while accepting eleven search params, so every filter combination was a crawlable near-duplicate of the page we most want to rank for "compare schools". All combinations now collapse onto /. Rankings and admissions had none either.

/compare?urns=… was indexable — ~317 million pairs of parameter space. Now noindex, follow with a canonical to the bare path, so its outbound links to each school page still count. Bare /compare stays indexable as the landing page for the head term.

The sitemap listed every URN with priority and changefreq (both ignored by Google) and no lastmod (which is not). Now drops schools with neither results nor Ofsted, adds /admissions which was never listed, and carries a real lastmod.

Split into a sitemap index with children under /sitemaps/, chunked at 10,000. For diagnostics, not size — Search Console reports coverage per submitted sitemap, so one file per family is what will make W2's location pages measurable.

Two judgement calls worth reviewing

lastmod comes from each school's Ofsted date and is omitted when unknown. An always-now lastmod is a claim Google learns to distrust; absent honestly means unknown. On the index it legitimately means "when this file changed", so generation time is correct there.

Children sit under /sitemaps/. Next only treats a whole bracketed path segment as dynamic — verified in Next's router source, where UrlNode._insert only reads a segment as dynamic if it startsWith('[') && endsWith(']'). A folder named sitemap-[...parts] would have been a literal static segment that never matched, and nothing would have errored — the children would just have 404'd in production.

Verification on the combined tree

Re-run after cherry-picking onto main, not carried over from #108:

  • Backend: 68 passed
  • Frontend: 219 passed, tsc --noEmit clean
  • e2e: 69 journeys register — #109's Beechwood and Gloucestershire-border tests coexist with W1's canonical and sitemap-index tests
  • next build green, /sitemaps/[...parts] listed as a dynamic route
  • Both mart filters intact after the cherry-pick: establishment type and #109's Welsh LA code range, in both dim_school and dim_location

Expect the URL count to drop

Staging currently reports 25,188 sitemap URLs. After this, expect roughly 21,300 across an index plus three children — that is the ~3,900 English schools with neither results nor Ofsted no longer being submitted. Not a regression.

⚠️ Staging sequencing

No mart changes here, so no Airflow run is strictly required. But the sitemap is cached in-process — hit POST /api/admin/regenerate-sitemap after deploy, or the e2e sitemap assertions read a stale flat sitemap and fail.

🤖 Generated with Claude Code

https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj

Re-targets #108, which merged into `feat/england-only-corpus` instead of `main` and so never reached the trunk. Same five commits, cherry-picked onto current `main` (which already carries #107 and #109). No content differences. Spec: `docs/superpowers/specs/2026-08-20-seo-programme-design.md`. Plan: `docs/superpowers/plans/2026-08-20-w1-crawl-hygiene.md`. ## What changed **Every canonical pointed at a redirect.** The apex 301s to `www` at Cloudflare, but `metadataBase`, the school-page canonical, `robots.txt`'s `Sitemap:` line and every sitemap `<loc>` named the apex — ~25,000 URLs each costing a hop. Now one shared `SITE_URL` in `nextjs-app/lib/site.ts`, matched by `BASE_URL` in the backend. **The homepage had no canonical at all**, while accepting eleven search params, so every filter combination was a crawlable near-duplicate of the page we most want to rank for "compare schools". All combinations now collapse onto `/`. Rankings and admissions had none either. **`/compare?urns=…` was indexable** — ~317 million pairs of parameter space. Now `noindex, follow` with a canonical to the bare path, so its outbound links to each school page still count. Bare `/compare` stays indexable as the landing page for the head term. **The sitemap** listed every URN with `priority` and `changefreq` (both ignored by Google) and no `lastmod` (which is not). Now drops schools with neither results nor Ofsted, adds `/admissions` which was never listed, and carries a real `lastmod`. **Split into a sitemap index** with children under `/sitemaps/`, chunked at 10,000. For diagnostics, not size — Search Console reports coverage per submitted sitemap, so one file per family is what will make W2's location pages measurable. ## Two judgement calls worth reviewing **`lastmod` comes from each school's Ofsted date and is omitted when unknown.** An always-`now` `lastmod` is a claim Google learns to distrust; absent honestly means unknown. On the *index* it legitimately means "when this file changed", so generation time is correct there. **Children sit under `/sitemaps/`.** Next only treats a whole bracketed path segment as dynamic — verified in Next's router source, where `UrlNode._insert` only reads a segment as dynamic if it `startsWith('[') && endsWith(']')`. A folder named `sitemap-[...parts]` would have been a literal static segment that never matched, and nothing would have errored — the children would just have 404'd in production. ## Verification on the combined tree Re-run after cherry-picking onto `main`, not carried over from #108: - Backend: 68 passed - Frontend: 219 passed, `tsc --noEmit` clean - e2e: 69 journeys register — #109's Beechwood and Gloucestershire-border tests coexist with W1's canonical and sitemap-index tests - `next build` green, `/sitemaps/[...parts]` listed as a dynamic route - Both mart filters intact after the cherry-pick: establishment type **and** #109's Welsh LA code range, in both `dim_school` and `dim_location` ## Expect the URL count to drop Staging currently reports 25,188 sitemap URLs. After this, expect roughly **21,300 across an index plus three children** — that is the ~3,900 English schools with neither results nor Ofsted no longer being submitted. Not a regression. ## ⚠️ Staging sequencing No mart changes here, so no Airflow run is strictly required. But the sitemap is cached in-process — hit `POST /api/admin/regenerate-sitemap` after deploy, or the e2e sitemap assertions read a stale flat sitemap and fail. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
tudor added 5 commits 2026-08-20 22:16:15 +00:00
The apex 301s to www at Cloudflare, but metadataBase, the school-page
canonical, robots.txt's Sitemap: line and the sitemap's own <loc> entries all
named the apex. Every one of those pointed Google at a redirect.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The homepage read eleven search params and declared no canonical, so every
filter combination was a crawlable near-duplicate of the page we most want to
rank. Rankings and admissions declared none either.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
25,193 schools make ~317 million pairs. The bare page stays indexable as the
landing page for the head term; the parameter space goes noindex, follow so
its outbound links still count.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Drops the schools with neither results nor an Ofsted grade, adds /admissions
which was never listed, replaces the invented priority and changefreq with a
lastmod taken from each school's Ofsted date.

lastmod is omitted where no date is known rather than defaulted to now. An
always-now lastmod is a claim Google learns to distrust; absent honestly
means unknown.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
feat(seo): split the sitemap into a per-family index
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m4s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 17s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m2s
2208ad93c1
Search Console reports coverage per submitted sitemap, so one file per page
family is what will make W2's location pages measurable when they land. The
index's lastmod is generation time, which is the correct semantic there —
unlike on a <url>, where it would be a claim we cannot support.

Children sit under /sitemaps/ because Next only treats a whole bracketed path
segment as dynamic; a route folder named sitemap-[...parts] would be read as a
literal static segment and never match. Confirmed by the build output, which
lists /sitemaps/[...parts] as a dynamic route.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj

🤖 AI Code Review (Claude Code)

This PR splits the single flat sitemap into an index plus chunked child sitemaps (static.xml, schools-N.xml), moves the canonical site origin from the apex domain to www across backend and frontend, and adds canonical/robots metadata for the homepage, rankings, admissions, and compare pages. The change is thorough and well covered by new backend, e2e, and Next.js unit tests, with no severe correctness, security, or deploy issues found.

🟡 Minor

  • backend/app.py: build_sitemap()'s docstring says it's 'Kept for lifespan and the admin endpoint', but both lifespan() and regenerate_sitemap() now call build_sitemaps() (plural) directly — the singular function is dead/unused outside tests, and the stale comment could mislead future maintainers into thinking it's still load-bearing.
  • backend/app.py: _school_sitemap_rows sorts by year descending and only evaluates _has_publishable_data on the single latest-year row per URN. A school with no rwm/attainment8/Ofsted data in its most recent year but valid data in an earlier year will be dropped from the sitemap entirely, even though its page has content worth indexing from that earlier year.
## 🤖 AI Code Review (Claude Code) This PR splits the single flat sitemap into an index plus chunked child sitemaps (static.xml, schools-N.xml), moves the canonical site origin from the apex domain to www across backend and frontend, and adds canonical/robots metadata for the homepage, rankings, admissions, and compare pages. The change is thorough and well covered by new backend, e2e, and Next.js unit tests, with no severe correctness, security, or deploy issues found. ### 🟡 Minor - **backend/app.py**: build_sitemap()'s docstring says it's 'Kept for lifespan and the admin endpoint', but both lifespan() and regenerate_sitemap() now call build_sitemaps() (plural) directly — the singular function is dead/unused outside tests, and the stale comment could mislead future maintainers into thinking it's still load-bearing. - **backend/app.py**: _school_sitemap_rows sorts by year descending and only evaluates _has_publishable_data on the single latest-year row per URN. A school with no rwm/attainment8/Ofsted data in its most recent year but valid data in an earlier year will be dropped from the sitemap entirely, even though its page has content worth indexing from that earlier year.
tudor merged commit f928a15c1e into main 2026-08-20 22:23:10 +00:00
Sign in to join this conversation.
No Reviewers
No labels
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: tudor/school_compare#110