Re-targets #108, which merged into feat/england-only-corpus instead of main and so never reached the trunk. Same five commits, cherry-picked onto current main (which already carries #107 and #109). No content differences.
Every canonical pointed at a redirect. The apex 301s to www at Cloudflare, but metadataBase, the school-page canonical, robots.txt's Sitemap: line and every sitemap <loc> named the apex — ~25,000 URLs each costing a hop. Now one shared SITE_URL in nextjs-app/lib/site.ts, matched by BASE_URL in the backend.
The homepage had no canonical at all, while accepting eleven search params, so every filter combination was a crawlable near-duplicate of the page we most want to rank for "compare schools". All combinations now collapse onto /. Rankings and admissions had none either.
/compare?urns=… was indexable — ~317 million pairs of parameter space. Now noindex, follow with a canonical to the bare path, so its outbound links to each school page still count. Bare /compare stays indexable as the landing page for the head term.
The sitemap listed every URN with priority and changefreq (both ignored by Google) and no lastmod (which is not). Now drops schools with neither results nor Ofsted, adds /admissions which was never listed, and carries a real lastmod.
Split into a sitemap index with children under /sitemaps/, chunked at 10,000. For diagnostics, not size — Search Console reports coverage per submitted sitemap, so one file per family is what will make W2's location pages measurable.
Two judgement calls worth reviewing
lastmod comes from each school's Ofsted date and is omitted when unknown. An always-nowlastmod is a claim Google learns to distrust; absent honestly means unknown. On the index it legitimately means "when this file changed", so generation time is correct there.
Children sit under /sitemaps/. Next only treats a whole bracketed path segment as dynamic — verified in Next's router source, where UrlNode._insert only reads a segment as dynamic if it startsWith('[') && endsWith(']'). A folder named sitemap-[...parts] would have been a literal static segment that never matched, and nothing would have errored — the children would just have 404'd in production.
Verification on the combined tree
Re-run after cherry-picking onto main, not carried over from #108:
Backend: 68 passed
Frontend: 219 passed, tsc --noEmit clean
e2e: 69 journeys register — #109's Beechwood and Gloucestershire-border tests coexist with W1's canonical and sitemap-index tests
next build green, /sitemaps/[...parts] listed as a dynamic route
Both mart filters intact after the cherry-pick: establishment type and#109's Welsh LA code range, in both dim_school and dim_location
Expect the URL count to drop
Staging currently reports 25,188 sitemap URLs. After this, expect roughly 21,300 across an index plus three children — that is the ~3,900 English schools with neither results nor Ofsted no longer being submitted. Not a regression.
⚠️ Staging sequencing
No mart changes here, so no Airflow run is strictly required. But the sitemap is cached in-process — hit POST /api/admin/regenerate-sitemap after deploy, or the e2e sitemap assertions read a stale flat sitemap and fail.
Re-targets #108, which merged into `feat/england-only-corpus` instead of `main` and so never reached the trunk. Same five commits, cherry-picked onto current `main` (which already carries #107 and #109). No content differences.
Spec: `docs/superpowers/specs/2026-08-20-seo-programme-design.md`. Plan: `docs/superpowers/plans/2026-08-20-w1-crawl-hygiene.md`.
## What changed
**Every canonical pointed at a redirect.** The apex 301s to `www` at Cloudflare, but `metadataBase`, the school-page canonical, `robots.txt`'s `Sitemap:` line and every sitemap `<loc>` named the apex — ~25,000 URLs each costing a hop. Now one shared `SITE_URL` in `nextjs-app/lib/site.ts`, matched by `BASE_URL` in the backend.
**The homepage had no canonical at all**, while accepting eleven search params, so every filter combination was a crawlable near-duplicate of the page we most want to rank for "compare schools". All combinations now collapse onto `/`. Rankings and admissions had none either.
**`/compare?urns=…` was indexable** — ~317 million pairs of parameter space. Now `noindex, follow` with a canonical to the bare path, so its outbound links to each school page still count. Bare `/compare` stays indexable as the landing page for the head term.
**The sitemap** listed every URN with `priority` and `changefreq` (both ignored by Google) and no `lastmod` (which is not). Now drops schools with neither results nor Ofsted, adds `/admissions` which was never listed, and carries a real `lastmod`.
**Split into a sitemap index** with children under `/sitemaps/`, chunked at 10,000. For diagnostics, not size — Search Console reports coverage per submitted sitemap, so one file per family is what will make W2's location pages measurable.
## Two judgement calls worth reviewing
**`lastmod` comes from each school's Ofsted date and is omitted when unknown.** An always-`now` `lastmod` is a claim Google learns to distrust; absent honestly means unknown. On the *index* it legitimately means "when this file changed", so generation time is correct there.
**Children sit under `/sitemaps/`.** Next only treats a whole bracketed path segment as dynamic — verified in Next's router source, where `UrlNode._insert` only reads a segment as dynamic if it `startsWith('[') && endsWith(']')`. A folder named `sitemap-[...parts]` would have been a literal static segment that never matched, and nothing would have errored — the children would just have 404'd in production.
## Verification on the combined tree
Re-run after cherry-picking onto `main`, not carried over from #108:
- Backend: 68 passed
- Frontend: 219 passed, `tsc --noEmit` clean
- e2e: 69 journeys register — #109's Beechwood and Gloucestershire-border tests coexist with W1's canonical and sitemap-index tests
- `next build` green, `/sitemaps/[...parts]` listed as a dynamic route
- Both mart filters intact after the cherry-pick: establishment type **and** #109's Welsh LA code range, in both `dim_school` and `dim_location`
## Expect the URL count to drop
Staging currently reports 25,188 sitemap URLs. After this, expect roughly **21,300 across an index plus three children** — that is the ~3,900 English schools with neither results nor Ofsted no longer being submitted. Not a regression.
## ⚠️ Staging sequencing
No mart changes here, so no Airflow run is strictly required. But the sitemap is cached in-process — hit `POST /api/admin/regenerate-sitemap` after deploy, or the e2e sitemap assertions read a stale flat sitemap and fail.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The apex 301s to www at Cloudflare, but metadataBase, the school-page
canonical, robots.txt's Sitemap: line and the sitemap's own <loc> entries all
named the apex. Every one of those pointed Google at a redirect.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
The homepage read eleven search params and declared no canonical, so every
filter combination was a crawlable near-duplicate of the page we most want to
rank. Rankings and admissions declared none either.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
25,193 schools make ~317 million pairs. The bare page stays indexable as the
landing page for the head term; the parameter space goes noindex, follow so
its outbound links still count.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Drops the schools with neither results nor an Ofsted grade, adds /admissions
which was never listed, replaces the invented priority and changefreq with a
lastmod taken from each school's Ofsted date.
lastmod is omitted where no date is known rather than defaulted to now. An
always-now lastmod is a claim Google learns to distrust; absent honestly
means unknown.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Search Console reports coverage per submitted sitemap, so one file per page
family is what will make W2's location pages measurable when they land. The
index's lastmod is generation time, which is the correct semantic there —
unlike on a <url>, where it would be a claim we cannot support.
Children sit under /sitemaps/ because Next only treats a whole bracketed path
segment as dynamic; a route folder named sitemap-[...parts] would be read as a
literal static segment and never match. Confirmed by the build output, which
lists /sitemaps/[...parts] as a dynamic route.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
This PR splits the single flat sitemap into an index plus chunked child sitemaps (static.xml, schools-N.xml), moves the canonical site origin from the apex domain to www across backend and frontend, and adds canonical/robots metadata for the homepage, rankings, admissions, and compare pages. The change is thorough and well covered by new backend, e2e, and Next.js unit tests, with no severe correctness, security, or deploy issues found.
🟡 Minor
backend/app.py: build_sitemap()'s docstring says it's 'Kept for lifespan and the admin endpoint', but both lifespan() and regenerate_sitemap() now call build_sitemaps() (plural) directly — the singular function is dead/unused outside tests, and the stale comment could mislead future maintainers into thinking it's still load-bearing.
backend/app.py: _school_sitemap_rows sorts by year descending and only evaluates _has_publishable_data on the single latest-year row per URN. A school with no rwm/attainment8/Ofsted data in its most recent year but valid data in an earlier year will be dropped from the sitemap entirely, even though its page has content worth indexing from that earlier year.
## 🤖 AI Code Review (Claude Code)
This PR splits the single flat sitemap into an index plus chunked child sitemaps (static.xml, schools-N.xml), moves the canonical site origin from the apex domain to www across backend and frontend, and adds canonical/robots metadata for the homepage, rankings, admissions, and compare pages. The change is thorough and well covered by new backend, e2e, and Next.js unit tests, with no severe correctness, security, or deploy issues found.
### 🟡 Minor
- **backend/app.py**: build_sitemap()'s docstring says it's 'Kept for lifespan and the admin endpoint', but both lifespan() and regenerate_sitemap() now call build_sitemaps() (plural) directly — the singular function is dead/unused outside tests, and the stale comment could mislead future maintainers into thinking it's still load-bearing.
- **backend/app.py**: _school_sitemap_rows sorts by year descending and only evaluates _has_publishable_data on the single latest-year row per URN. A school with no rwm/attainment8/Ofsted data in its most recent year but valid data in an earlier year will be dropped from the sitemap entirely, even though its page has content worth indexing from that earlier year.
tudor
merged commit f928a15c1e into main2026-08-20 22:23:10 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Re-targets #108, which merged into
feat/england-only-corpusinstead ofmainand so never reached the trunk. Same five commits, cherry-picked onto currentmain(which already carries #107 and #109). No content differences.Spec:
docs/superpowers/specs/2026-08-20-seo-programme-design.md. Plan:docs/superpowers/plans/2026-08-20-w1-crawl-hygiene.md.What changed
Every canonical pointed at a redirect. The apex 301s to
wwwat Cloudflare, butmetadataBase, the school-page canonical,robots.txt'sSitemap:line and every sitemap<loc>named the apex — ~25,000 URLs each costing a hop. Now one sharedSITE_URLinnextjs-app/lib/site.ts, matched byBASE_URLin the backend.The homepage had no canonical at all, while accepting eleven search params, so every filter combination was a crawlable near-duplicate of the page we most want to rank for "compare schools". All combinations now collapse onto
/. Rankings and admissions had none either./compare?urns=…was indexable — ~317 million pairs of parameter space. Nownoindex, followwith a canonical to the bare path, so its outbound links to each school page still count. Bare/comparestays indexable as the landing page for the head term.The sitemap listed every URN with
priorityandchangefreq(both ignored by Google) and nolastmod(which is not). Now drops schools with neither results nor Ofsted, adds/admissionswhich was never listed, and carries a reallastmod.Split into a sitemap index with children under
/sitemaps/, chunked at 10,000. For diagnostics, not size — Search Console reports coverage per submitted sitemap, so one file per family is what will make W2's location pages measurable.Two judgement calls worth reviewing
lastmodcomes from each school's Ofsted date and is omitted when unknown. An always-nowlastmodis a claim Google learns to distrust; absent honestly means unknown. On the index it legitimately means "when this file changed", so generation time is correct there.Children sit under
/sitemaps/. Next only treats a whole bracketed path segment as dynamic — verified in Next's router source, whereUrlNode._insertonly reads a segment as dynamic if itstartsWith('[') && endsWith(']'). A folder namedsitemap-[...parts]would have been a literal static segment that never matched, and nothing would have errored — the children would just have 404'd in production.Verification on the combined tree
Re-run after cherry-picking onto
main, not carried over from #108:tsc --noEmitcleannext buildgreen,/sitemaps/[...parts]listed as a dynamic routedim_schoolanddim_locationExpect the URL count to drop
Staging currently reports 25,188 sitemap URLs. After this, expect roughly 21,300 across an index plus three children — that is the ~3,900 English schools with neither results nor Ofsted no longer being submitted. Not a regression.
⚠️ Staging sequencing
No mart changes here, so no Airflow run is strictly required. But the sitemap is cached in-process — hit
POST /api/admin/regenerate-sitemapafter deploy, or the e2e sitemap assertions read a stale flat sitemap and fail.🤖 Generated with Claude Code
https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
🤖 AI Code Review (Claude Code)
This PR splits the single flat sitemap into an index plus chunked child sitemaps (static.xml, schools-N.xml), moves the canonical site origin from the apex domain to www across backend and frontend, and adds canonical/robots metadata for the homepage, rankings, admissions, and compare pages. The change is thorough and well covered by new backend, e2e, and Next.js unit tests, with no severe correctness, security, or deploy issues found.
🟡 Minor