Staging serves the same image as production off stx., with robots.txt saying Allow: / and no noindex — a fully crawlable duplicate of the site. Nothing appears indexed today, most likely because staging's pages canonicalise across to production, but that is a side effect rather than a control, and the surface area multiplies once W2's location pages land.
X-Robots-Tag, not a robots.txt Disallow
Worth being explicit, because Disallow: / is the intuitive answer and it is the wrong one.
Disallow blocks crawling, which is not the same as blocking indexing. A disallowed URL can still be indexed from external links — Search Console reports these as "Indexed, though blocked by robots.txt". Worse, blocking the crawl means Google never fetches the page and therefore never sees a noindex directive at all, so the two are actively counterproductive together.
Staging therefore stays crawlable and answers X-Robots-Tag: noindex, nofollow when crawled. There is a test asserting both halves in one case, because they only work together.
Why the host is matched explicitly
The rule matches stx.schoolcompare.co.uk by name rather than "any host that is not production".
The inverted form is tempting — it would cover any future preview or staging environment automatically. But its failure mode is deindexing production if the Host header ever arrives rewritten by a proxy, and I cannot verify from here what Host actually reaches the origin behind Cloudflare and the macvlan setup. This form's failure mode is a new environment being indexable until someone adds it to the list: recoverable, where the other is not.
Any new non-production hostname must be added to that list. The comment in next.config.js says so.
Mechanism
headers() is compiled into routes-manifest.json at build time, but has: [{ type: 'host' }] is evaluated per request — so one image still serves both environments, which is the constraint the rest of the config is built around. Verified by inspecting the built manifest:
The journeys only ever run against staging — deploy.yml passes STAGING_BASE_URL, and promote.yml only smoke-polls production without Playwright. So asserting the header cannot break promotion.
Staging serves the same image as production off `stx.`, with `robots.txt` saying `Allow: /` and no noindex — a fully crawlable duplicate of the site. Nothing appears indexed today, most likely because staging's pages canonicalise across to production, but that is a side effect rather than a control, and the surface area multiplies once W2's location pages land.
## X-Robots-Tag, not a robots.txt Disallow
Worth being explicit, because `Disallow: /` is the intuitive answer and it is the wrong one.
`Disallow` blocks **crawling**, which is not the same as blocking **indexing**. A disallowed URL can still be indexed from external links — Search Console reports these as "Indexed, though blocked by robots.txt". Worse, blocking the crawl means Google never fetches the page and therefore never sees a noindex directive at all, so the two are actively counterproductive together.
Staging therefore stays crawlable and answers `X-Robots-Tag: noindex, nofollow` when crawled. There is a test asserting both halves in one case, because they only work together.
## Why the host is matched explicitly
The rule matches `stx.schoolcompare.co.uk` by name rather than "any host that is not production".
The inverted form is tempting — it would cover any future preview or staging environment automatically. But its failure mode is **deindexing production** if the `Host` header ever arrives rewritten by a proxy, and I cannot verify from here what Host actually reaches the origin behind Cloudflare and the macvlan setup. This form's failure mode is a new environment being indexable until someone adds it to the list: recoverable, where the other is not.
**Any new non-production hostname must be added to that list.** The comment in `next.config.js` says so.
## Mechanism
`headers()` is compiled into `routes-manifest.json` at build time, but `has: [{ type: 'host' }]` is evaluated per request — so one image still serves both environments, which is the constraint the rest of the config is built around. Verified by inspecting the built manifest:
```json
{
"source": "/:path*",
"has": [{ "type": "host", "value": "stx.schoolcompare.co.uk" }],
"headers": [{ "key": "X-Robots-Tag", "value": "noindex, nofollow" }],
"regex": "^(?:/((?:[^/]+?)(?:/(?:[^/]+?))*))?(?:/)?$"
}
```
No middleware, so no per-request hop.
## Safety of the e2e assertion
The journeys only ever run against staging — `deploy.yml` passes `STAGING_BASE_URL`, and `promote.yml` only smoke-polls production without Playwright. So asserting the header cannot break promotion.
🤖 Generated with [Claude Code](https://claude.com/claude-code)
https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
Staging serves the same image as production off stx., with robots.txt saying
Allow: / and no noindex — a fully crawlable duplicate of the site. Nothing
appears indexed today, most likely because the pages canonicalise across to
production, but that is a side effect rather than a control.
X-Robots-Tag, not a robots.txt Disallow. Disallow blocks crawling, which is
not the same as blocking indexing: a disallowed URL can still be indexed from
external links, and blocking the crawl means Google never fetches the page and
so never sees a noindex at all. Staging stays crawlable and answers noindex.
Matched on the staging host explicitly rather than 'any host that is not
production'. The inverted form would cover future environments automatically,
but its failure mode is deindexing production if the Host header ever arrives
rewritten by a proxy — which cannot be verified from here. This form's failure
mode is a new environment being indexable until someone adds it, which is
recoverable. Any new non-production hostname must be added.
The journeys only ever run against staging (deploy.yml passes
STAGING_BASE_URL; promote.yml smoke-polls production without Playwright), so
asserting the header there is safe. The two assertions live in one test
because the halves only work together.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
79 remote branches deleted: 77 fully merged into main, plus
feat/seo-crawl-hygiene and feat/england-only-corpus, whose content is
preserved on feat/seo-crawl-hygiene-main (PR #110).
Each line carries the SHA, so any branch can be restored with
git push origin <sha>:refs/heads/<name>
The 14 branches left standing all carry content that differs from main and
none of them is mine to judge abandoned.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
This PR adds a documentation manifest of deleted/merged branches (informational only), plus a runtime fix that makes staging (stx.schoolcompare.co.uk) send X-Robots-Tag: noindex,nofollow via next.config.js headers() while staying crawlable (no robots.txt Disallow), with two new Playwright e2e tests asserting the header on the homepage and a school detail page. The implementation reasoning is sound (host-matched header, not an inverted 'not production' match, avoiding the Disallow-hides-noindex trap) and the staging hostname/architecture (same image promoted to both environments) checks out against other docs in the repo.
🟡 Minor
e2e/tests/journeys.spec.ts: The two new staging-noindex tests have no environment guard and will fail under the default/local BASE_URL (http://localhost:3000), unlike the rest of the suite; only an issue for local ad-hoc runs since CI always targets staging.
## 🤖 AI Code Review (Claude Code)
This PR adds a documentation manifest of deleted/merged branches (informational only), plus a runtime fix that makes staging (stx.schoolcompare.co.uk) send X-Robots-Tag: noindex,nofollow via next.config.js headers() while staying crawlable (no robots.txt Disallow), with two new Playwright e2e tests asserting the header on the homepage and a school detail page. The implementation reasoning is sound (host-matched header, not an inverted 'not production' match, avoiding the Disallow-hides-noindex trap) and the staging hostname/architecture (same image promoted to both environments) checks out against other docs in the repo.
### 🟡 Minor
- **e2e/tests/journeys.spec.ts**: The two new staging-noindex tests have no environment guard and will fail under the default/local BASE_URL (http://localhost:3000), unlike the rest of the suite; only an issue for local ad-hoc runs since CI always targets staging.
tudor
merged commit bb81337aba into main2026-08-20 22:39:38 +00:00
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Staging serves the same image as production off
stx., withrobots.txtsayingAllow: /and no noindex — a fully crawlable duplicate of the site. Nothing appears indexed today, most likely because staging's pages canonicalise across to production, but that is a side effect rather than a control, and the surface area multiplies once W2's location pages land.X-Robots-Tag, not a robots.txt Disallow
Worth being explicit, because
Disallow: /is the intuitive answer and it is the wrong one.Disallowblocks crawling, which is not the same as blocking indexing. A disallowed URL can still be indexed from external links — Search Console reports these as "Indexed, though blocked by robots.txt". Worse, blocking the crawl means Google never fetches the page and therefore never sees a noindex directive at all, so the two are actively counterproductive together.Staging therefore stays crawlable and answers
X-Robots-Tag: noindex, nofollowwhen crawled. There is a test asserting both halves in one case, because they only work together.Why the host is matched explicitly
The rule matches
stx.schoolcompare.co.ukby name rather than "any host that is not production".The inverted form is tempting — it would cover any future preview or staging environment automatically. But its failure mode is deindexing production if the
Hostheader ever arrives rewritten by a proxy, and I cannot verify from here what Host actually reaches the origin behind Cloudflare and the macvlan setup. This form's failure mode is a new environment being indexable until someone adds it to the list: recoverable, where the other is not.Any new non-production hostname must be added to that list. The comment in
next.config.jssays so.Mechanism
headers()is compiled intoroutes-manifest.jsonat build time, buthas: [{ type: 'host' }]is evaluated per request — so one image still serves both environments, which is the constraint the rest of the config is built around. Verified by inspecting the built manifest:No middleware, so no per-request hop.
Safety of the e2e assertion
The journeys only ever run against staging —
deploy.ymlpassesSTAGING_BASE_URL, andpromote.ymlonly smoke-polls production without Playwright. So asserting the header cannot break promotion.🤖 Generated with Claude Code
https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
🤖 AI Code Review (Claude Code)
This PR adds a documentation manifest of deleted/merged branches (informational only), plus a runtime fix that makes staging (stx.schoolcompare.co.uk) send X-Robots-Tag: noindex,nofollow via next.config.js headers() while staying crawlable (no robots.txt Disallow), with two new Playwright e2e tests asserting the header on the homepage and a school detail page. The implementation reasoning is sound (host-matched header, not an inverted 'not production' match, avoiding the Disallow-hides-noindex trap) and the staging hostname/architecture (same image promoted to both environments) checks out against other docs in the repo.
🟡 Minor