fix(seo): keep staging out of the search index #111

Merged
tudor merged 2 commits from fix/staging-noindex into main 2026-08-20 22:39:38 +00:00
Owner

Staging serves the same image as production off stx., with robots.txt saying Allow: / and no noindex — a fully crawlable duplicate of the site. Nothing appears indexed today, most likely because staging's pages canonicalise across to production, but that is a side effect rather than a control, and the surface area multiplies once W2's location pages land.

X-Robots-Tag, not a robots.txt Disallow

Worth being explicit, because Disallow: / is the intuitive answer and it is the wrong one.

Disallow blocks crawling, which is not the same as blocking indexing. A disallowed URL can still be indexed from external links — Search Console reports these as "Indexed, though blocked by robots.txt". Worse, blocking the crawl means Google never fetches the page and therefore never sees a noindex directive at all, so the two are actively counterproductive together.

Staging therefore stays crawlable and answers X-Robots-Tag: noindex, nofollow when crawled. There is a test asserting both halves in one case, because they only work together.

Why the host is matched explicitly

The rule matches stx.schoolcompare.co.uk by name rather than "any host that is not production".

The inverted form is tempting — it would cover any future preview or staging environment automatically. But its failure mode is deindexing production if the Host header ever arrives rewritten by a proxy, and I cannot verify from here what Host actually reaches the origin behind Cloudflare and the macvlan setup. This form's failure mode is a new environment being indexable until someone adds it to the list: recoverable, where the other is not.

Any new non-production hostname must be added to that list. The comment in next.config.js says so.

Mechanism

headers() is compiled into routes-manifest.json at build time, but has: [{ type: 'host' }] is evaluated per request — so one image still serves both environments, which is the constraint the rest of the config is built around. Verified by inspecting the built manifest:

{
 "source": "/:path*",
 "has": [{ "type": "host", "value": "stx.schoolcompare.co.uk" }],
 "headers": [{ "key": "X-Robots-Tag", "value": "noindex, nofollow" }],
 "regex": "^(?:/((?:[^/]+?)(?:/(?:[^/]+?))*))?(?:/)?$"
}

No middleware, so no per-request hop.

Safety of the e2e assertion

The journeys only ever run against staging — deploy.yml passes STAGING_BASE_URL, and promote.yml only smoke-polls production without Playwright. So asserting the header cannot break promotion.

🤖 Generated with Claude Code

https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj

Staging serves the same image as production off `stx.`, with `robots.txt` saying `Allow: /` and no noindex — a fully crawlable duplicate of the site. Nothing appears indexed today, most likely because staging's pages canonicalise across to production, but that is a side effect rather than a control, and the surface area multiplies once W2's location pages land. ## X-Robots-Tag, not a robots.txt Disallow Worth being explicit, because `Disallow: /` is the intuitive answer and it is the wrong one. `Disallow` blocks **crawling**, which is not the same as blocking **indexing**. A disallowed URL can still be indexed from external links — Search Console reports these as "Indexed, though blocked by robots.txt". Worse, blocking the crawl means Google never fetches the page and therefore never sees a noindex directive at all, so the two are actively counterproductive together. Staging therefore stays crawlable and answers `X-Robots-Tag: noindex, nofollow` when crawled. There is a test asserting both halves in one case, because they only work together. ## Why the host is matched explicitly The rule matches `stx.schoolcompare.co.uk` by name rather than "any host that is not production". The inverted form is tempting — it would cover any future preview or staging environment automatically. But its failure mode is **deindexing production** if the `Host` header ever arrives rewritten by a proxy, and I cannot verify from here what Host actually reaches the origin behind Cloudflare and the macvlan setup. This form's failure mode is a new environment being indexable until someone adds it to the list: recoverable, where the other is not. **Any new non-production hostname must be added to that list.** The comment in `next.config.js` says so. ## Mechanism `headers()` is compiled into `routes-manifest.json` at build time, but `has: [{ type: 'host' }]` is evaluated per request — so one image still serves both environments, which is the constraint the rest of the config is built around. Verified by inspecting the built manifest: ```json { "source": "/:path*", "has": [{ "type": "host", "value": "stx.schoolcompare.co.uk" }], "headers": [{ "key": "X-Robots-Tag", "value": "noindex, nofollow" }], "regex": "^(?:/((?:[^/]+?)(?:/(?:[^/]+?))*))?(?:/)?$" } ``` No middleware, so no per-request hop. ## Safety of the e2e assertion The journeys only ever run against staging — `deploy.yml` passes `STAGING_BASE_URL`, and `promote.yml` only smoke-polls production without Playwright. So asserting the header cannot break promotion. 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
tudor added 1 commit 2026-08-20 22:21:59 +00:00
fix(seo): keep staging out of the search index
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 13s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
1fc1e07d21
Staging serves the same image as production off stx., with robots.txt saying
Allow: / and no noindex — a fully crawlable duplicate of the site. Nothing
appears indexed today, most likely because the pages canonicalise across to
production, but that is a side effect rather than a control.

X-Robots-Tag, not a robots.txt Disallow. Disallow blocks crawling, which is
not the same as blocking indexing: a disallowed URL can still be indexed from
external links, and blocking the crawl means Google never fetches the page and
so never sees a noindex at all. Staging stays crawlable and answers noindex.

Matched on the staging host explicitly rather than 'any host that is not
production'. The inverted form would cover future environments automatically,
but its failure mode is deindexing production if the Host header ever arrives
rewritten by a proxy — which cannot be verified from here. This form's failure
mode is a new environment being indexable until someone adds it, which is
recoverable. Any new non-production hostname must be added.

The journeys only ever run against staging (deploy.yml passes
STAGING_BASE_URL; promote.yml smoke-polls production without Playwright), so
asserting the header there is safe. The two assertions live in one test
because the halves only work together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
tudor added 1 commit 2026-08-20 22:23:32 +00:00
chore: record the branch cleanup manifest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m30s
b34511e459
79 remote branches deleted: 77 fully merged into main, plus
feat/seo-crawl-hygiene and feat/england-only-corpus, whose content is
preserved on feat/seo-crawl-hygiene-main (PR #110).

Each line carries the SHA, so any branch can be restored with
  git push origin <sha>:refs/heads/<name>

The 14 branches left standing all carry content that differs from main and
none of them is mine to judge abandoned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj

🤖 AI Code Review (Claude Code)

This PR adds a documentation manifest of deleted/merged branches (informational only), plus a runtime fix that makes staging (stx.schoolcompare.co.uk) send X-Robots-Tag: noindex,nofollow via next.config.js headers() while staying crawlable (no robots.txt Disallow), with two new Playwright e2e tests asserting the header on the homepage and a school detail page. The implementation reasoning is sound (host-matched header, not an inverted 'not production' match, avoiding the Disallow-hides-noindex trap) and the staging hostname/architecture (same image promoted to both environments) checks out against other docs in the repo.

🟡 Minor

  • e2e/tests/journeys.spec.ts: The two new staging-noindex tests have no environment guard and will fail under the default/local BASE_URL (http://localhost:3000), unlike the rest of the suite; only an issue for local ad-hoc runs since CI always targets staging.
## 🤖 AI Code Review (Claude Code) This PR adds a documentation manifest of deleted/merged branches (informational only), plus a runtime fix that makes staging (stx.schoolcompare.co.uk) send X-Robots-Tag: noindex,nofollow via next.config.js headers() while staying crawlable (no robots.txt Disallow), with two new Playwright e2e tests asserting the header on the homepage and a school detail page. The implementation reasoning is sound (host-matched header, not an inverted 'not production' match, avoiding the Disallow-hides-noindex trap) and the staging hostname/architecture (same image promoted to both environments) checks out against other docs in the repo. ### 🟡 Minor - **e2e/tests/journeys.spec.ts**: The two new staging-noindex tests have no environment guard and will fail under the default/local BASE_URL (http://localhost:3000), unlike the rest of the suite; only an issue for local ad-hoc runs since CI always targets staging.
tudor merged commit bb81337aba into main 2026-08-20 22:39:38 +00:00
Sign in to join this conversation.
No Reviewers
No labels
2 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: tudor/school_compare#111