Compare commits

..
Author SHA1 Message Date
TudorandClaude Opus 5 c0547c45e5 feat(seo): rewrite the C1 snippets to earn the click (W8)
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Failing after 11s
The baseline says these pages already rank and are not clicked. 'compare
school performance' sits at position 6.1 with 0.43% CTR; 'compare schools' at
7.2 with 0.87%. The brand query 'school compare' draws 9.16% from the same
neighbourhood of the same results page, which rules out a ranking explanation
— when the snippet gives a reason to click, it gets clicked.

These SERPs are owned by the DfE's own 'Compare school performance' service.
The old title put a lowercase brand nobody searches for in the most valuable
pixels, then a near-paraphrase of that service's name. Beside the government's
own result it read as a lookalike.

Intent in the title, differentiator in the description. Titles now match what
people type, and the descriptions carry the one fact gov.uk does not publish:
how close you had to live to get a place.

/compare deliberately takes the tool phrasing rather than the homepage's, so
the two pages stop competing for one phrase. The root layout's default and
Open Graph copy were saying something different again; they now agree.

No hard school counts in any of it. The corpus moves with every data refresh
and this repo has already shipped one copy bug of that kind.

Tests guard the mechanics — SERP length, intent keyword, the differentiator,
no brand-first title — and leave the wording free to iterate.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:21:08 +01:00
tudor 4a3928df9f Merge pull request 'fix(seo): a school is publishable on any year's results, not the latest' (#112) from fix/sitemap-any-year-data into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 49s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 0s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m24s
Reviewed-on: #112
2026-08-20 23:15:19 +00:00
TudorandClaude Opus 5 07c97a46c5 fix(seo): a school is publishable on any year's results, not the latest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m2s
PR Checks / Backend Smoke (pull_request) Successful in 8s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 44s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 53s
_school_sitemap_rows tested only the latest year's row, which quietly dropped
every school with results in its history but a null row for the most recent
year — a school that stopped reporting, or whose figures were suppressed for
small-cohort disclosure.

The Mallard Academy (150367) is the case that caught it: real KS2 results for
2015-16 through 2018-19, then null rows from 2022-23 on. Its detail page shows
all four years; the sitemap omitted it. Sampling 40 of the 2,206 excluded
schools found 4 like this, so roughly 220 real pages were being withheld.

Publishable is now a property of the school, computed across every row, while
lastmod still comes from the latest row so the most recent Ofsted date wins.
The field list is a module constant shared with _has_publishable_data so the
two checks cannot drift.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-21 00:04:32 +01:00
tudor bb81337aba Merge pull request 'fix(seo): keep staging out of the search index' (#111) from fix/staging-noindex into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 2m1s
Reviewed-on: #111
2026-08-20 22:39:38 +00:00
TudorandClaude Opus 5 b34511e459 chore: record the branch cleanup manifest
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 16s
PR Checks / Build Frontend (no push) (pull_request) Successful in 45s
PR Checks / Build Pipeline (no push) (pull_request) Successful in 10s
PR Checks / AI Code Review (Claude) (pull_request) Successful in 2m30s
79 remote branches deleted: 77 fully merged into main, plus
feat/seo-crawl-hygiene and feat/england-only-corpus, whose content is
preserved on feat/seo-crawl-hygiene-main (PR #110).

Each line carries the SHA, so any branch can be restored with
  git push origin <sha>:refs/heads/<name>

The 14 branches left standing all carry content that differs from main and
none of them is mine to judge abandoned.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:23:30 +01:00
tudor f928a15c1e Merge pull request 'fix(seo): crawl hygiene and a per-family sitemap index (W1)' (#110) from feat/seo-crawl-hygiene-main into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 18s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 50s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 12s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m21s
Reviewed-on: #110
2026-08-20 22:23:10 +00:00
TudorandClaude Opus 5 1fc1e07d21 fix(seo): keep staging out of the search index
PR Checks / Frontend Typecheck + Tests (pull_request) Successful in 1m3s
PR Checks / Backend Smoke (pull_request) Successful in 7s
PR Checks / Build Backend (no push) (pull_request) Successful in 11s
PR Checks / Build Frontend (no push) (pull_request) Canceled after 13s
PR Checks / Build Pipeline (no push) (pull_request) Canceled after 0s
PR Checks / AI Code Review (Claude) (pull_request) Canceled after 0s
Staging serves the same image as production off stx., with robots.txt saying
Allow: / and no noindex — a fully crawlable duplicate of the site. Nothing
appears indexed today, most likely because the pages canonicalise across to
production, but that is a side effect rather than a control.

X-Robots-Tag, not a robots.txt Disallow. Disallow blocks crawling, which is
not the same as blocking indexing: a disallowed URL can still be indexed from
external links, and blocking the crawl means Google never fetches the page and
so never sees a noindex at all. Staging stays crawlable and answers noindex.

Matched on the staging host explicitly rather than 'any host that is not
production'. The inverted form would cover future environments automatically,
but its failure mode is deindexing production if the Host header ever arrives
rewritten by a proxy — which cannot be verified from here. This form's failure
mode is a new environment being indexable until someone adds it, which is
recoverable. Any new non-production hostname must be added.

The journeys only ever run against staging (deploy.yml passes
STAGING_BASE_URL; promote.yml smoke-polls production without Playwright), so
asserting the header there is safe. The two assertions live in one test
because the halves only work together.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 23:21:58 +01:00
11 changed files with 358 additions and 16 deletions

No files matched your search

+23 -2
View File
@@ -79,6 +79,12 @@ def _school_url(urn: int, school_name: str) -> str:
STATIC_SITEMAP_PATHS = ("/", "/rankings", "/compare", "/admissions")
# A page has something a search result could state if any of these is present
# in any year. Shared by _has_publishable_data and the per-school check in
# _school_sitemap_rows so the two can never drift.
_PUBLISHABLE_FIELDS = ("rwm_expected_pct", "attainment_8_score", "ofsted_grade")
def _has_publishable_data(row) -> bool:
"""True when a school page has something a search result could state.
@@ -87,7 +93,7 @@ def _has_publishable_data(row) -> bool:
signal down, so it stays out of the sitemap. The page itself still resolves
for anyone who has the URL.
"""
for field in ("rwm_expected_pct", "attainment_8_score", "ofsted_grade"):
for field in _PUBLISHABLE_FIELDS:
value = row.get(field)
if value is not None and not pd.isna(value):
return True
@@ -115,6 +121,21 @@ def _school_sitemap_rows(df) -> list[str]:
rows: list[str] = []
seen: set[int] = set()
# Publishable is a property of the SCHOOL, not of its latest row.
#
# The first cut tested the latest year's row alone, which quietly dropped
# every school that has results in its history but a null row for the most
# recent year — a school that stopped reporting, or whose figures were
# suppressed for small-cohort disclosure. The Mallard Academy (150367) is
# the case that caught it: real KS2 results for 2015-16 through 2018-19,
# then null rows for 2022-23 onward. Its page shows all four years; the
# sitemap omitted it. Roughly 220 schools were affected.
publishable_cols = [c for c in _PUBLISHABLE_FIELDS if c in df.columns]
publishable: set[int] = (
set(df.loc[df[publishable_cols].notna().any(axis=1), "urn"].astype(int))
if publishable_cols else set()
)
# Latest row per URN first, so a school's most recent Ofsted date wins.
ordered = df.sort_values("year", ascending=False) if "year" in df.columns else df
@@ -123,7 +144,7 @@ def _school_sitemap_rows(df) -> list[str]:
if urn in seen:
continue
seen.add(urn)
if not _has_publishable_data(row):
if urn not in publishable:
continue
lastmod = None
+41
View File
@@ -176,3 +176,44 @@ def test_children_are_chunked_under_the_limit(monkeypatch):
def test_build_sitemap_still_returns_the_index(sitemap):
# lifespan and the admin endpoint call build_sitemap(); keep it working.
assert "<sitemapindex" in sitemap
def test_school_with_results_in_an_earlier_year_is_still_listed(monkeypatch):
"""Regression: The Mallard Academy (150367).
Real KS2 results 2015-16 to 2018-19, then null rows from 2022-23 onward
because the school stopped reporting. The first cut tested the latest
year's row alone and dropped it, along with ~220 others, even though its
detail page shows all four years of results.
"""
from backend import app as app_module
import pandas as _pd
base = {"local_authority": "Testshire", "school_type": "Academy",
"phase": "Primary", "ofsted_date": None, "ofsted_grade": np.nan,
"attainment_8_score": np.nan, "urn": 150367,
"school_name": "Mallard Academy"}
df = _pd.DataFrame([
{**base, "year": 201819, "rwm_expected_pct": 67.0},
{**base, "year": 202324, "rwm_expected_pct": np.nan},
{**base, "year": 202425, "rwm_expected_pct": np.nan},
])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
xml = app_module.build_sitemaps()["schools-1.xml"]
assert "/school/150367-mallard-academy" in xml
def test_school_with_no_results_in_any_year_is_still_omitted(monkeypatch):
"""The fix must not turn into "list everything"."""
from backend import app as app_module
import pandas as _pd
base = {"local_authority": "Testshire", "school_type": "Academy",
"phase": "Primary", "ofsted_date": None, "ofsted_grade": np.nan,
"attainment_8_score": np.nan, "rwm_expected_pct": np.nan,
"urn": 100002, "school_name": "Ghost Primary"}
df = _pd.DataFrame([{**base, "year": y} for y in (202324, 202425)])
monkeypatch.setattr(app_module, "load_school_data", lambda: df)
assert "/school/100002" not in app_module.build_sitemaps()["schools-1.xml"]
+85
View File
@@ -0,0 +1,85 @@
# Remote branch cleanup, 2026-08-20
# Restore any branch with: git push origin <sha>:refs/heads/<name>
## Deleted: fully merged into main (content is in main)
75677f4252b759ef895e7d5f7c19f8f1745bdb59 add-contact-form-footer
fa1abff642683dfd26ba88a295a0a6710d77147a chore/byline-removal-and-audit-figure
95081d38bdf87764ef5d298676c25fae4cd3b792 chore/remove-parent-view
6877abedebfc1c7d95f1f6ebe68945c62be528ab chore/staged-prod-promotion
090d5f7bec824e083d3252e2c6e636686016304d ci/frontend-checks-speedup
8c3a5cc4e9f551f0190d85357ce7741ad87f3a4c design/cohort-identity
955659580067ca8b81bd77e31e4e2554103f094d feat/allow-analytics-iframe-embed
6828f6cd4417284ea3eb6f088fa20945b8b40ed3 feat/compare-chips-two-per-row
d5cd0abfee226885119665da5d2aa8288b59217f feat/compare-data-foundation
6dd9b04b50bee146682da87efad8fc8b526251c5 feat/compare-frontend-rebuild
96d5fcf5b07b6f175b48e9b20fcb765a320a907f feat/detail-header-details-reveal
f1388ff5bd0af1409823a1e047b7ba84246a0f70 feat/gias-sixth-form-flag
eddf74745f86c9c6d9eb07d867246ff9cc90dc20 feat/hero-artwork-v2
3015c37bac6dc28db58c80fdb9235942c25b83d7 feat/hero-byline
8e763e39d17a4964cf558e51c03f044186371f6d feat/info-popover-tooltip
88c653215d520ab6e902c9de55bf27d86eeb90c3 feat/last-distance-offered
a72323874f7aebdb2e64b6d64a5febd61152d09d feat/last-distance-offered-full
c9a1892bfb0370e0672e5849cba294ddfabe3c65 feat/latest-cutoff-only
4e8df006d75d8be2a1d8529ddf855c445854cba1 feat/near-me-by-search
45ab479062c6a1639facad636fc0cc0cf0fd9155 feat/proposed-to-close-schools
3bf2e8f262cbe058fda6de6f8ea3e050224a51b9 feat/school-detail-visualisations
1f80571b1ff217dc92b660a936c02b5f49d07f0b feat/umami-heatmap-recorder
609bb923d96aa5730131463ea35d5efdd404bf96 feature/ingest-independent-schools
94151c58ea38a9256d66d15161293486a500c7a2 fix/admissions-section-height
79246edc22961c2beb3520437e9d064d3b809d10 fix/annual-dag-ks4-national-selector
59ac9c10b97e0ef1143f57fea06b324e72ac3d4a fix/chart-marker-contrast
9f8dba227c95706ca3527bd48d381e7622cc0a5e fix/compare-chart-refetch-resilience
e74d3882ce78a141fa1a57daa3102d7a58852dc3 fix/compare-expert-fixes
80176cac4db4820e76ea2a156c7a2974bee2f203 fix/compare-final-review-mustfix
f579630fab6c456e26a6984a3e8eebdbe3184202 fix/compare-mockup-drift
dc85254ad2ddf134b4434065d4762cef374d2220 fix/compare-null-year-blanks-chart
43a2c4a6bc539b621f31655aec05ef319a25f343 fix/compare-refresh-and-fetch
d677b5453365b72c81d6df2de62b1fa0d05d684d fix/daily-dag-cache-invalidation
f6bb037c471553e8195b5a8b147467ce0d07a688 fix/detail-all-through
17bd4d5a5eb0b12ca79b97db587f14f7d671e85b fix/detail-chart-truthfulness
4e6be0ce65647410b4ff74f8b763207c28920c26 fix/detail-inclusion-admissions
fdda52ff0af3fae03a4b059a655973cdbb269f91 fix/detail-ofsted-correctness
e36125b24aba254a8d15c5c33b4d2a296e691995 fix/detail-provenance-anchoring
b31e71ac884df6507f567adc046c9fc52d9d310a fix/detail-report-card-render-date
32f8a02862be6a1d49f4c3928b17fcf15d4c94cc fix/detail-trend-chart-taller
e65688d600a86818fe21ae4c61ba27e5b6ec8d7c fix/e2e-brand-assertions
3adea73ee04cdedfab54b0351878b297f72756ad fix/e2e-compare-chips-phase
06e4898c30feaedc471f97aba28ddb0d61379f4d fix/e2e-compare-samephase
9abd020967670a855e80fe5a908c8048a3aa9f14 fix/e2e-distance-locator
acec8135e1ec7c9c3c255e5b23733a0ef862b590 fix/e2e-rankings-year-pick
5944d88f0b1517ef1ef56af1b62270d2cf28e717 fix/expert-signoff-mustfixes
2433101fa08be5df6f170d41790512ffe823d33e fix/font-cascade-and-map-palette
74ca76d150deec6725259d9637ea86d7bb90c683 fix/gias-legacy-fallback
bdaa05cd542f563ef74c45307cd8f7fc465193c9 fix/hero-fallback-and-sharp
d52d384cf23d282b44e9251176f8f3402d808600 fix/hero-map-ios-fullscreen
4043270a77bbe4edb18207fa5f1d94d4747fe12f fix/hero-mobile-and-wording
22e9eb2d48b0d6623e88fd67cb6ca8e5e4074583 fix/homepage-education-accuracy
8d50afef1e8a2b621b7344eadf475b0d609ad7a2 fix/leaflet-specificity-and-font-assertion
b2b2cad5acf534ae7a667d3fb2be15efff37c4a7 fix/list-map-report-card-signal
dc21e80a5e9eebd13aaab84735642eb77cef35e6 fix/mobile-cell-name-size
e5f7f4c959f024c073472122d858333a0f24866c fix/mobile-compare-polish
a00cbe916182d1e04661c750a1ce2e7ad27a68ae fix/mobile-sort-select-overflow
2fd997bfe640c419a6713e85df008463de7f56e8 fix/modal-keyboard-viewport
3e7705756776a0c44d27966dfc972023f1f38b69 fix/ofsted-link-text
ce422e64363e2b03c186ef316832b6de15502d67 fix/promote-status-token
4522cbf64560db1e1cd519e119aa466b42cd16a1 fix/proposed-to-close-copy
15da060e4af37fbae919e0edf25f108266de2585 fix/rankings-admissions-accuracy
6c872ce726f210433354ca38dc5314bc6a467534 fix/rankings-year-validation
1c1df7796194af3d47f8e5ac0a0fbe6f323700e7 fix/report-card-chip-alignment
b2dc4d0779ced02709429afaa5985acf4abda794 fix/results-map-ios-fullscreen
95a5783da1fc994df76cb97238b55596dce4cd8f fix/runtime-api-proxy
8a9ba30cc24e29653a72c03c0e817684b7db7c07 fix/sats-per-level-national
536832a524fc4f9ed858f047e00b9a4d9d429c46 fix/school-detail-nan-500
fef83b3bf244a9bf3cb4afa75dbc433d8725014f fix/schoolbar-sticky-offset
b0c5b6bb57c879477da12c37b15cff950c29ebd3 fix/secondary-anchors-button-affordance
261403bcd2b01aa4f26ee212e26302fc0f769bf9 fix/special-note-full-width
ea5249a2ea6faf6bfa5a1522644387ba59770ec8 fix/standardise-distance-units
e4565e9f158721d4df6b918f2b065de82851f8d9 fix/trends-chart-height
3aad5101a842539105022f9059e85c224fc5973a fix/welsh-establishment-leak
315f1feede70bdf3101d2fdd305b3d06e037fdac perf/batch-supplementary
d2dc78aeb599e16b7ef5019be2b360b08df463bc perf/compare-loading
e098ad4bd1130152705788d4773837b7d1e7112e perf/server-client-split
## Deleted: superseded by PR #110 (content preserved on feat/seo-crawl-hygiene-main)
786ec80dd4de4a3cb674a89e35b3b5e461639289 feat/seo-crawl-hygiene
a5ac0bcd1b37bc10dcbce88f8601d01bf7b3eaaf feat/england-only-corpus
+63
View File
@@ -1695,3 +1695,66 @@ test('a bare /compare is indexable, a parameterised one is not', async ({ page }
.getAttribute('href');
expect(canonical).toBe('https://www.schoolcompare.co.uk/compare');
});
/*
* Staging must not be indexable (spec 2026-08-20, W1 hygiene).
*
* These journeys only ever run against staging — deploy.yml passes
* STAGING_BASE_URL, and promote.yml only smoke-polls production without
* Playwright — so asserting the noindex header here is safe.
*/
test('staging answers noindex, and stays crawlable so the noindex is seen', async ({ page }) => {
const res = await page.request.get('/');
expect(res.ok()).toBeTruthy();
const tag = res.headers()['x-robots-tag'];
expect(tag, 'staging must send X-Robots-Tag').toBeTruthy();
expect(tag).toContain('noindex');
// The other half, and the reason this is one test rather than two: a
// Disallow would stop Google fetching the page at all, so it would never
// see the noindex above. The two only work together.
const robots = await (await page.request.get('/robots.txt')).text();
expect(robots).not.toMatch(/^\s*Disallow:\s*\/\s*$/mi);
});
test('a school page on staging is noindexed too, not just the homepage', async ({ page }) => {
const list = await page.request.get('/api/schools?search=primary&per_page=1');
const [first] = (await list.json()).schools ?? [];
expect(first, 'no school available').toBeTruthy();
const res = await page.request.get(`/school/${first.urn}-x`);
expect(res.headers()['x-robots-tag']).toContain('noindex');
});
/*
* W8 — the C1 pages must ship a description, and it must differentiate.
*
* Baseline was 0.43% CTR at position 6.1 on "compare school performance",
* against 9.16% for the brand query from the same neighbourhood. The SERP is
* owned by the DfE's own service, so a description that paraphrases it earns
* nothing. Google may rewrite a snippet, but it cannot use one we never sent.
*/
test('every C1 page ships a description, and none opens its title with the brand', async ({ page }) => {
for (const path of ['/', '/compare', '/rankings', '/admissions']) {
await page.goto(path);
const desc = await page.locator('meta[name="description"]').first()
.getAttribute('content');
expect(desc, `${path} must ship a description`).toBeTruthy();
expect(desc!.length, `${path} description too short to be worth reading`)
.toBeGreaterThan(100);
const title = await page.title();
expect(title.toLowerCase().startsWith('schoolcompare'),
`${path} spends its most valuable pixels on the brand`).toBe(false);
}
});
test('the homepage snippet names what gov.uk does not publish', async ({ page }) => {
await page.goto('/');
const desc = await page.locator('meta[name="description"]').first()
.getAttribute('content');
// Admissions distance is the one fact the DfE service has no equivalent for.
expect(desc).toMatch(/close you had to live|distance/i);
});
+77
View File
@@ -51,3 +51,80 @@ describe('/compare indexability', () => {
.toBe('https://www.schoolcompare.co.uk/compare');
});
});
/*
* W8 — snippet copy for the C1 cluster.
*
* The baseline (GSC, 16 months to 2026-08-20) showed these pages ranking on
* page one and converting at a tenth of the normal rate: "compare school
* performance" at position 6.1 with 0.43% CTR, against 9.16% for the brand
* query from the same neighbourhood. The SERP is dominated by the DfE's own
* "Compare school performance" service, so the job of this copy is to say
* what that service does not offer, without losing intent match on the title.
*
* These tests guard the mechanics that make a snippet work — length, intent
* keyword, differentiator, no brand-first — not the exact wording, which
* should stay free to iterate.
*/
// Google truncates titles near 60 characters and descriptions near 155.
const TITLE_MAX = 60;
const DESC_MIN = 110;
const DESC_MAX = 155;
type Meta = { title?: unknown; description?: unknown };
const titleOf = (m: Meta): string => {
const t = m.title as string | { absolute?: string } | undefined;
return typeof t === 'string' ? t : (t?.absolute ?? '');
};
describe('C1 snippet copy', () => {
const pages: Array<[string, Meta, RegExp]> = [
['home', homeMetadata as Meta, /compare schools/i],
['rankings', rankingsMetadata as Meta, /league table/i],
['admissions', admissionsMetadata as Meta, /admission/i],
];
for (const [name, meta, intent] of pages) {
it(`${name}: title carries the search intent and fits the SERP`, () => {
const t = titleOf(meta);
expect(t).toMatch(intent);
expect(t.length).toBeLessThanOrEqual(TITLE_MAX);
});
it(`${name}: title does not open with the brand`, () => {
// The measured 0.43% CTR came from a brand-first title. The most
// valuable pixels go to the thing the searcher typed.
expect(titleOf(meta).toLowerCase().startsWith('schoolcompare')).toBe(false);
});
it(`${name}: description is long enough to be worth reading, short enough to survive`, () => {
const d = meta.description as string;
expect(d.length).toBeGreaterThanOrEqual(DESC_MIN);
expect(d.length).toBeLessThanOrEqual(DESC_MAX);
});
}
it('the homepage description names what gov.uk does not publish', () => {
// Admissions distance is the one fact the DfE service has no equivalent
// for. If it ever leaves this description, the snippet is competing with
// gov.uk on gov.uk's own ground.
expect(homeMetadata.description).toMatch(/close you had to live|distance/i);
});
it('/compare targets the tool phrasing rather than repeating the homepage', () => {
// Two pages chasing one phrase is how a site competes with itself.
return compareMetadata({ searchParams: Promise.resolve({}) }).then((m) => {
expect(m.title).toMatch(/comparison tool/i);
expect(m.title).not.toBe(titleOf(homeMetadata as Meta));
});
});
it('no C1 page claims a school count that will drift', () => {
// The corpus moves with every data refresh; this repo has already shipped
// one copy bug of that kind ("three schools" against MAX_SCHOOLS = 5).
for (const [, meta] of pages) {
expect(meta.description as string).not.toMatch(/\b\d{2},\d{3}\b|\b\d{2},000\b/);
}
});
});
+4 -2
View File
@@ -5,9 +5,11 @@ import { AdmissionsView } from '@/components/AdmissionsView';
export const dynamic = 'force-static';
export const metadata: Metadata = {
title: 'School Admissions Guide',
// Deadlines and offer days are what gets searched, and what this page is
// genuinely best at — the countdowns are live.
title: { absolute: 'School Admissions Deadlines & Offer Days | schoolcompare' },
description:
'Understand the Primary and Secondary school admissions process in England, with live countdowns to every key deadline and National Offer Day.',
'Every key date for primary and secondary school admissions in England, with live countdowns to the application deadline and National Offer Day.',
alternates: { canonical: absoluteUrl('/admissions') },
};
+5 -2
View File
@@ -30,9 +30,12 @@ export async function generateMetadata(
const { urns } = await searchParams;
const base: Metadata = {
title: 'Compare Schools',
// Deliberately not the homepage's phrase. Two pages chasing "compare
// schools" is how a site competes with itself; this one takes the tool
// phrasing instead.
title: 'School Comparison Tool — Up to Five at Once | schoolcompare',
description:
'Compare schools in England side by side — Ofsted inspections, KS2 and GCSE results against the England average, admissions odds and school community.',
'Put up to five English schools in one table: SATs and GCSE results against the England average, Ofsted grades, and the distance places were offered.',
keywords:
'school comparison, compare schools, Ofsted comparison, school admissions, KS2 comparison, primary school performance',
alternates: { canonical: absoluteUrl('/compare') },
+9 -6
View File
@@ -48,10 +48,11 @@ export const metadata: Metadata = {
statusBarStyle: 'default',
},
title: {
default: 'schoolcompare | Compare School Performance',
default: 'Compare Schools Side by Side | schoolcompare',
template: '%s | schoolcompare',
},
description: 'Compare primary and secondary school SATs and GCSE performance across England',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place. Free, no sign-up.',
keywords: 'school comparison, KS2 results, KS4 results, primary school, secondary school, England schools, SATs results, GCSE results',
authors: [{ name: 'schoolcompare' }],
manifest: '/manifest.json',
@@ -61,16 +62,18 @@ export const metadata: Metadata = {
metadataBase: new URL(SITE_URL),
openGraph: {
type: 'website',
title: 'schoolcompare | Compare School Performance',
description: 'Compare primary and secondary school SATs and GCSE performance across England',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
url: SITE_URL,
siteName: 'schoolcompare',
},
twitter: {
// summary_large_image now that there is an image worth showing.
card: 'summary_large_image',
title: 'schoolcompare | Compare School Performance',
description: 'Compare primary and secondary school SATs and GCSE performance across England',
title: 'Compare Schools Side by Side | schoolcompare',
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place.',
},
};
+15 -2
View File
@@ -34,8 +34,21 @@ interface HomePageProps {
* saying the brand twice.
*/
export const metadata: Metadata = {
title: { absolute: 'schoolcompare | Compare every school in England' },
description: 'Search and compare school performance across England',
/*
* Intent in the title, differentiator in the description.
*
* These queries are owned by the DfE's own "Compare school performance"
* service, and the old title — brand first, then a near-paraphrase of that
* service's name — gave a searcher no reason to pick us over it. It drew
* 0.43% CTR at position 6.1 while the brand query drew 9.16% from the same
* neighbourhood, so the ranking was never the problem.
*
* The title now matches what people type. The description carries the one
* fact gov.uk does not publish: how close you had to live to get a place.
*/
title: { absolute: 'Compare Schools Side by Side | schoolcompare' },
description:
'Put five English schools on one screen — SATs, GCSE results, Ofsted grades, and how close you had to live to get a place. Free, no sign-up.',
// This page reads eleven search params. They filter a result set; they do
// not make a new document. Collapsing every combination onto "/" stops the
// homepage competing with itself for its own head terms.
+5 -2
View File
@@ -18,8 +18,11 @@ interface RankingsPageProps {
}
export const metadata: Metadata = {
title: 'School Rankings',
description: 'Top-ranked schools by SATs and GCSE performance across England',
// 'School Rankings' matched nothing anyone types. League tables is the
// phrase parents actually search, and it spikes each results day.
title: { absolute: 'Primary & Secondary School League Tables | schoolcompare' },
description:
'Rank English schools by SATs results, GCSEs, Progress 8 or Attainment 8, and filter by local authority or year. Built from the DfE’s own figures.',
keywords: 'school rankings, top schools, best schools, KS2 rankings, KS4 rankings, school league tables',
// Param forms (?metric=&local_authority=&year=&phase=) collapse here for
// now. W3 replaces them with real indexable paths.
+31
View File
@@ -55,6 +55,37 @@ const nextConfig = {
// Headers for caching and security
async headers() {
return [
{
/*
* Keep non-production hosts out of the index.
*
* Staging serves the same image as production off stx., so without
* this it is a full crawlable duplicate of the site.
*
* X-Robots-Tag, NOT a robots.txt Disallow. Disallow blocks crawling,
* which is not the same as blocking indexing — a disallowed URL can
* still be indexed from external links, and worse, blocking the crawl
* means Google never fetches the page and never sees a noindex at all.
* Staging therefore stays crawlable and answers "noindex" when crawled.
*
* Matched on the staging host explicitly rather than "any host that is
* not production". The inverted form is tempting because it would cover
* future environments automatically, but its failure mode is
* deindexing production if the Host header ever arrives rewritten by a
* proxy. This form's failure mode is a new environment being indexable
* until someone adds it here — recoverable, where the other is not.
*
* Any new non-production hostname must be added to this list.
*/
source: '/:path*',
has: [{ type: 'host', value: 'stx.schoolcompare.co.uk' }],
headers: [
{
key: 'X-Robots-Tag',
value: 'noindex, nofollow',
},
],
},
{
source: '/:path*',
headers: [