diff --git a/docs/superpowers/specs/2026-08-20-seo-programme-design.md b/docs/superpowers/specs/2026-08-20-seo-programme-design.md new file mode 100644 index 0000000..1986897 --- /dev/null +++ b/docs/superpowers/specs/2026-08-20-seo-programme-design.md @@ -0,0 +1,309 @@ +# SEO Programme — Design + +Date: 2026-08-20 +Status: awaiting review + +## Problem + +schoolcompare ranks second for "school compare" — an exact match for the +brand and the domain. It ranks poorly for "compare schools", "school +comparison" and "schools near me". The first is a naming artefact and +transfers to nothing; the rest are the queries that actually carry parent +demand. + +The cause is structural, not editorial. The site publishes five route +families: + + / /compare /rankings /admissions /school/[slug] + +Location intent has no landing page at all. Every competitor outranking us +on those queries wins with programmatic location pages: + +| Competitor | URL pattern | +|---------------|--------------------------------------------| +| School Guide | `/best-schools-in/manchester` | +| Locrating | `/the-best-primary-schools-in-Manchester_…`| +| FindMySchool | `/best-primary-schools/manchester` | +| Snobe | `/best-primary-schools/manchester` | +| School Atlas | `/guides/best-primary-schools-manchester` | + +"Schools near me" is a local-intent query. Google resolves it against the +user's coordinates and serves pages that are *about a place*. A national +homepage cannot win it. No title or description change fixes this; only +pages Google can localise will. + +## Baseline (measured 2026-08-20, production API) + +| Measure | Value | +|--------------------------------------------|---------| +| Unique schools | 27,229 | +| URLs in sitemap.xml | 27,232 | +| Schools with 2024/25 performance data | 21,266 | +| Schools with **no** current performance data| ~5,963 (22%) | +| Welsh establishments (all metrics null) | 1,569 | +| Overseas / offshore establishments | 467 | +| Static URLs in sitemap | 3 | +| Routes setting a canonical | 1 of 5 | + +Three findings from that table drive the plan. + +**We submit ~6,000 thin pages to Google.** `build_sitemap()` +(`backend/app.py:73`) enumerates every URN regardless of whether the school +has any data. Welsh establishments return `school_type: "Welsh +establishment"` with every performance metric, Ofsted grade and phase field +null. Overseas and offshore establishments ("BFPO Overseas Establishments", +"Gibraltar Overseas Establishments", "Jersey Offshore Establishments") are +in the local-authority list too. At 22% of the submitted corpus this is a +site-wide quality signal problem and a crawl-budget waste, not a rounding +error. + +W1 item 4 removes 2,036 of those — every non-England establishment — taking +the corpus to 25,193. The 3,927 that remain are English schools with no +current data: mostly newly opened, special, nursery or alternative provision. +Those are a template problem, not a corpus problem, and item 5 handles them +separately. + +**The homepage is its own competitor.** `app/page.tsx` accepts eleven search +params (`search`, `local_authority`, `school_type`, `phase`, `page`, +`postcode`, `radius`, `sort`, `gender`, `admissions_policy`, +`has_sixth_form`) and sets no canonical. Every filter combination is a +crawlable near-duplicate of the single page we are asking to rank for +"compare schools". + +**School pages are near-orphans.** Reachable from the sitemap and from site +search, but almost nothing links to them contextually, so they accrue no +internal authority. + +Also noted: the sitemap emits invented `priority` values and no `lastmod`. +Google ignores `priority` and `changefreq` entirely; `lastmod` is the field +it does read, and we omit it. + +## Keyword clusters + +Ranked by judgement of UK parent search behaviour and by the competitive +SERP evidence above. Google Search Console is connected, so cluster +priorities are to be re-derived from measured impressions before build +starts (see Workstream 0). + +**C1 — Head "compare" terms.** compare schools · school comparison · school +comparison tool · compare school performance · compare primary schools · +compare secondary schools · compare two schools + +**C2 — League tables and rankings.** primary school league tables · +secondary school league tables · school league tables 2026 · SATs results by +school · GCSE results by school · KS2 league tables · Progress 8 rankings · +best primary schools in [town] · top 10 primary schools in [LA] + +**C3 — Local / near me.** schools near me · primary schools near me · +secondary schools near me · best schools near me · good schools near me · +schools in [town] · primary schools in [LA] · schools near [postcode] · +[postcode] school catchment + +**C4 — Individual school long tail.** [school] ofsted · [school] SATs +results · [school] catchment area · [school] reviews · [school] URN + +**C5 — Admissions.** primary school admissions 2027 · national offer day +2027 · school application deadline · school admissions appeal · +oversubscription criteria · distance criteria school admissions · didn't get +first choice school · school admissions [LA] + +**C6 — Metric explainers.** what is a good SATs score · what is Progress 8 · +what is Attainment 8 · expected standard KS2 meaning · scaled score +explained · Ofsted grades explained · Ofsted report cards · pupil premium +explained + +**C7 — Head to head.** [school A] vs [school B] · academy vs community +school · grammar school vs comprehensive · faith school vs community school + +## Workstreams + +### W0 — Measure before touching anything + +Export a Google Search Console baseline: impressions, clicks, average +position and CTR by query and by page, for the trailing 16 months. Bucket +queries into C1–C7. This sets the counterfactual — without it, no later +claim about lift is defensible, because school-search traffic is strongly +seasonal (results day in December, offer day in March/April). + +Re-rank C1–C7 against measured impressions and adjust the sequence below if +the data disagrees with the judgement calls. + +### W1 — Crawl hygiene and index sanity + +Cheap, and it unblocks everything after it. Adding 5,000 pages on top of a +corpus that is 22% thin would compound the existing problem. + +1. Canonical on every route. `/`, `/rankings`, `/compare` and `/admissions` + currently set none. +2. The homepage canonicalises to `/` regardless of search params. +3. `/compare?urns=…` gets `noindex, follow` — it is an unbounded parameter + space with no standalone value. +4. **England only — DONE.** Wales, the Crown Dependencies, Gibraltar and the + service/overseas schools are removed from the corpus at the mart boundary, + not hidden at the view layer. `dim_school` and `dim_location` both exclude + `TypeOfEstablishment` in {25, 26, 30, 37} — Offshore schools, Service + children's education, Welsh establishment, British schools overseas — + listed once as `vars.non_england_school_type_codes` in `dbt_project.yml`. + That removes 2,036 establishments and 29 local authorities, and because + `build_sitemap()` reads the same marts, it drops them from the sitemap in + the same stroke. `assert_england_only_schools` fails the pipeline if a GIAS + refresh reintroduces them or if the two models drift apart. +5. Prune the remaining thin pages: exclude any school with no performance data + **and** no Ofsted record. Distinct from item 4 — these are English schools + with nothing yet to show, so the fix may be a better template rather than + removal. +6. Rebuild the sitemap as a sitemap **index**: one child per page family, + real `lastmod` from the data-load timestamp, `priority` and `changefreq` + dropped. + +### W2 — The location layer + +The dominant lever. `dim_location` already carries `town`, `county`, +`local_authority_name`, `parliamentary_constituency`, `latitude`, +`longitude` and `postcode`, so no new ingestion is required. + +Routes: + + /schools/[la] e.g. /schools/manchester + /schools/[la]/primary + /schools/[la]/secondary + /best-primary-schools/[town] + /best-secondary-schools/[town] + /schools/near/[outcode] e.g. /schools/near/m20 + /schools/near-me geolocating hub + +**Thin-page threshold: generate a town or outcode page only where at least +five schools have current performance data.** Below that, 301 to the parent +LA page. This is the single most important constraint in the workstream — +it is what separates a location layer from index bloat. + +Each page must earn its place with content a parent would actually use, not +a template shell: + +- H1 matching the query intent ("Best primary schools in Manchester") +- Counts framed usefully: "137 primary schools, 9 rated Outstanding" +- A ranked table of the top 20 on the headline metric +- Local average against the England average +- Ofsted grade distribution +- Map +- Links to neighbouring towns and to the parent LA +- An FAQ block (feeds `FAQPage` in W4) +- Links to every school page in scope — this is what de-orphans W1's corpus + +Sizing estimate: ~150 usable LAs × 3 ≈ 450; towns clearing the threshold +≈ 1,200 × 2 ≈ 2,400; outcodes ≈ 2,300. Roughly **5,000 new pages**, +comfortably inside a sitemap index and well under the per-file 50,000 limit. + +### W3 — Make rankings indexable + +`/rankings` is driven entirely by query params, so Google indexes +approximately one page where there should be hundreds. + + /rankings/[phase]/[metric] + /rankings/[phase]/[metric]/[la] + +The interactive filter UI stays; its state moves into real paths. Param +forms canonicalise to the clean path. This is the direct play for C2. + +### W4 — Structured data and internal linking + +- Replace the bare `EducationalOrganization` on school pages with `School`, + and populate it properly. +- `BreadcrumbList` site-wide. +- `ItemList` on every rankings and location page. +- `FAQPage` on admissions and on location pages. +- New school-page modules: "Other schools in [town]", "Nearby schools", + "Compare with similar schools". Each links out to W2 and W3 pages, which + is what circulates authority instead of stranding it. + +Explicitly **not** doing `Dataset` or `AggregateRating` — no review corpus +exists, and fabricating one would be both useless and a policy violation. + +### W5 — Admissions expansion + +One static page currently carries an entire cluster. + + /admissions/[la] per-authority deadlines and offer day + /admissions/appeals + /admissions/national-offer-day + +`school admissions [LA]` is high-intent and highly seasonal; per-authority +pages are the natural unit. + +### W6 — Explainer content + + /guides/progress-8 + /guides/attainment-8 + /guides/sats-scaled-scores + /guides/ofsted-grades + /guides/expected-standard + +Each links into the corresponding W3 rankings page. Cheap to build, and it +is what gives the metric vocabulary enough topical weight to support C1–C4. + +### W7 — Head-to-head pages + + /compare/[school-a]-vs-[school-b] + +C7 is uncontested and native to the product. It is also the easiest way to +destroy everything W1 fixes: 27,229 schools generate 370 million pairs. +**Curated pairs only** — same town, both with current data, both with real +search demand — capped in the low thousands. Gated behind evidence that W2 +is indexing cleanly. + +### W8 — Metadata rewrite for C1 + +Current homepage title is `schoolcompare | Compare every school in England`, +which spends the most valuable position on the brand. Rewrite the homepage, +rankings and compare titles and descriptions around C1 phrasing. Small +change, and the cheapest item in the programme. + +## Sequencing + + W0 → W1 → W2 → W3 → W4 → W5 → W6 → (W7 if W2 indexes cleanly) + +W1 before W2 is not negotiable: adding pages to a corpus that is 22% thin +compounds the problem rather than diluting it. + +This spec is a programme, not a single implementation plan. Each workstream +gets its own plan and its own PR; W2 will likely need several. Only W0 and W1 +are ready to plan against today — the rest should be re-read after W0's +Search Console baseline lands, because that data may reorder them. + +## Testing + +Per CLAUDE.md, user-facing behaviour changes extend the `e2e/` journeys in +the same PR. Each workstream adds: + +- W1: canonical present and correct on every route; `/compare?urns=` carries + `noindex`; sitemap excludes a known dataless URN. +- W2: a known LA, town and outcode page renders with the expected school + count; a below-threshold town redirects to its LA. +- W3: a clean rankings path renders; the param form canonicalises to it. +- W4: JSON-LD parses and validates against the declared types. + +## Risks + +**Index bloat.** The failure mode of every programmatic SEO programme. The +five-school threshold, the W1 prune and the W7 gate are the three controls. + +**Helpful-content exposure.** Google's stance on templated location pages +has hardened. The mitigation is that each page carries genuinely local +computed data — real counts, real distributions, real local-vs-national +comparison — rather than a name substituted into boilerplate. + +**Build cost.** School pages already use ISR with a 7-day revalidate and +`PRERENDER_SCHOOLS` gating full prerender. 5,000 more routes need the same +treatment; full static generation of 32,000 pages is likely impractical in +CI. + +**Seasonality.** Results day and offer day dominate the traffic curve. +Judging the programme on a mid-summer window would misread it in either +direction. W0's baseline must be year-on-year, not month-on-month. + +## Open questions + +1. Catchment areas are Locrating's moat and a strong C3 driver + (`[postcode] school catchment`). `fact_admissions` carries admission + distances. Is deriving approximate catchment a later workstream, or out + of scope? diff --git a/e2e/tests/journeys.spec.ts b/e2e/tests/journeys.spec.ts index 9c18dc8..c1c8a8b 100644 --- a/e2e/tests/journeys.spec.ts +++ b/e2e/tests/journeys.spec.ts @@ -1449,3 +1449,104 @@ test('no single section dominates the height of a school page', async ({ page }) + sections.map((s) => `${s.id}=${s.h}`).join(', '), ).toBeLessThan(2.5); }); + +/* + * England-only corpus. + * + * GIAS ships the whole UK plus overseas and offshore establishments, none of + * which carry comparable DfE performance data. dim_school and dim_location + * exclude them (vars.non_england_school_type_codes in dbt_project.yml), which + * keeps them out of the site, the filter lists and the sitemap together. + * + * These assert the published surface, not the warehouse: the dbt test + * assert_england_only_schools guards the marts, and these guard what the + * environment actually serves once the marts have been rebuilt. + */ + +const WELSH_AUTHORITIES = [ + 'Blaenau Gwent', 'Bridgend', 'Caerphilly', 'Cardiff', 'Carmarthenshire', + 'Ceredigion', 'Conwy', 'Denbighshire', 'Flintshire', 'Gwynedd', + 'Isle of Anglesey', 'Merthyr Tydfil', 'Monmouthshire', 'Neath Port Talbot', + 'Newport', 'Pembrokeshire', 'Powys', 'Rhondda Cynon Taf', 'Swansea', + 'Torfaen', 'Vale of Glamorgan', 'Wrexham', +]; + +const NON_ENGLAND_AUTHORITIES = [ + 'BFPO Overseas Establishments', 'Fieldwork Overseas Establishments', + 'Gibraltar Overseas Establishments', 'Guernsey Offshore Establishments', + 'Isle of Man Offshore Establishments', 'Jersey Offshore Establishments', + 'Scotland Offshore Establishments', +]; + +const NON_ENGLAND_TYPES = [ + 'Welsh establishment', 'Offshore schools', + "Service children's education", 'British schools overseas', +]; + +test('the local authority filter offers no Welsh or overseas authority', async ({ page }) => { + const res = await page.request.get('/api/filters'); + expect(res.ok()).toBeTruthy(); + const { local_authorities: las } = await res.json(); + + expect(Array.isArray(las)).toBeTruthy(); + // Guards against the list being empty, which would pass the check below + // for the wrong reason. + expect(las.length).toBeGreaterThan(100); + + const leaked = [...WELSH_AUTHORITIES, ...NON_ENGLAND_AUTHORITIES] + .filter((la) => las.includes(la)); + expect(leaked, `non-England authorities still offered: ${leaked.join(', ')}`) + .toEqual([]); +}); + +test('the school type filter offers no non-England establishment type', async ({ page }) => { + const res = await page.request.get('/api/filters'); + expect(res.ok()).toBeTruthy(); + const { school_types: types } = await res.json(); + + expect(Array.isArray(types)).toBeTruthy(); + expect(types.length).toBeGreaterThan(10); + + const leaked = NON_ENGLAND_TYPES.filter((t) => types.includes(t)); + expect(leaked, `non-England types still offered: ${leaked.join(', ')}`) + .toEqual([]); +}); + +test('searching a Welsh authority by name returns no schools', async ({ page }) => { + // Cardiff held 144 Welsh establishments and nothing else, so the authority + // should now be absent from the corpus entirely rather than merely thinned. + const res = await page.request.get('/api/schools?local_authority=Cardiff&page_size=1'); + expect(res.ok()).toBeTruthy(); + const body = await res.json(); + expect(body.total ?? (body.schools ?? []).length).toBe(0); +}); + +test('a Welsh school URL 404s while an English one still resolves', async ({ page }) => { + // Paired on purpose: the Welsh assertion alone would also pass if the whole + // site were down, which is the failure this test most needs to distinguish. + const english = await page.request.get('/api/schools?search=primary&per_page=1'); + expect(english.ok()).toBeTruthy(); + const [first] = (await english.json()).schools ?? []; + expect(first, 'no English school available to compare against').toBeTruthy(); + + const good = await page.goto(`/school/${first.urn}-x`); + expect(good?.status(), 'an English school should still resolve').toBeLessThan(400); + + // Adamsdown Primary School, Cardiff — a Welsh establishment (URN 401559). + const welsh = await page.goto('/school/401559-adamsdown-primary-school'); + expect(welsh?.status(), 'a Welsh school should no longer resolve').toBe(404); +}); + +test('the sitemap submits no Welsh or overseas school', async ({ page }) => { + const res = await page.request.get('/sitemap.xml'); + expect(res.ok()).toBeTruthy(); + const xml = await res.text(); + + const urlCount = (xml.match(//g) ?? []).length; + expect(urlCount, 'sitemap looks empty or truncated').toBeGreaterThan(1000); + + // 401559 (Cardiff) and 402426 (ACT Schools, Cardiff) were both submitted + // before the England-only filter landed. + expect(xml).not.toContain('/school/401559'); + expect(xml).not.toContain('/school/402426'); +}); diff --git a/pipeline/transform/dbt_project.yml b/pipeline/transform/dbt_project.yml index ba23641..ff62b6f 100644 --- a/pipeline/transform/dbt_project.yml +++ b/pipeline/transform/dbt_project.yml @@ -11,6 +11,19 @@ seed-paths: ["seeds"] target-path: "target" clean-targets: ["target", "dbt_packages"] +# schoolcompare publishes England only. GIAS ships the whole UK plus overseas +# and offshore establishments, none of which have comparable DfE performance +# data — Wales does not publish on the English measures at all. Excluding them +# at the mart boundary keeps them out of the site, the filter lists and the +# sitemap together. Codes are TypeOfEstablishment, see seeds/gias_code_names.csv. +# 25 = Offshore schools (Jersey, Guernsey, Isle of Man, Gibraltar) +# 26 = Service children's education (BFPO, Fieldwork Overseas) +# 30 = Welsh establishment +# 37 = British schools overseas +# dim_school, dim_location and assert_england_only_schools all read this list. +vars: + non_england_school_type_codes: [25, 26, 30, 37] + models: school_compare: staging: diff --git a/pipeline/transform/models/marts/dim_location.sql b/pipeline/transform/models/marts/dim_location.sql index b05e8a7..3495b14 100644 --- a/pipeline/transform/models/marts/dim_location.sql +++ b/pipeline/transform/models/marts/dim_location.sql @@ -31,5 +31,9 @@ select else null end as longitude from {{ ref('stg_gias_establishments') }} s --- Must match dim_school's status filter exactly (the API inner-joins the two). +-- Must match dim_school's status and England filters exactly (the API +-- inner-joins the two). where s.status_code in (1, 3) +-- coalesce, not a bare NOT IN: a null type code would make the predicate +-- null and drop the row silently. Unknown type is not grounds for exclusion. +and coalesce(s.school_type_code, -1) not in ({{ var('non_england_school_type_codes') | join(', ') }}) diff --git a/pipeline/transform/models/marts/dim_school.sql b/pipeline/transform/models/marts/dim_school.sql index 86dfa2e..bcb246c 100644 --- a/pipeline/transform/models/marts/dim_school.sql +++ b/pipeline/transform/models/marts/dim_school.sql @@ -91,3 +91,8 @@ left join {{ ref('int_ofsted_latest') }} o on s.urn = o.urn -- 1 = Open; 3 = Open, but proposed to close (still operating; drops out when -- GIAS flips to Closed — marts fully rebuild each run). where s.status_code in (1, 3) +-- England only. dim_location must apply this filter identically (the API +-- inner-joins the two). See vars in dbt_project.yml for what the codes are. +-- coalesce, not a bare NOT IN: a null type code would make the predicate +-- null and drop the row silently. Unknown type is not grounds for exclusion. +and coalesce(s.school_type_code, -1) not in ({{ var('non_england_school_type_codes') | join(', ') }}) diff --git a/pipeline/transform/tests/assert_england_only_schools.sql b/pipeline/transform/tests/assert_england_only_schools.sql new file mode 100644 index 0000000..0356e29 --- /dev/null +++ b/pipeline/transform/tests/assert_england_only_schools.sql @@ -0,0 +1,28 @@ +-- Custom test: the published corpus is England only. +-- +-- GIAS ships the whole UK plus overseas and offshore establishments. None of +-- them carry comparable DfE performance data, so dim_school and dim_location +-- filter them out (see vars.non_england_school_type_codes in dbt_project.yml). +-- This test is the guard: a GIAS refresh that reintroduces them, or an edit +-- that drops the filter from one of the two models, fails the pipeline rather +-- than quietly republishing ~2,000 dataless pages to the site and the sitemap. + +select + urn, + school_name, + school_type_code +from {{ ref('dim_school') }} +where school_type_code in ({{ var('non_england_school_type_codes') | join(', ') }}) + +union all + +-- dim_location is inner-joined to dim_school by the API, so it must filter +-- identically. Anything here that dim_school does not have means the two +-- models have drifted apart. +select + l.urn, + null::text as school_name, + null::integer as school_type_code +from {{ ref('dim_location') }} l +left join {{ ref('dim_school') }} s on l.urn = s.urn +where s.urn is null