Files
school_compare/docs/superpowers/specs/2026-08-20-seo-programme-design.md
T
TudorandClaude Opus 5 7650b16f62 feat(data): publish England only, dropping Welsh and overseas establishments
GIAS ships the whole UK plus overseas and offshore establishments. None of
them carry comparable DfE performance data — Wales does not publish on the
English measures at all — so every one of these pages rendered with null
results, null Ofsted and null phase. There were 2,036 of them: 1,569 Welsh,
123 offshore (Jersey, Guernsey, Isle of Man, Gibraltar), 316 British schools
overseas and 28 service children's schools. All 2,036 were being submitted to
search engines, alongside 29 local authorities that existed in the filters
purely to list them.

Filter at the mart boundary rather than the view layer. dim_school and
dim_location both exclude TypeOfEstablishment in {25, 26, 30, 37}, listed once
as vars.non_england_school_type_codes. Everything downstream reads those two
marts — search, the school page, /api/filters, rankings, Typesense and
build_sitemap() — so one filter removes them from the site and the sitemap
together, and Typesense drops them on its next rebuild since it recreates the
collection and swaps the alias rather than upserting in place.

coalesce rather than a bare NOT IN: a null type code would make the predicate
null and drop the row silently, and an unknown type is not grounds for
exclusion. No establishment has a null type today, but a future GIAS refresh
could ship one and the loss would be invisible.

assert_england_only_schools guards both directions: no excluded type survives
in dim_school, and dim_location holds no URN dim_school lacks — the API
inner-joins them, so the two filters drifting apart would silently shrink the
corpus.

Corpus goes from 27,229 schools to 25,193, and the authority list from 182 to
153. The 1,569 Welsh URLs now 404.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015mWQnpye9F299NVRCCSRvj
2026-08-20 21:46:28 +01:00

310 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SEO Programme — Design
Date: 2026-08-20
Status: awaiting review
## Problem
schoolcompare ranks second for "school compare" — an exact match for the
brand and the domain. It ranks poorly for "compare schools", "school
comparison" and "schools near me". The first is a naming artefact and
transfers to nothing; the rest are the queries that actually carry parent
demand.
The cause is structural, not editorial. The site publishes five route
families:
/ /compare /rankings /admissions /school/[slug]
Location intent has no landing page at all. Every competitor outranking us
on those queries wins with programmatic location pages:
| Competitor | URL pattern |
|---------------|--------------------------------------------|
| School Guide | `/best-schools-in/manchester` |
| Locrating | `/the-best-primary-schools-in-Manchester_…`|
| FindMySchool | `/best-primary-schools/manchester` |
| Snobe | `/best-primary-schools/manchester` |
| School Atlas | `/guides/best-primary-schools-manchester` |
"Schools near me" is a local-intent query. Google resolves it against the
user's coordinates and serves pages that are *about a place*. A national
homepage cannot win it. No title or description change fixes this; only
pages Google can localise will.
## Baseline (measured 2026-08-20, production API)
| Measure | Value |
|--------------------------------------------|---------|
| Unique schools | 27,229 |
| URLs in sitemap.xml | 27,232 |
| Schools with 2024/25 performance data | 21,266 |
| Schools with **no** current performance data| ~5,963 (22%) |
| Welsh establishments (all metrics null) | 1,569 |
| Overseas / offshore establishments | 467 |
| Static URLs in sitemap | 3 |
| Routes setting a canonical | 1 of 5 |
Three findings from that table drive the plan.
**We submit ~6,000 thin pages to Google.** `build_sitemap()`
(`backend/app.py:73`) enumerates every URN regardless of whether the school
has any data. Welsh establishments return `school_type: "Welsh
establishment"` with every performance metric, Ofsted grade and phase field
null. Overseas and offshore establishments ("BFPO Overseas Establishments",
"Gibraltar Overseas Establishments", "Jersey Offshore Establishments") are
in the local-authority list too. At 22% of the submitted corpus this is a
site-wide quality signal problem and a crawl-budget waste, not a rounding
error.
W1 item 4 removes 2,036 of those — every non-England establishment — taking
the corpus to 25,193. The 3,927 that remain are English schools with no
current data: mostly newly opened, special, nursery or alternative provision.
Those are a template problem, not a corpus problem, and item 5 handles them
separately.
**The homepage is its own competitor.** `app/page.tsx` accepts eleven search
params (`search`, `local_authority`, `school_type`, `phase`, `page`,
`postcode`, `radius`, `sort`, `gender`, `admissions_policy`,
`has_sixth_form`) and sets no canonical. Every filter combination is a
crawlable near-duplicate of the single page we are asking to rank for
"compare schools".
**School pages are near-orphans.** Reachable from the sitemap and from site
search, but almost nothing links to them contextually, so they accrue no
internal authority.
Also noted: the sitemap emits invented `priority` values and no `lastmod`.
Google ignores `priority` and `changefreq` entirely; `lastmod` is the field
it does read, and we omit it.
## Keyword clusters
Ranked by judgement of UK parent search behaviour and by the competitive
SERP evidence above. Google Search Console is connected, so cluster
priorities are to be re-derived from measured impressions before build
starts (see Workstream 0).
**C1 — Head "compare" terms.** compare schools · school comparison · school
comparison tool · compare school performance · compare primary schools ·
compare secondary schools · compare two schools
**C2 — League tables and rankings.** primary school league tables ·
secondary school league tables · school league tables 2026 · SATs results by
school · GCSE results by school · KS2 league tables · Progress 8 rankings ·
best primary schools in [town] · top 10 primary schools in [LA]
**C3 — Local / near me.** schools near me · primary schools near me ·
secondary schools near me · best schools near me · good schools near me ·
schools in [town] · primary schools in [LA] · schools near [postcode] ·
[postcode] school catchment
**C4 — Individual school long tail.** [school] ofsted · [school] SATs
results · [school] catchment area · [school] reviews · [school] URN
**C5 — Admissions.** primary school admissions 2027 · national offer day
2027 · school application deadline · school admissions appeal ·
oversubscription criteria · distance criteria school admissions · didn't get
first choice school · school admissions [LA]
**C6 — Metric explainers.** what is a good SATs score · what is Progress 8 ·
what is Attainment 8 · expected standard KS2 meaning · scaled score
explained · Ofsted grades explained · Ofsted report cards · pupil premium
explained
**C7 — Head to head.** [school A] vs [school B] · academy vs community
school · grammar school vs comprehensive · faith school vs community school
## Workstreams
### W0 — Measure before touching anything
Export a Google Search Console baseline: impressions, clicks, average
position and CTR by query and by page, for the trailing 16 months. Bucket
queries into C1–C7. This sets the counterfactual — without it, no later
claim about lift is defensible, because school-search traffic is strongly
seasonal (results day in December, offer day in March/April).
Re-rank C1–C7 against measured impressions and adjust the sequence below if
the data disagrees with the judgement calls.
### W1 — Crawl hygiene and index sanity
Cheap, and it unblocks everything after it. Adding 5,000 pages on top of a
corpus that is 22% thin would compound the existing problem.
1. Canonical on every route. `/`, `/rankings`, `/compare` and `/admissions`
currently set none.
2. The homepage canonicalises to `/` regardless of search params.
3. `/compare?urns=…` gets `noindex, follow` — it is an unbounded parameter
space with no standalone value.
4. **England only — DONE.** Wales, the Crown Dependencies, Gibraltar and the
service/overseas schools are removed from the corpus at the mart boundary,
not hidden at the view layer. `dim_school` and `dim_location` both exclude
`TypeOfEstablishment` in {25, 26, 30, 37} — Offshore schools, Service
children's education, Welsh establishment, British schools overseas —
listed once as `vars.non_england_school_type_codes` in `dbt_project.yml`.
That removes 2,036 establishments and 29 local authorities, and because
`build_sitemap()` reads the same marts, it drops them from the sitemap in
the same stroke. `assert_england_only_schools` fails the pipeline if a GIAS
refresh reintroduces them or if the two models drift apart.
5. Prune the remaining thin pages: exclude any school with no performance data
**and** no Ofsted record. Distinct from item 4 — these are English schools
with nothing yet to show, so the fix may be a better template rather than
removal.
6. Rebuild the sitemap as a sitemap **index**: one child per page family,
real `lastmod` from the data-load timestamp, `priority` and `changefreq`
dropped.
### W2 — The location layer
The dominant lever. `dim_location` already carries `town`, `county`,
`local_authority_name`, `parliamentary_constituency`, `latitude`,
`longitude` and `postcode`, so no new ingestion is required.
Routes:
/schools/[la] e.g. /schools/manchester
/schools/[la]/primary
/schools/[la]/secondary
/best-primary-schools/[town]
/best-secondary-schools/[town]
/schools/near/[outcode] e.g. /schools/near/m20
/schools/near-me geolocating hub
**Thin-page threshold: generate a town or outcode page only where at least
five schools have current performance data.** Below that, 301 to the parent
LA page. This is the single most important constraint in the workstream —
it is what separates a location layer from index bloat.
Each page must earn its place with content a parent would actually use, not
a template shell:
- H1 matching the query intent ("Best primary schools in Manchester")
- Counts framed usefully: "137 primary schools, 9 rated Outstanding"
- A ranked table of the top 20 on the headline metric
- Local average against the England average
- Ofsted grade distribution
- Map
- Links to neighbouring towns and to the parent LA
- An FAQ block (feeds `FAQPage` in W4)
- Links to every school page in scope — this is what de-orphans W1's corpus
Sizing estimate: ~150 usable LAs × 3 ≈ 450; towns clearing the threshold
≈ 1,200 × 2 ≈ 2,400; outcodes ≈ 2,300. Roughly **5,000 new pages**,
comfortably inside a sitemap index and well under the per-file 50,000 limit.
### W3 — Make rankings indexable
`/rankings` is driven entirely by query params, so Google indexes
approximately one page where there should be hundreds.
/rankings/[phase]/[metric]
/rankings/[phase]/[metric]/[la]
The interactive filter UI stays; its state moves into real paths. Param
forms canonicalise to the clean path. This is the direct play for C2.
### W4 — Structured data and internal linking
- Replace the bare `EducationalOrganization` on school pages with `School`,
and populate it properly.
- `BreadcrumbList` site-wide.
- `ItemList` on every rankings and location page.
- `FAQPage` on admissions and on location pages.
- New school-page modules: "Other schools in [town]", "Nearby schools",
"Compare with similar schools". Each links out to W2 and W3 pages, which
is what circulates authority instead of stranding it.
Explicitly **not** doing `Dataset` or `AggregateRating` — no review corpus
exists, and fabricating one would be both useless and a policy violation.
### W5 — Admissions expansion
One static page currently carries an entire cluster.
/admissions/[la] per-authority deadlines and offer day
/admissions/appeals
/admissions/national-offer-day
`school admissions [LA]` is high-intent and highly seasonal; per-authority
pages are the natural unit.
### W6 — Explainer content
/guides/progress-8
/guides/attainment-8
/guides/sats-scaled-scores
/guides/ofsted-grades
/guides/expected-standard
Each links into the corresponding W3 rankings page. Cheap to build, and it
is what gives the metric vocabulary enough topical weight to support C1–C4.
### W7 — Head-to-head pages
/compare/[school-a]-vs-[school-b]
C7 is uncontested and native to the product. It is also the easiest way to
destroy everything W1 fixes: 27,229 schools generate 370 million pairs.
**Curated pairs only** — same town, both with current data, both with real
search demand — capped in the low thousands. Gated behind evidence that W2
is indexing cleanly.
### W8 — Metadata rewrite for C1
Current homepage title is `schoolcompare | Compare every school in England`,
which spends the most valuable position on the brand. Rewrite the homepage,
rankings and compare titles and descriptions around C1 phrasing. Small
change, and the cheapest item in the programme.
## Sequencing
W0 → W1 → W2 → W3 → W4 → W5 → W6 → (W7 if W2 indexes cleanly)
W1 before W2 is not negotiable: adding pages to a corpus that is 22% thin
compounds the problem rather than diluting it.
This spec is a programme, not a single implementation plan. Each workstream
gets its own plan and its own PR; W2 will likely need several. Only W0 and W1
are ready to plan against today — the rest should be re-read after W0's
Search Console baseline lands, because that data may reorder them.
## Testing
Per CLAUDE.md, user-facing behaviour changes extend the `e2e/` journeys in
the same PR. Each workstream adds:
- W1: canonical present and correct on every route; `/compare?urns=` carries
`noindex`; sitemap excludes a known dataless URN.
- W2: a known LA, town and outcode page renders with the expected school
count; a below-threshold town redirects to its LA.
- W3: a clean rankings path renders; the param form canonicalises to it.
- W4: JSON-LD parses and validates against the declared types.
## Risks
**Index bloat.** The failure mode of every programmatic SEO programme. The
five-school threshold, the W1 prune and the W7 gate are the three controls.
**Helpful-content exposure.** Google's stance on templated location pages
has hardened. The mitigation is that each page carries genuinely local
computed data — real counts, real distributions, real local-vs-national
comparison — rather than a name substituted into boilerplate.
**Build cost.** School pages already use ISR with a 7-day revalidate and
`PRERENDER_SCHOOLS` gating full prerender. 5,000 more routes need the same
treatment; full static generation of 32,000 pages is likely impractical in
CI.
**Seasonality.** Results day and offer day dominate the traffic curve.
Judging the programme on a mid-summer window would misread it in either
direction. W0's baseline must be year-on-year, not month-on-month.
## Open questions
1. Catchment areas are Locrating's moat and a strong C3 driver
(`[postcode] school catchment`). `fact_admissions` carries admission
distances. Is deriving approximate catchment a later workstream, or out
of scope?