Files
school_compare/docs/superpowers/specs/2026-08-20-seo-programme-design.md
T

309 lines
13 KiB
Markdown
Raw Normal View History

# SEO Programme — Design
Date: 2026-08-20
Status: awaiting review
## Problem
schoolcompare ranks second for "school compare" — an exact match for the
brand and the domain. It ranks poorly for "compare schools", "school
comparison" and "schools near me". The first is a naming artefact and
transfers to nothing; the rest are the queries that actually carry parent
demand.
The cause is structural, not editorial. The site publishes five route
families:
/ /compare /rankings /admissions /school/[slug]
Location intent has no landing page at all. Every competitor outranking us
on those queries wins with programmatic location pages:
| Competitor | URL pattern |
|---------------|--------------------------------------------|
| School Guide | `/best-schools-in/manchester` |
| Locrating | `/the-best-primary-schools-in-Manchester_…`|
| FindMySchool | `/best-primary-schools/manchester` |
| Snobe | `/best-primary-schools/manchester` |
| School Atlas | `/guides/best-primary-schools-manchester` |
"Schools near me" is a local-intent query. Google resolves it against the
user's coordinates and serves pages that are *about a place*. A national
homepage cannot win it. No title or description change fixes this; only
pages Google can localise will.
## Baseline (measured 2026-08-20, production API)
| Measure | Value |
|--------------------------------------------|---------|
| Unique schools | 27,229 |
| URLs in sitemap.xml | 27,232 |
| Schools with 2024/25 performance data | 21,266 |
| Schools with **no** current performance data| ~5,963 (22%) |
| Welsh establishments (all metrics null) | 1,569 |
| Overseas / offshore establishments | 467 |
| Static URLs in sitemap | 3 |
| Routes setting a canonical | 1 of 5 |
Three findings from that table drive the plan.
**We submit ~6,000 thin pages to Google.** `build_sitemap()`
(`backend/app.py:73`) enumerates every URN regardless of whether the school
has any data. Welsh establishments return `school_type: "Welsh
establishment"` with every performance metric, Ofsted grade and phase field
null. Overseas and offshore establishments ("BFPO Overseas Establishments",
"Gibraltar Overseas Establishments", "Jersey Offshore Establishments") are
in the local-authority list too. At 22% of the submitted corpus this is a
site-wide quality signal problem and a crawl-budget waste, not a rounding
error.
W1 item 4 removes 2,036 of those — every non-England establishment — taking
the corpus to 25,193. The 3,927 that remain are English schools with no
current data: mostly newly opened, special, nursery or alternative provision.
Those are a template problem, not a corpus problem, and item 5 handles them
separately.
**The homepage is its own competitor.** `app/page.tsx` accepts eleven search
params (`search`, `local_authority`, `school_type`, `phase`, `page`,
`postcode`, `radius`, `sort`, `gender`, `admissions_policy`,
`has_sixth_form`) and sets no canonical. Every filter combination is a
crawlable near-duplicate of the single page we are asking to rank for
"compare schools".
**School pages are near-orphans.** Reachable from the sitemap and from site
search, but almost nothing links to them contextually, so they accrue no
internal authority.
Also noted: the sitemap emits invented `priority` values and no `lastmod`.
Google ignores `priority` and `changefreq` entirely; `lastmod` is the field
it does read, and we omit it.
## Keyword clusters
Ranked by judgement of UK parent search behaviour and by the competitive
SERP evidence above. Google Search Console is connected, so cluster
priorities are to be re-derived from measured impressions before build
starts (see Workstream 0).
**C1 — Head "compare" terms.** compare schools · school comparison · school
comparison tool · compare school performance · compare primary schools ·
compare secondary schools · compare two schools
**C2 — League tables and rankings.** primary school league tables ·
secondary school league tables · school league tables 2026 · SATs results by
school · GCSE results by school · KS2 league tables · Progress 8 rankings ·
best primary schools in [town] · top 10 primary schools in [LA]
**C3 — Local / near me.** schools near me · primary schools near me ·
secondary schools near me · best schools near me · good schools near me ·
schools in [town] · primary schools in [LA] · schools near [postcode] ·
[postcode] school catchment
**C4 — Individual school long tail.** [school] ofsted · [school] SATs
results · [school] catchment area · [school] reviews · [school] URN
**C5 — Admissions.** primary school admissions 2027 · national offer day
2027 · school application deadline · school admissions appeal ·
oversubscription criteria · distance criteria school admissions · didn't get
first choice school · school admissions [LA]
**C6 — Metric explainers.** what is a good SATs score · what is Progress 8 ·
what is Attainment 8 · expected standard KS2 meaning · scaled score
explained · Ofsted grades explained · Ofsted report cards · pupil premium
explained
**C7 — Head to head.** [school A] vs [school B] · academy vs community
school · grammar school vs comprehensive · faith school vs community school
## Workstreams
### W0 — Measure before touching anything
Export a Google Search Console baseline: impressions, clicks, average
position and CTR by query and by page, for the trailing 16 months. Bucket
queries into C1–C7. This sets the counterfactual — without it, no later
claim about lift is defensible, because school-search traffic is strongly
seasonal (results day in December, offer day in March/April).
Re-rank C1–C7 against measured impressions and adjust the sequence below if
the data disagrees with the judgement calls.
### W1 — Crawl hygiene and index sanity
Cheap, and it unblocks everything after it. Adding 5,000 pages on top of a
corpus that is 22% thin would compound the existing problem.
1. Canonical on every route. `/`, `/rankings`, `/compare` and `/admissions`
currently set none.
2. The homepage canonicalises to `/` regardless of search params.
3. `/compare?urns=…` gets `noindex, follow` — it is an unbounded parameter
space with no standalone value.
4. **England only — DONE.** Wales, the Crown Dependencies, Gibraltar and the
service/overseas schools are removed from the corpus at the mart boundary,
not hidden at the view layer. `dim_school` and `dim_location` both exclude
`TypeOfEstablishment` in {25, 26, 30, 37} — Offshore schools, Service
children's education, Welsh establishment, British schools overseas —
listed once as `vars.non_england_school_type_codes` in `dbt_project.yml`.
That removes 2,036 establishments and 29 local authorities, and because
`build_sitemap()` reads the same marts, it drops them from the sitemap in
the same stroke. `assert_england_only_schools` fails the pipeline if a GIAS
refresh reintroduces them or if the two models drift apart.
5. Prune the remaining thin pages: exclude any school with no performance data
**and** no Ofsted record. Distinct from item 4 — these are English schools
with nothing yet to show, so the fix may be a better template rather than
removal.
6. Rebuild the sitemap as a sitemap **index**: one child per page family,
real `lastmod` from the data-load timestamp, `priority` and `changefreq`
dropped.
### W2 — The location layer
The dominant lever. `dim_location` already carries `town`, `county`,
`local_authority_name`, `parliamentary_constituency`, `latitude`,
`longitude` and `postcode`, so no new ingestion is required.
Routes:
/schools/[la] e.g. /schools/manchester
/schools/[la]/primary
/schools/[la]/secondary
/best-primary-schools/[town]
/best-secondary-schools/[town]
/schools/near/[outcode] e.g. /schools/near/m20
/schools/near-me geolocating hub
**Thin-page threshold: generate a town or outcode page only where at least
five schools have current performance data.** Below that, 301 to the parent
LA page. This is the single most important constraint in the workstream —
it is what separates a location layer from index bloat.
Each page must earn its place with content a parent would actually use, not
a template shell:
- H1 matching the query intent ("Best primary schools in Manchester")
- Counts framed usefully: "137 primary schools, 9 rated Outstanding"
- A ranked table of the top 20 on the headline metric
- Local average against the England average
- Ofsted grade distribution
- Map
- Links to neighbouring towns and to the parent LA
- An FAQ block (feeds `FAQPage` in W4)
- Links to every school page in scope — this is what de-orphans W1's corpus
Sizing estimate: ~150 usable LAs × 3 ≈ 450; towns clearing the threshold
≈ 1,200 × 2 ≈ 2,400; outcodes ≈ 2,300. Roughly **5,000 new pages**,
comfortably inside a sitemap index and well under the per-file 50,000 limit.
### W3 — Make rankings indexable
`/rankings` is driven entirely by query params, so Google indexes
approximately one page where there should be hundreds.
/rankings/[phase]/[metric]
/rankings/[phase]/[metric]/[la]
The interactive filter UI stays; its state moves into real paths. Param
forms canonicalise to the clean path. This is the direct play for C2.
### W4 — Structured data and internal linking
- Replace the bare `EducationalOrganization` on school pages with `School`,
and populate it properly.
- `BreadcrumbList` site-wide.
- `ItemList` on every rankings and location page.
- `FAQPage` on admissions and on location pages.
- New school-page modules: "Other schools in [town]", "Nearby schools",
"Compare with similar schools". Each links out to W2 and W3 pages, which
is what circulates authority instead of stranding it.
Explicitly **not** doing `Dataset` or `AggregateRating` — no review corpus
exists, and fabricating one would be both useless and a policy violation.
### W5 — Admissions expansion
One static page currently carries an entire cluster.
/admissions/[la] per-authority deadlines and offer day
/admissions/appeals
/admissions/national-offer-day
`school admissions [LA]` is high-intent and highly seasonal; per-authority
pages are the natural unit.
### W6 — Explainer content
/guides/progress-8
/guides/attainment-8
/guides/sats-scaled-scores
/guides/ofsted-grades
/guides/expected-standard
Each links into the corresponding W3 rankings page. Cheap to build, and it
is what gives the metric vocabulary enough topical weight to support C1–C4.
### W7 — Head-to-head pages
/compare/[school-a]-vs-[school-b]
C7 is uncontested and native to the product. It is also the easiest way to
destroy everything W1 fixes: 27,229 schools generate 370 million pairs.
**Curated pairs only** — same town, both with current data, both with real
search demand — capped in the low thousands. Gated behind evidence that W2
is indexing cleanly.
### W8 — Metadata rewrite for C1
Current homepage title is `schoolcompare | Compare every school in England`,
which spends the most valuable position on the brand. Rewrite the homepage,
rankings and compare titles and descriptions around C1 phrasing. Small
change, and the cheapest item in the programme.
## Sequencing
W0 → W1 → W2 → W3 → W4 → W5 → W6 → (W7 if W2 indexes cleanly)
W1 before W2 is not negotiable: adding pages to a corpus that is 22% thin
compounds the problem rather than diluting it.
This spec is a programme, not a single implementation plan. Each workstream
gets its own plan and its own PR; W2 will likely need several. Only W0 and W1
are ready to plan against today — the rest should be re-read after W0's
Search Console baseline lands, because that data may reorder them.
## Testing
Per CLAUDE.md, user-facing behaviour changes extend the `e2e/` journeys in
the same PR. Each workstream adds:
- W1: canonical present and correct on every route; `/compare?urns=` carries
`noindex`; sitemap excludes a known dataless URN.
- W2: a known LA, town and outcode page renders with the expected school
count; a below-threshold town redirects to its LA.
- W3: a clean rankings path renders; the param form canonicalises to it.
- W4: JSON-LD parses and validates against the declared types.
## Risks
**Index bloat.** The failure mode of every programmatic SEO programme. The
five-school threshold, the W1 prune and the W7 gate are the three controls.
**Helpful-content exposure.** Google's stance on templated location pages
has hardened. The mitigation is that each page carries genuinely local
computed data — real counts, real distributions, real local-vs-national
comparison — rather than a name substituted into boilerplate.
**Build cost.** School pages already use ISR with a 7-day revalidate and
`PRERENDER_SCHOOLS` gating full prerender. 5,000 more routes need the same
treatment; full static generation of 32,000 pages is likely impractical in
CI.
**Seasonality.** Results day and offer day dominate the traffic curve.
Judging the programme on a mid-summer window would misread it in either
direction. W0's baseline must be year-on-year, not month-on-month.
## Open questions
1. Catchment areas are Locrating's moat and a strong C3 driver
(`[postcode] school catchment`). `fact_admissions` carries admission
distances. Is deriving approximate catchment a later workstream, or out
of scope?