Merge pull request 'feat(data): publish England only, dropping Welsh and overseas establishments' (#107) from feat/england-only-corpus into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m15s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m27s

Reviewed-on: #107
This commit was merged in pull request #107.
This commit is contained in:
tudor committed 2026-08-20 21:20:13 +00:00
commit c47fe38971
7 files changed
+1716 -1

No files matched your search

File diff suppressed because it is too large. Load diff
@@ -0,0 +1,309 @@
# SEO Programme — Design
Date: 2026-08-20
Status: awaiting review
## Problem
schoolcompare ranks second for "school compare" — an exact match for the
brand and the domain. It ranks poorly for "compare schools", "school
comparison" and "schools near me". The first is a naming artefact and
transfers to nothing; the rest are the queries that actually carry parent
demand.
The cause is structural, not editorial. The site publishes five route
families:
/ /compare /rankings /admissions /school/[slug]
Location intent has no landing page at all. Every competitor outranking us
on those queries wins with programmatic location pages:
| Competitor | URL pattern |
|---------------|--------------------------------------------|
| School Guide | `/best-schools-in/manchester` |
| Locrating | `/the-best-primary-schools-in-Manchester_…`|
| FindMySchool | `/best-primary-schools/manchester` |
| Snobe | `/best-primary-schools/manchester` |
| School Atlas | `/guides/best-primary-schools-manchester` |
"Schools near me" is a local-intent query. Google resolves it against the
user's coordinates and serves pages that are *about a place*. A national
homepage cannot win it. No title or description change fixes this; only
pages Google can localise will.
## Baseline (measured 2026-08-20, production API)
| Measure | Value |
|--------------------------------------------|---------|
| Unique schools | 27,229 |
| URLs in sitemap.xml | 27,232 |
| Schools with 2024/25 performance data | 21,266 |
| Schools with **no** current performance data| ~5,963 (22%) |
| Welsh establishments (all metrics null) | 1,569 |
| Overseas / offshore establishments | 467 |
| Static URLs in sitemap | 3 |
| Routes setting a canonical | 1 of 5 |
Three findings from that table drive the plan.
**We submit ~6,000 thin pages to Google.** `build_sitemap()`
(`backend/app.py:73`) enumerates every URN regardless of whether the school
has any data. Welsh establishments return `school_type: "Welsh
establishment"` with every performance metric, Ofsted grade and phase field
null. Overseas and offshore establishments ("BFPO Overseas Establishments",
"Gibraltar Overseas Establishments", "Jersey Offshore Establishments") are
in the local-authority list too. At 22% of the submitted corpus this is a
site-wide quality signal problem and a crawl-budget waste, not a rounding
error.
W1 item 4 removes 2,036 of those — every non-England establishment — taking
the corpus to 25,193. The 3,927 that remain are English schools with no
current data: mostly newly opened, special, nursery or alternative provision.
Those are a template problem, not a corpus problem, and item 5 handles them
separately.
**The homepage is its own competitor.** `app/page.tsx` accepts eleven search
params (`search`, `local_authority`, `school_type`, `phase`, `page`,
`postcode`, `radius`, `sort`, `gender`, `admissions_policy`,
`has_sixth_form`) and sets no canonical. Every filter combination is a
crawlable near-duplicate of the single page we are asking to rank for
"compare schools".
**School pages are near-orphans.** Reachable from the sitemap and from site
search, but almost nothing links to them contextually, so they accrue no
internal authority.
Also noted: the sitemap emits invented `priority` values and no `lastmod`.
Google ignores `priority` and `changefreq` entirely; `lastmod` is the field
it does read, and we omit it.
## Keyword clusters
Ranked by judgement of UK parent search behaviour and by the competitive
SERP evidence above. Google Search Console is connected, so cluster
priorities are to be re-derived from measured impressions before build
starts (see Workstream 0).
**C1 — Head "compare" terms.** compare schools · school comparison · school
comparison tool · compare school performance · compare primary schools ·
compare secondary schools · compare two schools
**C2 — League tables and rankings.** primary school league tables ·
secondary school league tables · school league tables 2026 · SATs results by
school · GCSE results by school · KS2 league tables · Progress 8 rankings ·
best primary schools in [town] · top 10 primary schools in [LA]
**C3 — Local / near me.** schools near me · primary schools near me ·
secondary schools near me · best schools near me · good schools near me ·
schools in [town] · primary schools in [LA] · schools near [postcode] ·
[postcode] school catchment
**C4 — Individual school long tail.** [school] ofsted · [school] SATs
results · [school] catchment area · [school] reviews · [school] URN
**C5 — Admissions.** primary school admissions 2027 · national offer day
2027 · school application deadline · school admissions appeal ·
oversubscription criteria · distance criteria school admissions · didn't get
first choice school · school admissions [LA]
**C6 — Metric explainers.** what is a good SATs score · what is Progress 8 ·
what is Attainment 8 · expected standard KS2 meaning · scaled score
explained · Ofsted grades explained · Ofsted report cards · pupil premium
explained
**C7 — Head to head.** [school A] vs [school B] · academy vs community
school · grammar school vs comprehensive · faith school vs community school
## Workstreams
### W0 — Measure before touching anything
Export a Google Search Console baseline: impressions, clicks, average
position and CTR by query and by page, for the trailing 16 months. Bucket
queries into C1–C7. This sets the counterfactual — without it, no later
claim about lift is defensible, because school-search traffic is strongly
seasonal (results day in December, offer day in March/April).
Re-rank C1–C7 against measured impressions and adjust the sequence below if
the data disagrees with the judgement calls.
### W1 — Crawl hygiene and index sanity
Cheap, and it unblocks everything after it. Adding 5,000 pages on top of a
corpus that is 22% thin would compound the existing problem.
1. Canonical on every route. `/`, `/rankings`, `/compare` and `/admissions`
currently set none.
2. The homepage canonicalises to `/` regardless of search params.
3. `/compare?urns=…` gets `noindex, follow` — it is an unbounded parameter
space with no standalone value.
4. **England only — DONE.** Wales, the Crown Dependencies, Gibraltar and the
service/overseas schools are removed from the corpus at the mart boundary,
not hidden at the view layer. `dim_school` and `dim_location` both exclude
`TypeOfEstablishment` in {25, 26, 30, 37} — Offshore schools, Service
children's education, Welsh establishment, British schools overseas —
listed once as `vars.non_england_school_type_codes` in `dbt_project.yml`.
That removes 2,036 establishments and 29 local authorities, and because
`build_sitemap()` reads the same marts, it drops them from the sitemap in
the same stroke. `assert_england_only_schools` fails the pipeline if a GIAS
refresh reintroduces them or if the two models drift apart.
5. Prune the remaining thin pages: exclude any school with no performance data
**and** no Ofsted record. Distinct from item 4 — these are English schools
with nothing yet to show, so the fix may be a better template rather than
removal.
6. Rebuild the sitemap as a sitemap **index**: one child per page family,
real `lastmod` from the data-load timestamp, `priority` and `changefreq`
dropped.
### W2 — The location layer
The dominant lever. `dim_location` already carries `town`, `county`,
`local_authority_name`, `parliamentary_constituency`, `latitude`,
`longitude` and `postcode`, so no new ingestion is required.
Routes:
/schools/[la] e.g. /schools/manchester
/schools/[la]/primary
/schools/[la]/secondary
/best-primary-schools/[town]
/best-secondary-schools/[town]
/schools/near/[outcode] e.g. /schools/near/m20
/schools/near-me geolocating hub
**Thin-page threshold: generate a town or outcode page only where at least
five schools have current performance data.** Below that, 301 to the parent
LA page. This is the single most important constraint in the workstream —
it is what separates a location layer from index bloat.
Each page must earn its place with content a parent would actually use, not
a template shell:
- H1 matching the query intent ("Best primary schools in Manchester")
- Counts framed usefully: "137 primary schools, 9 rated Outstanding"
- A ranked table of the top 20 on the headline metric
- Local average against the England average
- Ofsted grade distribution
- Map
- Links to neighbouring towns and to the parent LA
- An FAQ block (feeds `FAQPage` in W4)
- Links to every school page in scope — this is what de-orphans W1's corpus
Sizing estimate: ~150 usable LAs × 3 ≈ 450; towns clearing the threshold
≈ 1,200 × 2 ≈ 2,400; outcodes ≈ 2,300. Roughly **5,000 new pages**,
comfortably inside a sitemap index and well under the per-file 50,000 limit.
### W3 — Make rankings indexable
`/rankings` is driven entirely by query params, so Google indexes
approximately one page where there should be hundreds.
/rankings/[phase]/[metric]
/rankings/[phase]/[metric]/[la]
The interactive filter UI stays; its state moves into real paths. Param
forms canonicalise to the clean path. This is the direct play for C2.
### W4 — Structured data and internal linking
- Replace the bare `EducationalOrganization` on school pages with `School`,
and populate it properly.
- `BreadcrumbList` site-wide.
- `ItemList` on every rankings and location page.
- `FAQPage` on admissions and on location pages.
- New school-page modules: "Other schools in [town]", "Nearby schools",
"Compare with similar schools". Each links out to W2 and W3 pages, which
is what circulates authority instead of stranding it.
Explicitly **not** doing `Dataset` or `AggregateRating` — no review corpus
exists, and fabricating one would be both useless and a policy violation.
### W5 — Admissions expansion
One static page currently carries an entire cluster.
/admissions/[la] per-authority deadlines and offer day
/admissions/appeals
/admissions/national-offer-day
`school admissions [LA]` is high-intent and highly seasonal; per-authority
pages are the natural unit.
### W6 — Explainer content
/guides/progress-8
/guides/attainment-8
/guides/sats-scaled-scores
/guides/ofsted-grades
/guides/expected-standard
Each links into the corresponding W3 rankings page. Cheap to build, and it
is what gives the metric vocabulary enough topical weight to support C1–C4.
### W7 — Head-to-head pages
/compare/[school-a]-vs-[school-b]
C7 is uncontested and native to the product. It is also the easiest way to
destroy everything W1 fixes: 27,229 schools generate 370 million pairs.
**Curated pairs only** — same town, both with current data, both with real
search demand — capped in the low thousands. Gated behind evidence that W2
is indexing cleanly.
### W8 — Metadata rewrite for C1
Current homepage title is `schoolcompare | Compare every school in England`,
which spends the most valuable position on the brand. Rewrite the homepage,
rankings and compare titles and descriptions around C1 phrasing. Small
change, and the cheapest item in the programme.
## Sequencing
W0 → W1 → W2 → W3 → W4 → W5 → W6 → (W7 if W2 indexes cleanly)
W1 before W2 is not negotiable: adding pages to a corpus that is 22% thin
compounds the problem rather than diluting it.
This spec is a programme, not a single implementation plan. Each workstream
gets its own plan and its own PR; W2 will likely need several. Only W0 and W1
are ready to plan against today — the rest should be re-read after W0's
Search Console baseline lands, because that data may reorder them.
## Testing
Per CLAUDE.md, user-facing behaviour changes extend the `e2e/` journeys in
the same PR. Each workstream adds:
- W1: canonical present and correct on every route; `/compare?urns=` carries
`noindex`; sitemap excludes a known dataless URN.
- W2: a known LA, town and outcode page renders with the expected school
count; a below-threshold town redirects to its LA.
- W3: a clean rankings path renders; the param form canonicalises to it.
- W4: JSON-LD parses and validates against the declared types.
## Risks
**Index bloat.** The failure mode of every programmatic SEO programme. The
five-school threshold, the W1 prune and the W7 gate are the three controls.
**Helpful-content exposure.** Google's stance on templated location pages
has hardened. The mitigation is that each page carries genuinely local
computed data — real counts, real distributions, real local-vs-national
comparison — rather than a name substituted into boilerplate.
**Build cost.** School pages already use ISR with a 7-day revalidate and
`PRERENDER_SCHOOLS` gating full prerender. 5,000 more routes need the same
treatment; full static generation of 32,000 pages is likely impractical in
CI.
**Seasonality.** Results day and offer day dominate the traffic curve.
Judging the programme on a mid-summer window would misread it in either
direction. W0's baseline must be year-on-year, not month-on-month.
## Open questions
1. Catchment areas are Locrating's moat and a strong C3 driver
(`[postcode] school catchment`). `fact_admissions` carries admission
distances. Is deriving approximate catchment a later workstream, or out
of scope?
+101
View File
@@ -1473,3 +1473,104 @@ test('no single section dominates the height of a school page', async ({ page })
+ sections.map((s) => `${s.id}=${s.h}`).join(', '),
).toBeLessThan(2.5);
});
/*
* England-only corpus.
*
* GIAS ships the whole UK plus overseas and offshore establishments, none of
* which carry comparable DfE performance data. dim_school and dim_location
* exclude them (vars.non_england_school_type_codes in dbt_project.yml), which
* keeps them out of the site, the filter lists and the sitemap together.
*
* These assert the published surface, not the warehouse: the dbt test
* assert_england_only_schools guards the marts, and these guard what the
* environment actually serves once the marts have been rebuilt.
*/
const WELSH_AUTHORITIES = [
'Blaenau Gwent', 'Bridgend', 'Caerphilly', 'Cardiff', 'Carmarthenshire',
'Ceredigion', 'Conwy', 'Denbighshire', 'Flintshire', 'Gwynedd',
'Isle of Anglesey', 'Merthyr Tydfil', 'Monmouthshire', 'Neath Port Talbot',
'Newport', 'Pembrokeshire', 'Powys', 'Rhondda Cynon Taf', 'Swansea',
'Torfaen', 'Vale of Glamorgan', 'Wrexham',
];
const NON_ENGLAND_AUTHORITIES = [
'BFPO Overseas Establishments', 'Fieldwork Overseas Establishments',
'Gibraltar Overseas Establishments', 'Guernsey Offshore Establishments',
'Isle of Man Offshore Establishments', 'Jersey Offshore Establishments',
'Scotland Offshore Establishments',
];
const NON_ENGLAND_TYPES = [
'Welsh establishment', 'Offshore schools',
"Service children's education", 'British schools overseas',
];
test('the local authority filter offers no Welsh or overseas authority', async ({ page }) => {
const res = await page.request.get('/api/filters');
expect(res.ok()).toBeTruthy();
const { local_authorities: las } = await res.json();
expect(Array.isArray(las)).toBeTruthy();
// Guards against the list being empty, which would pass the check below
// for the wrong reason.
expect(las.length).toBeGreaterThan(100);
const leaked = [...WELSH_AUTHORITIES, ...NON_ENGLAND_AUTHORITIES]
.filter((la) => las.includes(la));
expect(leaked, `non-England authorities still offered: ${leaked.join(', ')}`)
.toEqual([]);
});
test('the school type filter offers no non-England establishment type', async ({ page }) => {
const res = await page.request.get('/api/filters');
expect(res.ok()).toBeTruthy();
const { school_types: types } = await res.json();
expect(Array.isArray(types)).toBeTruthy();
expect(types.length).toBeGreaterThan(10);
const leaked = NON_ENGLAND_TYPES.filter((t) => types.includes(t));
expect(leaked, `non-England types still offered: ${leaked.join(', ')}`)
.toEqual([]);
});
test('searching a Welsh authority by name returns no schools', async ({ page }) => {
// Cardiff held 144 Welsh establishments and nothing else, so the authority
// should now be absent from the corpus entirely rather than merely thinned.
const res = await page.request.get('/api/schools?local_authority=Cardiff&page_size=1');
expect(res.ok()).toBeTruthy();
const body = await res.json();
expect(body.total ?? (body.schools ?? []).length).toBe(0);
});
test('a Welsh school URL 404s while an English one still resolves', async ({ page }) => {
// Paired on purpose: the Welsh assertion alone would also pass if the whole
// site were down, which is the failure this test most needs to distinguish.
const english = await page.request.get('/api/schools?search=primary&per_page=1');
expect(english.ok()).toBeTruthy();
const [first] = (await english.json()).schools ?? [];
expect(first, 'no English school available to compare against').toBeTruthy();
const good = await page.goto(`/school/${first.urn}-x`);
expect(good?.status(), 'an English school should still resolve').toBeLessThan(400);
// Adamsdown Primary School, Cardiff — a Welsh establishment (URN 401559).
const welsh = await page.goto('/school/401559-adamsdown-primary-school');
expect(welsh?.status(), 'a Welsh school should no longer resolve').toBe(404);
});
test('the sitemap submits no Welsh or overseas school', async ({ page }) => {
const res = await page.request.get('/sitemap.xml');
expect(res.ok()).toBeTruthy();
const xml = await res.text();
const urlCount = (xml.match(/<url>/g) ?? []).length;
expect(urlCount, 'sitemap looks empty or truncated').toBeGreaterThan(1000);
// 401559 (Cardiff) and 402426 (ACT Schools, Cardiff) were both submitted
// before the England-only filter landed.
expect(xml).not.toContain('/school/401559');
expect(xml).not.toContain('/school/402426');
});
+13
View File
@@ -11,6 +11,19 @@ seed-paths: ["seeds"]
target-path: "target"
clean-targets: ["target", "dbt_packages"]
# schoolcompare publishes England only. GIAS ships the whole UK plus overseas
# and offshore establishments, none of which have comparable DfE performance
# data — Wales does not publish on the English measures at all. Excluding them
# at the mart boundary keeps them out of the site, the filter lists and the
# sitemap together. Codes are TypeOfEstablishment, see seeds/gias_code_names.csv.
# 25 = Offshore schools (Jersey, Guernsey, Isle of Man, Gibraltar)
# 26 = Service children's education (BFPO, Fieldwork Overseas)
# 30 = Welsh establishment
# 37 = British schools overseas
# dim_school, dim_location and assert_england_only_schools all read this list.
vars:
non_england_school_type_codes: [25, 26, 30, 37]
models:
school_compare:
staging:
@@ -31,5 +31,9 @@ select
else null
end as longitude
from {{ ref('stg_gias_establishments') }} s
-- Must match dim_school's status filter exactly (the API inner-joins the two).
-- Must match dim_school's status and England filters exactly (the API
-- inner-joins the two).
where s.status_code in (1, 3)
-- coalesce, not a bare NOT IN: a null type code would make the predicate
-- null and drop the row silently. Unknown type is not grounds for exclusion.
and coalesce(s.school_type_code, -1) not in ({{ var('non_england_school_type_codes') | join(', ') }})
@@ -91,3 +91,8 @@ left join {{ ref('int_ofsted_latest') }} o on s.urn = o.urn
-- 1 = Open; 3 = Open, but proposed to close (still operating; drops out when
-- GIAS flips to Closed — marts fully rebuild each run).
where s.status_code in (1, 3)
-- England only. dim_location must apply this filter identically (the API
-- inner-joins the two). See vars in dbt_project.yml for what the codes are.
-- coalesce, not a bare NOT IN: a null type code would make the predicate
-- null and drop the row silently. Unknown type is not grounds for exclusion.
and coalesce(s.school_type_code, -1) not in ({{ var('non_england_school_type_codes') | join(', ') }})
@@ -0,0 +1,28 @@
-- Custom test: the published corpus is England only.
--
-- GIAS ships the whole UK plus overseas and offshore establishments. None of
-- them carry comparable DfE performance data, so dim_school and dim_location
-- filter them out (see vars.non_england_school_type_codes in dbt_project.yml).
-- This test is the guard: a GIAS refresh that reintroduces them, or an edit
-- that drops the filter from one of the two models, fails the pipeline rather
-- than quietly republishing ~2,000 dataless pages to the site and the sitemap.
select
urn,
school_name,
school_type_code
from {{ ref('dim_school') }}
where school_type_code in ({{ var('non_england_school_type_codes') | join(', ') }})
union all
-- dim_location is inner-joined to dim_school by the API, so it must filter
-- identically. Anything here that dim_school does not have means the two
-- models have drifted apart.
select
l.urn,
null::text as school_name,
null::integer as school_type_code
from {{ ref('dim_location') }} l
left join {{ ref('dim_school') }} s on l.urn = s.urn
where s.urn is null