Merge pull request 'feat(data): publish England only, dropping Welsh and overseas establishments' (#107) from feat/england-only-corpus into main
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m15s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m27s
Stage (build -> staging -> E2E gate) / Build Backend (FastAPI) (push) Successful in 13s
Stage (build -> staging -> E2E gate) / Build Frontend (Next.js) (push) Successful in 51s
Stage (build -> staging -> E2E gate) / Build Pipeline (Meltano + dbt + Airflow) (push) Successful in 1m15s
Stage (build -> staging -> E2E gate) / Deploy to Staging (push) Successful in 1s
Stage (build -> staging -> E2E gate) / E2E Journeys against Staging (push) Failing after 1m27s
Reviewed-on: #107
This commit was merged in pull request #107.
This commit is contained in:
commit
c47fe38971
7 files changed
+1716
-1
No files matched your search
File diff suppressed because it is too large.
Load diff
@@ -0,0 +1,309 @@
|
||||
# SEO Programme — Design
|
||||
|
||||
Date: 2026-08-20
|
||||
Status: awaiting review
|
||||
|
||||
## Problem
|
||||
|
||||
schoolcompare ranks second for "school compare" — an exact match for the
|
||||
brand and the domain. It ranks poorly for "compare schools", "school
|
||||
comparison" and "schools near me". The first is a naming artefact and
|
||||
transfers to nothing; the rest are the queries that actually carry parent
|
||||
demand.
|
||||
|
||||
The cause is structural, not editorial. The site publishes five route
|
||||
families:
|
||||
|
||||
/ /compare /rankings /admissions /school/[slug]
|
||||
|
||||
Location intent has no landing page at all. Every competitor outranking us
|
||||
on those queries wins with programmatic location pages:
|
||||
|
||||
| Competitor | URL pattern |
|
||||
|---------------|--------------------------------------------|
|
||||
| School Guide | `/best-schools-in/manchester` |
|
||||
| Locrating | `/the-best-primary-schools-in-Manchester_…`|
|
||||
| FindMySchool | `/best-primary-schools/manchester` |
|
||||
| Snobe | `/best-primary-schools/manchester` |
|
||||
| School Atlas | `/guides/best-primary-schools-manchester` |
|
||||
|
||||
"Schools near me" is a local-intent query. Google resolves it against the
|
||||
user's coordinates and serves pages that are *about a place*. A national
|
||||
homepage cannot win it. No title or description change fixes this; only
|
||||
pages Google can localise will.
|
||||
|
||||
## Baseline (measured 2026-08-20, production API)
|
||||
|
||||
| Measure | Value |
|
||||
|--------------------------------------------|---------|
|
||||
| Unique schools | 27,229 |
|
||||
| URLs in sitemap.xml | 27,232 |
|
||||
| Schools with 2024/25 performance data | 21,266 |
|
||||
| Schools with **no** current performance data| ~5,963 (22%) |
|
||||
| Welsh establishments (all metrics null) | 1,569 |
|
||||
| Overseas / offshore establishments | 467 |
|
||||
| Static URLs in sitemap | 3 |
|
||||
| Routes setting a canonical | 1 of 5 |
|
||||
|
||||
Three findings from that table drive the plan.
|
||||
|
||||
**We submit ~6,000 thin pages to Google.** `build_sitemap()`
|
||||
(`backend/app.py:73`) enumerates every URN regardless of whether the school
|
||||
has any data. Welsh establishments return `school_type: "Welsh
|
||||
establishment"` with every performance metric, Ofsted grade and phase field
|
||||
null. Overseas and offshore establishments ("BFPO Overseas Establishments",
|
||||
"Gibraltar Overseas Establishments", "Jersey Offshore Establishments") are
|
||||
in the local-authority list too. At 22% of the submitted corpus this is a
|
||||
site-wide quality signal problem and a crawl-budget waste, not a rounding
|
||||
error.
|
||||
|
||||
W1 item 4 removes 2,036 of those — every non-England establishment — taking
|
||||
the corpus to 25,193. The 3,927 that remain are English schools with no
|
||||
current data: mostly newly opened, special, nursery or alternative provision.
|
||||
Those are a template problem, not a corpus problem, and item 5 handles them
|
||||
separately.
|
||||
|
||||
**The homepage is its own competitor.** `app/page.tsx` accepts eleven search
|
||||
params (`search`, `local_authority`, `school_type`, `phase`, `page`,
|
||||
`postcode`, `radius`, `sort`, `gender`, `admissions_policy`,
|
||||
`has_sixth_form`) and sets no canonical. Every filter combination is a
|
||||
crawlable near-duplicate of the single page we are asking to rank for
|
||||
"compare schools".
|
||||
|
||||
**School pages are near-orphans.** Reachable from the sitemap and from site
|
||||
search, but almost nothing links to them contextually, so they accrue no
|
||||
internal authority.
|
||||
|
||||
Also noted: the sitemap emits invented `priority` values and no `lastmod`.
|
||||
Google ignores `priority` and `changefreq` entirely; `lastmod` is the field
|
||||
it does read, and we omit it.
|
||||
|
||||
## Keyword clusters
|
||||
|
||||
Ranked by judgement of UK parent search behaviour and by the competitive
|
||||
SERP evidence above. Google Search Console is connected, so cluster
|
||||
priorities are to be re-derived from measured impressions before build
|
||||
starts (see Workstream 0).
|
||||
|
||||
**C1 — Head "compare" terms.** compare schools · school comparison · school
|
||||
comparison tool · compare school performance · compare primary schools ·
|
||||
compare secondary schools · compare two schools
|
||||
|
||||
**C2 — League tables and rankings.** primary school league tables ·
|
||||
secondary school league tables · school league tables 2026 · SATs results by
|
||||
school · GCSE results by school · KS2 league tables · Progress 8 rankings ·
|
||||
best primary schools in [town] · top 10 primary schools in [LA]
|
||||
|
||||
**C3 — Local / near me.** schools near me · primary schools near me ·
|
||||
secondary schools near me · best schools near me · good schools near me ·
|
||||
schools in [town] · primary schools in [LA] · schools near [postcode] ·
|
||||
[postcode] school catchment
|
||||
|
||||
**C4 — Individual school long tail.** [school] ofsted · [school] SATs
|
||||
results · [school] catchment area · [school] reviews · [school] URN
|
||||
|
||||
**C5 — Admissions.** primary school admissions 2027 · national offer day
|
||||
2027 · school application deadline · school admissions appeal ·
|
||||
oversubscription criteria · distance criteria school admissions · didn't get
|
||||
first choice school · school admissions [LA]
|
||||
|
||||
**C6 — Metric explainers.** what is a good SATs score · what is Progress 8 ·
|
||||
what is Attainment 8 · expected standard KS2 meaning · scaled score
|
||||
explained · Ofsted grades explained · Ofsted report cards · pupil premium
|
||||
explained
|
||||
|
||||
**C7 — Head to head.** [school A] vs [school B] · academy vs community
|
||||
school · grammar school vs comprehensive · faith school vs community school
|
||||
|
||||
## Workstreams
|
||||
|
||||
### W0 — Measure before touching anything
|
||||
|
||||
Export a Google Search Console baseline: impressions, clicks, average
|
||||
position and CTR by query and by page, for the trailing 16 months. Bucket
|
||||
queries into C1–C7. This sets the counterfactual — without it, no later
|
||||
claim about lift is defensible, because school-search traffic is strongly
|
||||
seasonal (results day in December, offer day in March/April).
|
||||
|
||||
Re-rank C1–C7 against measured impressions and adjust the sequence below if
|
||||
the data disagrees with the judgement calls.
|
||||
|
||||
### W1 — Crawl hygiene and index sanity
|
||||
|
||||
Cheap, and it unblocks everything after it. Adding 5,000 pages on top of a
|
||||
corpus that is 22% thin would compound the existing problem.
|
||||
|
||||
1. Canonical on every route. `/`, `/rankings`, `/compare` and `/admissions`
|
||||
currently set none.
|
||||
2. The homepage canonicalises to `/` regardless of search params.
|
||||
3. `/compare?urns=…` gets `noindex, follow` — it is an unbounded parameter
|
||||
space with no standalone value.
|
||||
4. **England only — DONE.** Wales, the Crown Dependencies, Gibraltar and the
|
||||
service/overseas schools are removed from the corpus at the mart boundary,
|
||||
not hidden at the view layer. `dim_school` and `dim_location` both exclude
|
||||
`TypeOfEstablishment` in {25, 26, 30, 37} — Offshore schools, Service
|
||||
children's education, Welsh establishment, British schools overseas —
|
||||
listed once as `vars.non_england_school_type_codes` in `dbt_project.yml`.
|
||||
That removes 2,036 establishments and 29 local authorities, and because
|
||||
`build_sitemap()` reads the same marts, it drops them from the sitemap in
|
||||
the same stroke. `assert_england_only_schools` fails the pipeline if a GIAS
|
||||
refresh reintroduces them or if the two models drift apart.
|
||||
5. Prune the remaining thin pages: exclude any school with no performance data
|
||||
**and** no Ofsted record. Distinct from item 4 — these are English schools
|
||||
with nothing yet to show, so the fix may be a better template rather than
|
||||
removal.
|
||||
6. Rebuild the sitemap as a sitemap **index**: one child per page family,
|
||||
real `lastmod` from the data-load timestamp, `priority` and `changefreq`
|
||||
dropped.
|
||||
|
||||
### W2 — The location layer
|
||||
|
||||
The dominant lever. `dim_location` already carries `town`, `county`,
|
||||
`local_authority_name`, `parliamentary_constituency`, `latitude`,
|
||||
`longitude` and `postcode`, so no new ingestion is required.
|
||||
|
||||
Routes:
|
||||
|
||||
/schools/[la] e.g. /schools/manchester
|
||||
/schools/[la]/primary
|
||||
/schools/[la]/secondary
|
||||
/best-primary-schools/[town]
|
||||
/best-secondary-schools/[town]
|
||||
/schools/near/[outcode] e.g. /schools/near/m20
|
||||
/schools/near-me geolocating hub
|
||||
|
||||
**Thin-page threshold: generate a town or outcode page only where at least
|
||||
five schools have current performance data.** Below that, 301 to the parent
|
||||
LA page. This is the single most important constraint in the workstream —
|
||||
it is what separates a location layer from index bloat.
|
||||
|
||||
Each page must earn its place with content a parent would actually use, not
|
||||
a template shell:
|
||||
|
||||
- H1 matching the query intent ("Best primary schools in Manchester")
|
||||
- Counts framed usefully: "137 primary schools, 9 rated Outstanding"
|
||||
- A ranked table of the top 20 on the headline metric
|
||||
- Local average against the England average
|
||||
- Ofsted grade distribution
|
||||
- Map
|
||||
- Links to neighbouring towns and to the parent LA
|
||||
- An FAQ block (feeds `FAQPage` in W4)
|
||||
- Links to every school page in scope — this is what de-orphans W1's corpus
|
||||
|
||||
Sizing estimate: ~150 usable LAs × 3 ≈ 450; towns clearing the threshold
|
||||
≈ 1,200 × 2 ≈ 2,400; outcodes ≈ 2,300. Roughly **5,000 new pages**,
|
||||
comfortably inside a sitemap index and well under the per-file 50,000 limit.
|
||||
|
||||
### W3 — Make rankings indexable
|
||||
|
||||
`/rankings` is driven entirely by query params, so Google indexes
|
||||
approximately one page where there should be hundreds.
|
||||
|
||||
/rankings/[phase]/[metric]
|
||||
/rankings/[phase]/[metric]/[la]
|
||||
|
||||
The interactive filter UI stays; its state moves into real paths. Param
|
||||
forms canonicalise to the clean path. This is the direct play for C2.
|
||||
|
||||
### W4 — Structured data and internal linking
|
||||
|
||||
- Replace the bare `EducationalOrganization` on school pages with `School`,
|
||||
and populate it properly.
|
||||
- `BreadcrumbList` site-wide.
|
||||
- `ItemList` on every rankings and location page.
|
||||
- `FAQPage` on admissions and on location pages.
|
||||
- New school-page modules: "Other schools in [town]", "Nearby schools",
|
||||
"Compare with similar schools". Each links out to W2 and W3 pages, which
|
||||
is what circulates authority instead of stranding it.
|
||||
|
||||
Explicitly **not** doing `Dataset` or `AggregateRating` — no review corpus
|
||||
exists, and fabricating one would be both useless and a policy violation.
|
||||
|
||||
### W5 — Admissions expansion
|
||||
|
||||
One static page currently carries an entire cluster.
|
||||
|
||||
/admissions/[la] per-authority deadlines and offer day
|
||||
/admissions/appeals
|
||||
/admissions/national-offer-day
|
||||
|
||||
`school admissions [LA]` is high-intent and highly seasonal; per-authority
|
||||
pages are the natural unit.
|
||||
|
||||
### W6 — Explainer content
|
||||
|
||||
/guides/progress-8
|
||||
/guides/attainment-8
|
||||
/guides/sats-scaled-scores
|
||||
/guides/ofsted-grades
|
||||
/guides/expected-standard
|
||||
|
||||
Each links into the corresponding W3 rankings page. Cheap to build, and it
|
||||
is what gives the metric vocabulary enough topical weight to support C1–C4.
|
||||
|
||||
### W7 — Head-to-head pages
|
||||
|
||||
/compare/[school-a]-vs-[school-b]
|
||||
|
||||
C7 is uncontested and native to the product. It is also the easiest way to
|
||||
destroy everything W1 fixes: 27,229 schools generate 370 million pairs.
|
||||
**Curated pairs only** — same town, both with current data, both with real
|
||||
search demand — capped in the low thousands. Gated behind evidence that W2
|
||||
is indexing cleanly.
|
||||
|
||||
### W8 — Metadata rewrite for C1
|
||||
|
||||
Current homepage title is `schoolcompare | Compare every school in England`,
|
||||
which spends the most valuable position on the brand. Rewrite the homepage,
|
||||
rankings and compare titles and descriptions around C1 phrasing. Small
|
||||
change, and the cheapest item in the programme.
|
||||
|
||||
## Sequencing
|
||||
|
||||
W0 → W1 → W2 → W3 → W4 → W5 → W6 → (W7 if W2 indexes cleanly)
|
||||
|
||||
W1 before W2 is not negotiable: adding pages to a corpus that is 22% thin
|
||||
compounds the problem rather than diluting it.
|
||||
|
||||
This spec is a programme, not a single implementation plan. Each workstream
|
||||
gets its own plan and its own PR; W2 will likely need several. Only W0 and W1
|
||||
are ready to plan against today — the rest should be re-read after W0's
|
||||
Search Console baseline lands, because that data may reorder them.
|
||||
|
||||
## Testing
|
||||
|
||||
Per CLAUDE.md, user-facing behaviour changes extend the `e2e/` journeys in
|
||||
the same PR. Each workstream adds:
|
||||
|
||||
- W1: canonical present and correct on every route; `/compare?urns=` carries
|
||||
`noindex`; sitemap excludes a known dataless URN.
|
||||
- W2: a known LA, town and outcode page renders with the expected school
|
||||
count; a below-threshold town redirects to its LA.
|
||||
- W3: a clean rankings path renders; the param form canonicalises to it.
|
||||
- W4: JSON-LD parses and validates against the declared types.
|
||||
|
||||
## Risks
|
||||
|
||||
**Index bloat.** The failure mode of every programmatic SEO programme. The
|
||||
five-school threshold, the W1 prune and the W7 gate are the three controls.
|
||||
|
||||
**Helpful-content exposure.** Google's stance on templated location pages
|
||||
has hardened. The mitigation is that each page carries genuinely local
|
||||
computed data — real counts, real distributions, real local-vs-national
|
||||
comparison — rather than a name substituted into boilerplate.
|
||||
|
||||
**Build cost.** School pages already use ISR with a 7-day revalidate and
|
||||
`PRERENDER_SCHOOLS` gating full prerender. 5,000 more routes need the same
|
||||
treatment; full static generation of 32,000 pages is likely impractical in
|
||||
CI.
|
||||
|
||||
**Seasonality.** Results day and offer day dominate the traffic curve.
|
||||
Judging the programme on a mid-summer window would misread it in either
|
||||
direction. W0's baseline must be year-on-year, not month-on-month.
|
||||
|
||||
## Open questions
|
||||
|
||||
1. Catchment areas are Locrating's moat and a strong C3 driver
|
||||
(`[postcode] school catchment`). `fact_admissions` carries admission
|
||||
distances. Is deriving approximate catchment a later workstream, or out
|
||||
of scope?
|
||||
@@ -1473,3 +1473,104 @@ test('no single section dominates the height of a school page', async ({ page })
|
||||
+ sections.map((s) => `${s.id}=${s.h}`).join(', '),
|
||||
).toBeLessThan(2.5);
|
||||
});
|
||||
|
||||
/*
|
||||
* England-only corpus.
|
||||
*
|
||||
* GIAS ships the whole UK plus overseas and offshore establishments, none of
|
||||
* which carry comparable DfE performance data. dim_school and dim_location
|
||||
* exclude them (vars.non_england_school_type_codes in dbt_project.yml), which
|
||||
* keeps them out of the site, the filter lists and the sitemap together.
|
||||
*
|
||||
* These assert the published surface, not the warehouse: the dbt test
|
||||
* assert_england_only_schools guards the marts, and these guard what the
|
||||
* environment actually serves once the marts have been rebuilt.
|
||||
*/
|
||||
|
||||
const WELSH_AUTHORITIES = [
|
||||
'Blaenau Gwent', 'Bridgend', 'Caerphilly', 'Cardiff', 'Carmarthenshire',
|
||||
'Ceredigion', 'Conwy', 'Denbighshire', 'Flintshire', 'Gwynedd',
|
||||
'Isle of Anglesey', 'Merthyr Tydfil', 'Monmouthshire', 'Neath Port Talbot',
|
||||
'Newport', 'Pembrokeshire', 'Powys', 'Rhondda Cynon Taf', 'Swansea',
|
||||
'Torfaen', 'Vale of Glamorgan', 'Wrexham',
|
||||
];
|
||||
|
||||
const NON_ENGLAND_AUTHORITIES = [
|
||||
'BFPO Overseas Establishments', 'Fieldwork Overseas Establishments',
|
||||
'Gibraltar Overseas Establishments', 'Guernsey Offshore Establishments',
|
||||
'Isle of Man Offshore Establishments', 'Jersey Offshore Establishments',
|
||||
'Scotland Offshore Establishments',
|
||||
];
|
||||
|
||||
const NON_ENGLAND_TYPES = [
|
||||
'Welsh establishment', 'Offshore schools',
|
||||
"Service children's education", 'British schools overseas',
|
||||
];
|
||||
|
||||
test('the local authority filter offers no Welsh or overseas authority', async ({ page }) => {
|
||||
const res = await page.request.get('/api/filters');
|
||||
expect(res.ok()).toBeTruthy();
|
||||
const { local_authorities: las } = await res.json();
|
||||
|
||||
expect(Array.isArray(las)).toBeTruthy();
|
||||
// Guards against the list being empty, which would pass the check below
|
||||
// for the wrong reason.
|
||||
expect(las.length).toBeGreaterThan(100);
|
||||
|
||||
const leaked = [...WELSH_AUTHORITIES, ...NON_ENGLAND_AUTHORITIES]
|
||||
.filter((la) => las.includes(la));
|
||||
expect(leaked, `non-England authorities still offered: ${leaked.join(', ')}`)
|
||||
.toEqual([]);
|
||||
});
|
||||
|
||||
test('the school type filter offers no non-England establishment type', async ({ page }) => {
|
||||
const res = await page.request.get('/api/filters');
|
||||
expect(res.ok()).toBeTruthy();
|
||||
const { school_types: types } = await res.json();
|
||||
|
||||
expect(Array.isArray(types)).toBeTruthy();
|
||||
expect(types.length).toBeGreaterThan(10);
|
||||
|
||||
const leaked = NON_ENGLAND_TYPES.filter((t) => types.includes(t));
|
||||
expect(leaked, `non-England types still offered: ${leaked.join(', ')}`)
|
||||
.toEqual([]);
|
||||
});
|
||||
|
||||
test('searching a Welsh authority by name returns no schools', async ({ page }) => {
|
||||
// Cardiff held 144 Welsh establishments and nothing else, so the authority
|
||||
// should now be absent from the corpus entirely rather than merely thinned.
|
||||
const res = await page.request.get('/api/schools?local_authority=Cardiff&page_size=1');
|
||||
expect(res.ok()).toBeTruthy();
|
||||
const body = await res.json();
|
||||
expect(body.total ?? (body.schools ?? []).length).toBe(0);
|
||||
});
|
||||
|
||||
test('a Welsh school URL 404s while an English one still resolves', async ({ page }) => {
|
||||
// Paired on purpose: the Welsh assertion alone would also pass if the whole
|
||||
// site were down, which is the failure this test most needs to distinguish.
|
||||
const english = await page.request.get('/api/schools?search=primary&per_page=1');
|
||||
expect(english.ok()).toBeTruthy();
|
||||
const [first] = (await english.json()).schools ?? [];
|
||||
expect(first, 'no English school available to compare against').toBeTruthy();
|
||||
|
||||
const good = await page.goto(`/school/${first.urn}-x`);
|
||||
expect(good?.status(), 'an English school should still resolve').toBeLessThan(400);
|
||||
|
||||
// Adamsdown Primary School, Cardiff — a Welsh establishment (URN 401559).
|
||||
const welsh = await page.goto('/school/401559-adamsdown-primary-school');
|
||||
expect(welsh?.status(), 'a Welsh school should no longer resolve').toBe(404);
|
||||
});
|
||||
|
||||
test('the sitemap submits no Welsh or overseas school', async ({ page }) => {
|
||||
const res = await page.request.get('/sitemap.xml');
|
||||
expect(res.ok()).toBeTruthy();
|
||||
const xml = await res.text();
|
||||
|
||||
const urlCount = (xml.match(/<url>/g) ?? []).length;
|
||||
expect(urlCount, 'sitemap looks empty or truncated').toBeGreaterThan(1000);
|
||||
|
||||
// 401559 (Cardiff) and 402426 (ACT Schools, Cardiff) were both submitted
|
||||
// before the England-only filter landed.
|
||||
expect(xml).not.toContain('/school/401559');
|
||||
expect(xml).not.toContain('/school/402426');
|
||||
});
|
||||
@@ -11,6 +11,19 @@ seed-paths: ["seeds"]
|
||||
target-path: "target"
|
||||
clean-targets: ["target", "dbt_packages"]
|
||||
|
||||
# schoolcompare publishes England only. GIAS ships the whole UK plus overseas
|
||||
# and offshore establishments, none of which have comparable DfE performance
|
||||
# data — Wales does not publish on the English measures at all. Excluding them
|
||||
# at the mart boundary keeps them out of the site, the filter lists and the
|
||||
# sitemap together. Codes are TypeOfEstablishment, see seeds/gias_code_names.csv.
|
||||
# 25 = Offshore schools (Jersey, Guernsey, Isle of Man, Gibraltar)
|
||||
# 26 = Service children's education (BFPO, Fieldwork Overseas)
|
||||
# 30 = Welsh establishment
|
||||
# 37 = British schools overseas
|
||||
# dim_school, dim_location and assert_england_only_schools all read this list.
|
||||
vars:
|
||||
non_england_school_type_codes: [25, 26, 30, 37]
|
||||
|
||||
models:
|
||||
school_compare:
|
||||
staging:
|
||||
|
||||
@@ -31,5 +31,9 @@ select
|
||||
else null
|
||||
end as longitude
|
||||
from {{ ref('stg_gias_establishments') }} s
|
||||
-- Must match dim_school's status filter exactly (the API inner-joins the two).
|
||||
-- Must match dim_school's status and England filters exactly (the API
|
||||
-- inner-joins the two).
|
||||
where s.status_code in (1, 3)
|
||||
-- coalesce, not a bare NOT IN: a null type code would make the predicate
|
||||
-- null and drop the row silently. Unknown type is not grounds for exclusion.
|
||||
and coalesce(s.school_type_code, -1) not in ({{ var('non_england_school_type_codes') | join(', ') }})
|
||||
@@ -91,3 +91,8 @@ left join {{ ref('int_ofsted_latest') }} o on s.urn = o.urn
|
||||
-- 1 = Open; 3 = Open, but proposed to close (still operating; drops out when
|
||||
-- GIAS flips to Closed — marts fully rebuild each run).
|
||||
where s.status_code in (1, 3)
|
||||
-- England only. dim_location must apply this filter identically (the API
|
||||
-- inner-joins the two). See vars in dbt_project.yml for what the codes are.
|
||||
-- coalesce, not a bare NOT IN: a null type code would make the predicate
|
||||
-- null and drop the row silently. Unknown type is not grounds for exclusion.
|
||||
and coalesce(s.school_type_code, -1) not in ({{ var('non_england_school_type_codes') | join(', ') }})
|
||||
@@ -0,0 +1,28 @@
|
||||
-- Custom test: the published corpus is England only.
|
||||
--
|
||||
-- GIAS ships the whole UK plus overseas and offshore establishments. None of
|
||||
-- them carry comparable DfE performance data, so dim_school and dim_location
|
||||
-- filter them out (see vars.non_england_school_type_codes in dbt_project.yml).
|
||||
-- This test is the guard: a GIAS refresh that reintroduces them, or an edit
|
||||
-- that drops the filter from one of the two models, fails the pipeline rather
|
||||
-- than quietly republishing ~2,000 dataless pages to the site and the sitemap.
|
||||
|
||||
select
|
||||
urn,
|
||||
school_name,
|
||||
school_type_code
|
||||
from {{ ref('dim_school') }}
|
||||
where school_type_code in ({{ var('non_england_school_type_codes') | join(', ') }})
|
||||
|
||||
union all
|
||||
|
||||
-- dim_location is inner-joined to dim_school by the API, so it must filter
|
||||
-- identically. Anything here that dim_school does not have means the two
|
||||
-- models have drifted apart.
|
||||
select
|
||||
l.urn,
|
||||
null::text as school_name,
|
||||
null::integer as school_type_code
|
||||
from {{ ref('dim_location') }} l
|
||||
left join {{ ref('dim_school') }} s on l.urn = s.urn
|
||||
where s.urn is null
|
||||
Reference in new issue
Block a user